Short answer: a photograph can help name a food and approximate its visible size. It cannot directly measure the recipe, fat percentage, absorbed oil or ingredients mixed inside the dish. Photo-calorie systems must infer those details, restrict the problem or ask the user to correct the result.
The important question is therefore not “Can AI recognize this plate?” It is “What information produced the calorie number?”
What the best research systems show#
Photo estimation is not one technology. The results change with the camera, number of views, allowed food list and amount of human correction.
In one 2020 study, a depth-sensing iPhone system estimated energy across 48 meals with a mean absolute error of 12.7% Herzig 2020. That is a strong result. It also came from a narrower task than a one-tap consumer promise: the food type was manually selected from a predefined list, food recognition was not tested, and segmentation could be corrected by hand.
A 2025 study tested a different workflow in daily life. The system identified 189 of 220 dishes (86%) and accurately reported 136 of 200 recognized dishes (68%) Sahoo 2025. Users could correct names, ingredients, cooking methods and portions, and add missing dishes. The study measured identification and reporting, not the calorie error of each final meal.
A systematic review reached the broader conclusion: AI systems sometimes matched or exceeded human estimates, but the methods were too different for a pooled accuracy number, and the tools still needed development before stand-alone research or clinical use Shonkoff 2023.
The evidence does not say that photo logging is useless. It says that “one photo” and “accurate meal” are different claims.
What a flat photo can observe#
A picture can provide useful evidence:
- visible food identity: rice, salad, bread, chicken, soup
- rough count: one egg or two, one patty or two
- visible area and shape
- color and surface texture
- context from the plate, bowl or utensils
Depth sensing, a second angle or a known reference object can improve volume estimation. Recognition models can also narrow a long text search to a few likely foods.
Those are observations. The calorie total also depends on information the photo does not contain.
What the image cannot directly measure#
Recipe#
The same curry can use tomato, coconut milk, cream or ghee. A soup can be broth-based or finished with cream. A photograph records the finished surface, not the ingredients that produced it.
Fat percentage and cut#
Lean and fatty ground meat can look alike after cooking. So can different cuts of pork or chicken with fat rendered into the dish. The hidden calories guide shows a verified USDA example: raw 95% lean and 80/20 ground beef differ by 111 kcal per 100 g.
Cooking fat#
Oil can coat vegetables, remain in breading or disappear into rice and sauce. A glossy surface does not reveal how many grams were added or how much remained in the pan.
Ingredients inside or underneath#
Sugar in a glaze, butter in rice, cheese inside a patty and a second layer beneath the sauce may be absent from the image altogether.
Real-world scale#
A single flat image does not contain a reliable physical scale by itself. Plate size, camera distance, angle and occlusion all affect the inferred portion.
Recognition is the start of the workflow#
The Sahoo study makes the gap concrete: 86% of dishes were identified, while 68% were accurately reported after a correction workflow Sahoo 2025.
When a photo result needs repair, the user may still have to decide:
- which dish the image contains
- which ingredients are missing or wrong
- how it was cooked
- whether oil, sauce or sugar was included
- what portion was actually eaten
That correction is not a minor detail outside the method. It supplies the information the image did not contain.
Why one correction factor cannot fix every meal#
The missing variables change from meal to meal. One photo may infer too much dressing; the next may miss absorbed oil. A portion can be close while the recipe is wrong, or the food name can be right while the amount is wrong.
That makes “my camera is usually 15% low” a weak calibration rule. The direction and size of the error depend on which information was missing this time.
What Calk does differently#
Calk begins with the dish rather than an image of it. A meal template contains the choices that commonly change:
- protein or main ingredient
- cooking method
- oil, sauce, sugar, salt and toppings
- portion
Choose the template from the home screen. If the meal is as usual, tap “Eat.” If chicken became salmon, change the protein and the entire meal recalculates.
The method is deterministic: the same choices produce the same estimate. When the meal changes, the changed ingredient is named. There is no recognition result to reverse-engineer before you can edit it.
See how meal templates work or the broader comparison of calorie-tracking methods.
The summary#
A photo is useful evidence for visible food identity and approximate size. It is not a direct measurement of recipe, fat percentage, absorbed oil or hidden ingredients.
Research systems improve the result by adding depth sensing, predefined food selection and manual correction. Calk takes the other route: start from an editable meal, make the important assumptions explicit and reuse it.


