Wisely
Accuracy · Test design

Benchmark methodology

This page explains the test behind Wisely's published accuracy score: what is included, how references are selected, how results are scored, and where the benchmark does not generalise.

Last reviewed August 16, 2026

What the benchmark measures

The benchmark measures the first calorie and macronutrient estimate returned for 70 meals entered as plain text. The same written input is assessed against a chosen reference value. Wisely's headline calorie result is also compared with the top MyFitnessPal search result for that input, using the same references and formula.

The current published result is 90/100 for Wisely calorie accuracy and 71/100 for MyFitnessPal on this test set. Wisely's combined score, which also weights protein, carbohydrates, and fat, is 88/100. Every test row and first-pass result is available on the full accuracy benchmark.

Test-set design

The 70 inputs are divided across six kinds of food-logging task:

  1. single-ingredient foods with defined serving sizes;
  2. multi-ingredient, home-cooked meals;
  3. named restaurant and chain items;
  4. international dishes;
  5. intentionally misspelled inputs; and
  6. foods entered with exact weights or measurements.

Inputs are sent to Wisely's normal food-analysis endpoint as text, without a meal photo and without a manual correction. This isolates text interpretation and allows the same wording to be used for a database-search comparison. It does not directly test Wisely's camera-first flow.

Reference hierarchy

A score is only as useful as its reference. References are selected from the most direct and authoritative source available for the item:

The published result rows expose the reference calories and macros used for each input. The benchmark currently names USDA, official restaurant pages, and government databases including sources covering Japan, Korea, Thailand, the Philippines, Türkiye, FAO, and CARICOM foods.

Scoring

Each calorie result starts with absolute percentage error: 100 × |estimate − reference| / reference. Accuracy is then:

Calorie score = 100 − absolute percentage error

Scores are expressed on a 0–100 scale and clamped at zero so an error greater than 100% cannot produce a negative score. The headline result is the arithmetic mean of the 70 per-meal calorie scores. Higher is better.

Wisely's combined macro score applies the same comparison to each output and weights the components as follows: calories 40%, protein 25%, carbohydrates 20%, and fat 15%. Calories remain the headline measure for the cross-product comparison because the same value can be retrieved and scored for every tested tool.

Comparison procedure

Each MyFitnessPal figure is the top search result shown for the same plain-text input. It is not manually adjusted to improve the match. Wisely also receives no manual correction. This models a first-result, low-friction logging path; it does not test every search strategy, barcode result, premium feature, or correction a skilled user might apply.

Wisely and the benchmark are both published by EthicsLab Limited. This is therefore a company-run benchmark, not a peer-reviewed study or a test commissioned from an independent laboratory. Publishing the inputs, references, outputs, and fixed formula is intended to make the claim inspectable despite that conflict of interest.

Limitations

Corrections and challenges

If a test input, reference, transcription, calculation, or description appears wrong, email support@wisely.fit with the row, the suspected error, and a source. We review the underlying source and calculation. A confirmed substantive error is corrected in the benchmark, any affected aggregate is recalculated, and the page's update date is changed.

We do not remove an unfavourable result merely because a different plausible reference exists. Where food variation makes one value contestable, the preferred response is to explain the choice or update the reference consistently—not to select whichever value improves Wisely's score.

Update policy

The benchmark may be rerun when Wisely's food-analysis model changes materially, when the test harness changes, when a relevant source is corrected, or when a comparison product changes enough to make the published result misleading. A rerun should preserve the same public inputs and scoring formula unless the methodology itself is being revised.

If the test set, reference hierarchy, comparison procedure, or formula changes materially, the methodology and benchmark should say what changed and update their review dates. Current figures should not be presented as directly comparable with older figures when the underlying procedure is materially different.

Read the evidence