GenGrades

Methodology

How GenGrades grades AI tools and models

The GenGrades methodology is a fixed test set per dimension, an anchored 0 to 10 scale, public weights per category, and a rule that nothing is scored by impression. Every model in a category runs the same tests, so a 7 means the same thing on every page, and every result can be reproduced from the prompts and outputs we publish.

Last updated

Seven rules every grade follows

  1. One test set per dimension

    Each dimension has a fixed set of prompts and inputs and a rubric with written anchors. We do not score from a demo or from memory of what a model did last month.

  2. Two layers: model and product

    Model capability is judged on the outputs. Product experience is judged on using the tool: onboarding, pricing, limits, export. A tool page shows both.

  3. Weights per category, totals as weighted means

    A category decides how much each dimension matters and publishes it. No universal champion: the home page labels what each tool is best for, not one winner for everything.

  4. Blind where possible, humans decide

    Outputs are scored with the model name hidden whenever the output itself does not give it away. A large language model may pre-screen, a person sets the final score.

  5. Every score is dated and versioned

    A new version gets a new run and a new score. Old scores stay published, so improvements and regressions are visible.

  6. Every review gives a verdict and a dealbreaker

    One-sentence conclusion, who it is for, who it is not for, and the single flaw most likely to make you walk away. The not-for list is never softened.

  7. Prompts and raw outputs are public

    The exact prompts and the outputs they produced are published, watermarked, so readers can check our work. This is where our credibility comes from.

The 0 to 10 scale

Every rubric is written against these five bands, so a score is a statement about usability, not a ranking position.

BandLabelMeaning
0 to 2UnusableFails the task outright or produces something that cannot be used without redoing it.
3 to 4PoorGets part of the way; the result needs substantial correction before it is usable.
5 to 6AdequateDoes the job for everyday use; visible flaws, but nothing that stops you.
7 to 8GoodReliable, with only minor flaws; what most people would be happy to ship.
9 to 10Best in classAs good as anything available on this dimension at the time of testing.

What we measure

Nine common dimensions apply to every model. Each modality adds its own, with its own test set. Tools are also graded on product experience.

Common to every model

Instruction following
Does the output do what the prompt asked, including counts, relations and constraints?
Output quality
Judged against the rubric for the modality: fidelity, coherence, polish.
Consistency
Variance across repeated runs of the same prompt. A model that is brilliant one time in five scores low here.
Speed
Time to first token or first frame, and time to completion, measured by us.
Price
Per image, second, minute or million tokens, from the vendor price list, with the date we checked it.
Failure rate
Errors, timeouts and refusals of ordinary requests across the test set.
Content policy strictness
How often harmless prompts are refused; both over-blocking and under-blocking count against.
Licensing
Whether outputs can be used commercially and who owns them, from the terms we read.
Privacy
Whether your inputs are used for training by default, and whether you can opt out.

By modality

Image generation
Prompt adherence (object counts, relations, attributes), text rendering, aesthetics and composition, realism and range of styles, anatomy and hands, subject consistency across images, resolution and aspect ratios, reference image and style control.
Image editing
Instruction fidelity (change only what was asked), preservation of untouched regions (identity, background), edge and mask precision, text editing, multi-turn editing, resolution preservation, inpainting and outpainting, layer separation.
Video generation
Motion quality and physical plausibility, temporal consistency (flicker, morphing), prompt adherence, camera control, duration, resolution and frame rate, fidelity to the source image in image-to-video, character consistency across shots, native audio, success rate.
Video editing tools
Feature coverage (cut, captions, background removal, upscaling), quality of the core edits, export formats and limits, processing speed, collaboration.
Speech and text-to-speech
Naturalness, pronunciation of names, numbers and abbreviations, emotion and style control, cloning quality and its consent policy, languages and accents, latency for real-time use, consistency over long texts, control granularity such as SSML.
Audio processing and transcription
Word error rate under accents and noise, speaker diarisation, timestamps, number of languages, speed.
Music
Musicality and structure, prompt adherence (genre, mood, tempo), vocal and lyric accuracy, duration and stem export, originality and copyright safety, commercial licence.
Language models and writing
Instruction following, factuality and hallucination rate, reasoning, long context, writing quality (style control, naturalness, how much it reads like a machine), structured output and tool calling, multilingual ability, context length, throughput and latency.

Product experience, for tools

  • Ease of getting started
  • Workflow fit: plugins, API, integrations
  • Pricing transparency and value
  • Free tier and its limits
  • Reliability
  • Documentation and support
  • Data privacy terms
  • Team features
  • Export and ownership of your work
  • Update cadence

Weights and the total

The total is the weighted mean of the dimension scores, with the weights of the category the entry is graded in. Weights are published on each category page and can be changed, which recomputes every total in that category. Below is the set used for image generation as an example; the category page is authoritative.

DimensionWeight
Prompt adherence25%
Output quality25%
Price15%
Text rendering10%
Consistency10%
Speed10%
Licensing5%

Re-testing, freshness and what we refuse to do

New version, new run

When a model ships a new version we re-run its test set and publish a new, dated score. Old scores stay. Each re-run has a cost cap.

Facts carry a date

Every price, limit or claim about a vendor shows its source and the day we checked it. A stale fact is worse than none, so the date is always visible.

No invented numbers

No aggregate rating until real users have voted. No score without a test behind it. No grade is for sale, and affiliate links, where they exist, are marked and change nothing.

Conflicts of interest

Where we stand

GenGrades is built by a team that also makes AI tools, among them LayerGrab, MagicRemover and the apimodels model gateway, which we also use to run many of the tests. We disclose that relationship wherever one of those products appears, and they go through the same test sets, the same blind scoring and the same published outputs as everything else. If you think a grade is wrong, the prompts and outputs are public: show us where.

Questions about the method

How do you score a model?

Every dimension has a fixed test set and a rubric with anchors on a 0 to 10 scale. The model runs the whole set, humans score the outputs against the anchors, blind to the model name where the output allows it, and the dimension score is the mean. A large language model may pre-screen outputs, but it never sets a final score.

Why does a tool show two sets of scores?

Because a tool is a model plus a product. Model capability scores describe the output; product experience scores describe using it, from pricing transparency to export. A tool page shows both so you can tell a great model behind a poor product from the reverse.

How is the total calculated?

The total is the weighted mean of the dimension scores, using the weights of the category the entry is graded in. Weights are published on every category page, and changing a weight recomputes every total in that category.

What happens when a model releases a new version?

We re-run the test set on the new version and publish a new score with its own date and version label. The old score stays on the record, so you can see whether a model is improving. Each re-run has a cost cap so that no single model can distort our budget.

Do you publish user ratings or aggregate scores?

No. GenGrades does not publish an aggregate rating until there are real user votes behind it, and it never shows a number that was not measured. The only ratings on the site are ours, labelled as reviews by GenGrades.

Can I reproduce your results?

Yes, that is the point. The prompts and input files for every test set are published, as are the raw outputs, watermarked. Run the same prompts on the same version and you should see the same kind of results; if you do not, tell us and we will re-test.