- Image generation
- Prompt adherence (object counts, relations, attributes), text rendering, aesthetics and composition, realism and range of styles, anatomy and hands, subject consistency across images, resolution and aspect ratios, reference image and style control.
- Image editing
- Instruction fidelity (change only what was asked), preservation of untouched regions (identity, background), edge and mask precision, text editing, multi-turn editing, resolution preservation, inpainting and outpainting, layer separation.
- Video generation
- Motion quality and physical plausibility, temporal consistency (flicker, morphing), prompt adherence, camera control, duration, resolution and frame rate, fidelity to the source image in image-to-video, character consistency across shots, native audio, success rate.
- Video editing tools
- Feature coverage (cut, captions, background removal, upscaling), quality of the core edits, export formats and limits, processing speed, collaboration.
- Speech and text-to-speech
- Naturalness, pronunciation of names, numbers and abbreviations, emotion and style control, cloning quality and its consent policy, languages and accents, latency for real-time use, consistency over long texts, control granularity such as SSML.
- Audio processing and transcription
- Word error rate under accents and noise, speaker diarisation, timestamps, number of languages, speed.
- Music
- Musicality and structure, prompt adherence (genre, mood, tempo), vocal and lyric accuracy, duration and stem export, originality and copyright safety, commercial licence.
- Language models and writing
- Instruction following, factuality and hallucination rate, reasoning, long context, writing quality (style control, naturalness, how much it reads like a machine), structured output and tool calling, multilingual ability, context length, throughput and latency.