Omneky has introduced TASTE BENCH, an evaluation suite designed to measure how effectively generative AI models turn advertising briefs and brand assets into usable creative. The system tests image and video models against advertising-specific quality, brand, copy, placement, and compliance requirements, adding a structured evaluation layer to AI-powered ad production.
Generative AI has made it easier for advertisers to produce images and video at scale, but producing an asset that can actually enter a campaign is a different problem. Creative teams still have to check typography, logos, product accuracy, language, claims, composition, and whether the final asset works as an advertisement.
Omneky’s TASTE BENCH is designed to quantify that gap.
The evaluation suite compares AI-generated advertising creative using identical briefs and brand assets. Its current image evaluation covers eight models, 59 ad briefs, four brands, and five languages: English, Japanese, Arabic, Hindi, and Spanish. The tests generated 469 ads and produced 1,371 valid blind judge reviews.
For the current image benchmark, GPT Image 2.5 Sunburst recorded a 54.2% ready-to-run rate, with 32 of 59 tested outputs meeting the benchmark’s criteria. Its average quality score was 7.38 out of 10. Across the eight models, ready-to-run rates ranged from 22% to 54.2%.
Those numbers need to be viewed within the scope of the test. These are first-attempt results with Omneky’s own review process disabled. The benchmark is also based on automated creative evaluation rather than campaign conversion, revenue, or return on ad spend. Omneky itself describes the results as creative judgments rather than performance outcomes.
The methodology is built around advertising-specific checks. Models receive the same prompts and assets, while a blind panel of AI judges from OpenAI, Anthropic, and Google evaluates outputs across eight quality dimensions, including brand fit, typography, and ad effectiveness. Eleven pass/fail checks examine issues such as exact copy, language accuracy, logo and product fidelity, safe-zone placement, fabricated claims, and visual artifacts.
That approach addresses a growing problem in AI advertising workflows: generation quality and advertising readiness are not synonymous.
An image can look visually polished while containing incorrect product details, malformed text, an inaccurate logo, or an unsupported claim. For advertisers, those failures can create another layer of manual review even when the underlying generation model produces creative quickly.
TASTE BENCH therefore positions evaluation as part of the advertising technology stack. Omneky says its AI Growth Agent can select among image and video models from multiple AI labs for individual advertising tasks. The benchmark provides a framework for making those model-selection decisions based on advertising-specific criteria rather than relying solely on general-purpose model capabilities.
The development also points toward a broader change in creative operations. As advertisers use multiple generative models simultaneously, agencies and marketing teams need ways to compare outputs on requirements specific to paid media.
That does not make TASTE BENCH an independent certification system. The methodology, judging framework, model selection, and benchmark design come from Omneky. Its results are most useful as a documented evaluation of how the tested models performed under this particular advertising workflow.
For AdTech, however, the larger development is significant: creative generation is increasingly becoming an automated media workflow, and evaluation is emerging as a necessary control layer between AI generation and campaign deployment.
Market Landscape
Generative AI is increasingly being incorporated into advertising workflows for image generation, video creation, copy development, personalization, and campaign iteration. The challenge is shifting from simply generating more creative to determining which outputs satisfy brand and advertising requirements.
TASTE BENCH reflects that transition by treating creative evaluation as a measurable component of ad production. Its checks around copy, logos, products, claims, language, and placement resemble the operational controls advertisers already apply during human creative review.
Strategic Outlook
As advertisers adopt multiple generative models, model selection may increasingly depend on the requirements of a particular campaign rather than a single model being used for every task. Evaluation frameworks can provide one mechanism for comparing models against those requirements.
The harder industry question will be whether automated creative scores correlate with actual campaign outcomes. Omneky’s current benchmark does not establish that relationship because it measures creative readiness and quality rather than conversions or ROAS.
Top Insights
- Omneky’s TASTE BENCH evaluates generative AI models against advertising-specific creative quality, brand, compliance, and ready-to-run requirements.
- The current benchmark covers eight image models, 59 briefs, four brands, five languages, 469 generated ads, and 1,371 blind reviews.
- GPT Image 2.5 Sunburst recorded a 54.2% ready-to-run rate in Omneky’s current 59-brief evaluation.
- Eleven hard checks test issues including copy accuracy, logos, products, safe zones, fabricated claims, and visual artifacts.
- The benchmark measures automated creative judgments rather than campaign conversions, revenue, or return on advertising spend.
Get in touch with our Adtech experts
