01

Freeze the fixture

Store the exact prompt, negative prompt if supported, model identifier, size or aspect ratio, seed behavior, safety settings, and number of requested outputs. Keep the fixture separate from conclusions so it can be rerun when a model changes.

02

Preserve the original result

Download expiring assets, keep the original format, and link each asset to a machine-readable run record. Record submission time, completion time, attempts, task ID when safe to publish, and the provider-reported cost. Never include credentials or signed URLs.

03

Review with declared criteria

  1. Instruction following and composition.
  2. Preservation of required product or character attributes.
  3. Text, anatomy, marks, and unwanted-object defects.
  4. Latency, failure, retry, and effective cost evidence.
04

Keep the conclusion narrow

One prompt and one accepted output demonstrate a workflow, not a universal winner. Publish the test date and limitations beside the finding. Explore the open benchmark hub for fifty fixtures, thirty raw outputs, and dated comparison pages.

Run your own test

Use supported models through one APIMART account.

Confirm current model availability, pricing, and limits before routing production traffic.

Create an APIMART account