Artificial Analysis has launched Optima, a platform that allows users to create custom AI benchmarks tailored to specific use cases. The platform enables users to test models against their own data or described scenarios, evaluating performance on quality, cost, and speed. According to Artificial Analysis, this addresses the limited usefulness of general-purpose benchmarks for real-world applications. Optima is now available for use, offering a solution to the shortcomings of existing AI benchmarking systems.
Users can build benchmarks using their own data, workflows, or descriptions of a use case, and then run them across leading current models. The platform supports multiple data sources, including existing evaluation datasets from Hugging Face, AI agent traces from platforms like Arize, Braintrust, or Langfuse, and skills that gather information from coding environments. For users without data, Optima generates suggested test inputs, evaluation criteria, and example tasks, which can be refined through feedback before running the benchmark. Two scoring approaches are available: rubric-based evaluation against objective criteria, or a pairwise comparison method used in benchmarks like GDPval-AA and AA-Briefcase.
Optima also tracks cost per task and time per task as standalone metrics, enabling users to assess whether performance gains justify higher costs or longer processing times. Early testers built benchmarks for finance and accounting agents, finding models that could reduce costs by up to 10 times without major quality loss. The platform charges only the actual token costs of the models used, with no markup, and offers pricing for rubric-based and pairwise evaluations. Billing is based on actual usage and evaluation costs incurred. According to Artificial Analysis, Optima addresses the problem of general benchmarks failing to capture a specific use case.
Source: thedecoder