Are Open Source Models Teaching to the Test?

OS World Benchmark

In the last week, there has been a lot of talk about how open source models are nearly as good as the frontier models on benchmarks and are 1/10th of the cost. Our experience with our use cases have caused us to question the benchmarking results for the open source models.

The primary capability of the model we need is Computer Use (CU). The test for this use case is the OSworld benchmark shown here. As you can see, Claude Opus 4.6 came close to or started exceeding human performance (>72.9%) starting February 2026. GPT followed with similar performance starting with GPT 5.4 in Mar 2026. Kimi published that they are at 73.1% in April 2026.

The results for our use cases match the benchmark results published by Anthropic and OpenAI. Our results also match with Gemini’s poor performance. Our results on Kimi have not been consistent with their benchmark results.

 

Last fall, we started testing with OpenAI, Gemini and Anthropic models. Starting in Q1, the Anthropic models - both Opus and Sonnet - started performing really well on our internal benchmarks. In Q2, the latest GPT models from OpenAI also started performing well. Both the OpenAI and Anthropic models were consistent with the benchmarking data. On two different benchmarks, we were able to test Anthropic models against both Google Vertex and Amazon Bedrock. We were also able to test against GPT 5.5. The results for all these runs are captured below. No consistent pattern in terms of speed, except Opus seems to be faster than Sonnet on both GCP and AWS.

Based on these results and alignment with the OSworld benchmarks, we tried the Kimi 2.6 model. Unfortunately the agents did not complete the tasks, so we are not able to report performance results. We have reached out to the Kimi/Moonshot team to understand if they have any guidance for us.

For now, the results for our use cases match the benchmark results published by Anthropic, OpenAI and Gemini. Our results with Kimi have not been consistent with the benchmarks.

Amitabh Sinha

Co-Founder & CEO

LinkedIn

Previous
Previous

Accuracy is the Key to AI Agents

Next
Next

Anyone Can Rapidly Create Agents Using Screenshots