In section Releases

ITSMBench Exposes AI Limitations in Enterprise IT Service Management

Even the most advanced frontier AI models currently fail to complete half of the tasks required by enterprise IT service teams. A new open-source benchmark, ITSMBench, reveals that top-tier models from OpenAI, Anthropic, and xAI often struggle with the complex, multi-layered workflows inherent in corporate environments.

ITSMBench Exposes AI Limitations in Enterprise IT Service Management

Developed by Atomicwork and New Measure, the benchmark simulates real-world enterprise conditions using 42 mocked software systems, 1,800 database tables, and over 2,000 REST endpoints. It tests 89 specific service desk tasks ranging from identity management to complex infrastructure escalations. The results indicate that while models excel at specific sub-tasks, they frequently falter when navigating the security requirements and interconnected systems of a modern business.

Performance metrics highlight a fragmented landscape: Grok 4.5 demonstrates superior tool discovery at 83%, yet struggles to finalize resolutions. Conversely, Opus 5 exhibits the highest execution rate at 63.5% once tools are identified. The findings suggest that relying on a single "frontier" model is insufficient for enterprise needs. Instead, the developers advocate for an orchestrated approach, where different models are paired with specific agent frameworks to balance accuracy, cost, and speed. The full methodology and environment are now open-source, providing CIOs with a standardized way to evaluate AI reliability before deploying agents into high-stakes production workflows.

Share:on TelegramXFacebook

Subscribe to our newsletter

Once a week — the best stories from our editors, no ads or push notifications. Delivered Sunday morning.

Comments (0)

Leave a comment

No comments yet. Be the first!