Signal Comparative assessment of publicly benchmarked Indian foundation models
Summary
The paper presents a benchmark-based comparison of Indian foundation models against global frontier and comparable-scale models. The comparison spans eight capability domains, including general-purpose reasoning, coding, agentic AI, cybersecurity and Indic-language ability. Using only publicly reported benchmark results, the authors find Indian models score well on established but now-saturated benchmarks such as MMLU and MATH-500, while participating far less in newer agentic and domain-specialized evaluations. Participation varies widely across organizations, with Sarvam AI showing the broadest benchmark coverage. The authors propose an exploratory Benchmark Maturity Index built on four dimensions, namely standardization, participation, independent verification and national coverage. They argue that many apparent capability gaps in the public record may reflect gaps in the evaluation ecosystem rather than true capability gaps, with direct implications for how national AI programs design monitoring and funding criteria.
Classification
Evidence 1
- Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework arXiv (cs.CY, cs.AI, cs.HC) 2026-08-12 accessed 2026-08-13T13:49:29+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-849a0aa0ce57
