Signal JEV versus LLMs on seven political science replications
Summary
This paper tests whether JEV, a commercial model that returns decisions and probability distributions over a fixed answer set, matches LLMs on social science text tasks. The authors compare JEV with LLMs and human coders from published research across seven political science replications. Comparators include a mid-tier commercial LLM (GPT-6 Luna) and an open-weight model (Qwen3.8-27B). JEV matches or comes close to both LLMs across a range of tasks. At OpenAI's batch prices, however, it shows no cost advantage over GPT-6 Luna. Its probabilities are better calibrated than GPT-6 Luna's token probabilities when each question is asked once, but not consistently better than Qwen3.8-27B's.
Classification
Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-10-07)
Last updated2026-10-07 08:51 KST
Evidence 1
- JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications arXiv (cs.CY) 2026-10-05 accessed 2026-10-06T23:00:14+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-35b608d40154
