Reference-agent full sweep across commercial and open LLMs on all 100 tasks (single run), sorted by accuracy β a bilingual, multi-modal benchmark over a real telecom operator's data, scored by exact string matching. Duration and turns are per task. β run under rate limits.