NVIDIA AVO and ASR Benchmark Results #115
Today's Letter
NVIDIA AVO, ARC-AGI-3 100% score achieved

- NVIDIA’s Agentic Variation Operators (AVO) system achieved a 100.00 RHAE score on the ARC-AGI-3 public set
- AVO completed all 183 levels across 25 environments
- The architecture combines persistent memory, supervision, tool use, iterative planning, execution, and evaluation
- In GPU-kernel optimization, AVO ran autonomously for seven days, explored over 500 directions, and committed 40 kernel versions
- The resulting attention kernels were up to 3.5% faster than cuDNN and 10.5% faster than FlashAttention-4 on DGX B200
- The results position agent performance as a system-level property rather than a model-capability metric alone
Source: developer.nvidia.com
HumeAI, ASR benchmark optimization measured

- HumeAI introduced three tests to measure benchmark optimization in speech recognition
- The study evaluated 11 widely used open-source ASR models
- Several models reproduced VoxPopuli and LibriSpeech reference transcripts even when audio contradicted them
- Six of 11 models omitted an audible “Thank you” to match an erroneous VoxPopuli transcript
- The method flagged potential reference errors in 40% of analyzed VoxPopuli clips, covering roughly 3% of reference words
- Benchmark-optimized models reproduced erroneous references 18–30% of the time
- Held-out sets in Real World VoiceEQ and the Open-ASR and Far-field ASR Leaderboards are intended to measure real-world performance
Source: huggingface.co
More: unite.ai · gigazine.net
Jocoletter curates AI, software, and product trends for developers and builders.
#HumeAI #NVIDIA