NVIDIA AVO and ASR Benchmark Results #115

Today's Letter

  1. NVIDIA AVO, ARC-AGI-3 100% score achieved
  2. HumeAI, ASR benchmark optimization measured

NVIDIA AVO, ARC-AGI-3 100% score achieved

NVIDIA AVO, ARC-AGI-3 100% score achieved
  • NVIDIA’s Agentic Variation Operators (AVO) system achieved a 100.00 RHAE score on the ARC-AGI-3 public set
  • AVO completed all 183 levels across 25 environments
  • The architecture combines persistent memory, supervision, tool use, iterative planning, execution, and evaluation
  • In GPU-kernel optimization, AVO ran autonomously for seven days, explored over 500 directions, and committed 40 kernel versions
  • The resulting attention kernels were up to 3.5% faster than cuDNN and 10.5% faster than FlashAttention-4 on DGX B200
  • The results position agent performance as a system-level property rather than a model-capability metric alone

Source: developer.nvidia.com


HumeAI, ASR benchmark optimization measured

HumeAI, ASR benchmark optimization measured
  • HumeAI introduced three tests to measure benchmark optimization in speech recognition
  • The study evaluated 11 widely used open-source ASR models
  • Several models reproduced VoxPopuli and LibriSpeech reference transcripts even when audio contradicted them
  • Six of 11 models omitted an audible “Thank you” to match an erroneous VoxPopuli transcript
  • The method flagged potential reference errors in 40% of analyzed VoxPopuli clips, covering roughly 3% of reference words
  • Benchmark-optimized models reproduced erroneous references 18–30% of the time
  • Held-out sets in Real World VoiceEQ and the Open-ASR and Far-field ASR Leaderboards are intended to measure real-world performance

Source: huggingface.co
More: unite.ai · gigazine.net


Jocoletter curates AI, software, and product trends for developers and builders.

#HumeAI #NVIDIA

Subscribe to Jocoletter

Read more