AI model and developer tooling updates #123
Today's Letter
NVIDIA, Qwen3.8-Flash-Next on GB300 NVL72

- NVIDIA published deployment results for Alibaba’s Qwen3.8-Flash-Next on the GB300 NVL72 platform on August 26, 2026
- The multimodal MoE model has 125B parameters, activates 6B per token, and supports a native 262,144-token context window extendable to 1M with YaRN
- Its hybrid architecture combines Gated DeltaNet in three of every four layers with Qwen Sparse Attention for full-context retrieval
- Alibaba benchmarks report up to 7.6x faster prefill and 4.9x faster decoding than full attention at 1M-token context
- NVIDIA reports throughput above 16K tokens per second per GPU and above 200 tokens per second per user on GB300 NVL72
- Developers can fine-tune checkpoints with NeMo AutoModel and reinforcement learning recipes from NeMo RL
- Open-source deployment recipes are available for SGLang, vLLM, and TokenSpeed, with weights distributed through Hugging Face and ModelScope
Source: developer.nvidia.com
GitHub Copilot in Visual Studio, August update

- GitHub added organization-level custom agents for Visual Studio 2026, available across repositories in GitHub organizations
- The agent picker shows each agent's description and organization source
- Supported models now provide Low, Medium, and High thinking-effort controls to balance reasoning depth and token usage
- Users can pin models, collapse less-used models, and review capabilities, context windows, and cost information
- The Git agent can review uncommitted changes and commits, with inline findings and navigable results in Git Changes
- Reviews support repositories hosted on GitHub and Azure DevOps
- Usage details and clearer limit notifications are available from the Copilot context window
- The update is available across Copilot Free, Student, Pro, Pro+, Max, Business, and Enterprise plans
Source: github.blog
Jocoletter curates AI, software, and product trends for developers and builders.
#GitHub #NVIDIA