AI Infrastructure and Developer Tooling Updates #119
Today's Letter
NVIDIA, Qwen3.8-2.4T-A95B Serving on GB300

- NVIDIA reports serving Alibaba’s open-weight Qwen3.8-2.4T-A95B model on the GB300 NVL72 platform
- The model has 2.4T total parameters, 95B activated per token, and a context window of up to one million tokens
- Its hybrid full-attention and linear-attention architecture bounds memory growth for long-context workloads
- Configurable reasoning levels—low, high, and xhigh—allow per-request tradeoffs between latency and reasoning depth
- NVIDIA reports over 4K tokens per second per GPU and over 350 tokens per second per user in FP8 precision
- The GB300 NVL72 combines 72 Blackwell Ultra GPUs with a 130 TB/s NVLink communication domain
- Deployment options include SGLang, vLLM, NVIDIA Dynamo, NVIDIA NIM, and NeMo AutoModel
Source: developer.nvidia.com
Vercel Chat SDK, Slack Enterprise Grid support
- Chat SDK's Slack adapter now supports Slack Enterprise Grid installations
- Organization-wide bots work across every workspace with enterprise-scoped token resolution
- Installation records include enterprise ID, installation type, and workspace identity fields
- Event routing uses authorization data from incoming envelopes, including Slack Connect channels
- Caches are scoped per installation to prevent cross-tenant profile and mention resolution
- Retried event deliveries are deduplicated for 24 hours, including Slack delayed-event redeliveries
- W-prefixed Grid user IDs are supported for outgoing mentions; single-workspace installations remain unchanged
Source: vercel.com
Jocoletter curates AI, software, and product trends for developers and builders.
#NVIDIA #Vercel