AI Infrastructure and Developer Tooling Updates #119

Today's Letter

  1. NVIDIA, Qwen3.8-2.4T-A95B Serving on GB300
  2. Vercel Chat SDK, Slack Enterprise Grid support

NVIDIA, Qwen3.8-2.4T-A95B Serving on GB300

NVIDIA, Qwen3.8-2.4T-A95B Serving on GB300
  • NVIDIA reports serving Alibaba’s open-weight Qwen3.8-2.4T-A95B model on the GB300 NVL72 platform
  • The model has 2.4T total parameters, 95B activated per token, and a context window of up to one million tokens
  • Its hybrid full-attention and linear-attention architecture bounds memory growth for long-context workloads
  • Configurable reasoning levels—low, high, and xhigh—allow per-request tradeoffs between latency and reasoning depth
  • NVIDIA reports over 4K tokens per second per GPU and over 350 tokens per second per user in FP8 precision
  • The GB300 NVL72 combines 72 Blackwell Ultra GPUs with a 130 TB/s NVLink communication domain
  • Deployment options include SGLang, vLLM, NVIDIA Dynamo, NVIDIA NIM, and NeMo AutoModel

Source: developer.nvidia.com


Vercel Chat SDK, Slack Enterprise Grid support

  • Chat SDK's Slack adapter now supports Slack Enterprise Grid installations
  • Organization-wide bots work across every workspace with enterprise-scoped token resolution
  • Installation records include enterprise ID, installation type, and workspace identity fields
  • Event routing uses authorization data from incoming envelopes, including Slack Connect channels
  • Caches are scoped per installation to prevent cross-tenant profile and mention resolution
  • Retried event deliveries are deduplicated for 24 hours, including Slack delayed-event redeliveries
  • W-prefixed Grid user IDs are supported for outgoing mentions; single-workspace installations remain unchanged

Source: vercel.com


Jocoletter curates AI, software, and product trends for developers and builders.

#NVIDIA #Vercel

Subscribe to Jocoletter

Read more