Browsing Tag

AI Inference

5 posts

AI inference infrastructure, model serving, token generation, inference clouds, latency, throughput, and production AI workloads.

OpenAI CEO Sam Altman and Broadcom CEO Hock Tan holding a display with the Jalapeño inference chip wafer

OpenAI’s Jalapeño Benchmarks Turn Inference Power Into the AI Chip Fight

OpenAI published first benchmark results for Jalapeño, its Broadcom-built inference chip, claiming higher performance per watt and lower latency than leading commercial AI systems. The useful question is not whether it replaces Nvidia immediately, but whether custom inference silicon can turn power limits into a product advantage for ChatGPT, Codex, and agentic AI workloads.
Read More
Close-up of a computer chip on a circuit board

Qualcomm’s Modular Deal Is a $3.9 Billion Bet on AI Software Portability

Qualcomm agreed to acquire Modular in a nearly $4 billion stock deal, giving its AI data center push a software layer built around portable model deployment. The move is aimed at a practical bottleneck in AI infrastructure: making models run efficiently across CPUs, GPUs, NPUs, and custom accelerators without locking developers into one hardware stack.
Read More