Latest AI and tech news
Voice agents have a latency problem that shows up as soon as they have to do real work. Within five days, Google and OpenAI shipped two very different fixes....
Fiona McDonnell · 15 Sep 2026
Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize output within the factory’s limited power budget. This makes performance per watt—rather than...
Since about 2020, AI has largely focused on training bigger and better models. Large language models (LLMs) ballooned from millions of parameters to trillions. This proved effective: The largest version of OpenAI’s GPT-3, released in 2020, correctly ...
TL:DR – AxeleraScript allows customers to compile their own models, not just supported models with their own weights, but their own models, on Axelera hardware!...
The model is only one part of the inference pipeline. At scale, the work happening around it, from routing requests to coordinating…...
Picking an inference vendor used to be a short conversation. You wanted Llama behind an HTTP endpoint, three companies served it, and their prices sat close enough that the decision came down to whoever had capacity. The market has since split into a dozen serious operators running truly different b
Selecting the right serving engine for your embedding model can dramatically outperform hardware upgrades, yielding up to an 11x throughput increase on the same GPU.
By Chander Chadha, Director of Product Marketing, Storage Products, Marvell
Getting a model serving on day one is the easy part. The expensive half of running a model catalog is maintaining every model, precision, GPU, and inference engine combination as the stack underneath keeps moving.
Novita AI has open-sourced Chord, a high-performance W4A16 MoE CUDA kernel for Kimi K2.x serving shapes, with a Humming-compatible indexed path and grouped SM90 operators.
Rubin is the first platform co-designed across six products for the agentic era: Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, and Spectrum-6. Today we are publishing the first verified agentic inference results for Rubin, measured o...
So far, AI has mostly lived behind a screen. Chatbots answered questions. Then agents started driving software and finishing multi-step tasks on their own. The next step is AI that acts in the physical world, and the biggest piece of that is robots. ...
When we introduced co-operative time-slicing in llm-d, we made a claim: if RL phases become schedulable units, independent jobs can share accelerators with near-zero waste. Today we're backing that claim with a measured, end-to-end proof. For researc...
DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model from deepseek-ai that accepts text and images and generates text. Its defining feature is memory and inference efficiency for long, input-heavy workloads: it has 552B backbone parameters bu...
Last October, the P99 conference — the online gathering for developers focused on high-performance, low-latency applications — featured a cracking keynote from Chip Huyen....
The best AI TTS tools for ops leaders ranked on latency, concurrency, and compliance so your 2026 campaigns never fail where demos can't show you.
Rank the best TTS for AI voice agents before latency kills your live calls. Tested for enterprise pipelines so production never fails.
Kimi K3 serving optimizations across scheduling, KDA prefix caching, ReplaySSM state recovery, PD disaggregation and state offload, parallelism, MoE, and GPU kernels.
A simple guide to the real GPU metrics and monitoring tools you need to run AI inference without wasting money or losing performance....
How one inherited configuration flag shaped GLM-4.5-Air’s performance....
Hey yall, welcome back to ThursdAI, this is Alex, let me catch you up!...
Jobs reach their wall-time limits and lose hours or days of in-memory progress. However, the limit itself is not the problem.
Lei Pan | Senior Software Engineer; Salina Wu | Senior Software Engineer; Cristian Lopez | Software Engineer I; Guangtong Bai | Staff Software Engineer; Soam Acharya | Principal Engineer; Saurabh Vishwas Joshi | Principal Engineer; Chia-Wei Chen | St...
Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA BioNeMo Inference Runtime (BioIR) helps accelerate supported biomolecular structure-prediction...