NewsNTech
In AI inference, the economics turn on how much useful output a chip delivers per unit of energy and cost.
OpenAI's Jalapeño chip has beaten Nvidia's Blackwell systems on key inference-efficiency tests, applying direct pressure to Nvidia's margins at the layer of the stack where that competition now concentrates.
The result arrives as major technology companies broadly accelerate their development of custom AI silicon. The inference constraint Training a large AI model is a bounded, periodic event.
Serving a model's output against live traffic runs continuously, and its cost accumulates in direct proportion to how efficiently a chip converts energy and compute into tokens delivered.
Keep reading