AI inference is expensive - but it doesn’t have to be.
In this #InfoQ talk, Meryem Arik breaks down how to systematically reduce cost per token across AI workloads, with real-world examples from data transformation, offline agents, and aggregated insights.
🔹 Measure and optimize inference costs across NVIDIA and AMD GPUs 🔹 Understand key tradeoffs in vLLM, SGLang, and Dynamo
🔗 Watch now: https://bit.ly/4xyPRNN
Comments (0)