AI inference is expensive - but it doesn’t have to be.

In this #InfoQ talk, Meryem Arik breaks down how to systematically reduce cost per token across AI workloads, with real-world examples from data transformation, offline agents, and aggregated insights.

🔹 Measure and optimize inference costs across NVIDIA and AMD GPUs 🔹 Understand key tradeoffs in vLLM, SGLang, and Dynamo

🔗 Watch now: https://bit.ly/4xyPRNN

#AI #LLM #CostOptimization