The LLM Tax, Revisited: Cost, Latency, and Where to Run the Model
Quality, cost, and latency. You can usually have two. Here is the framework I use to pick which two on purpose, with a voice assistant, an RSS scorer, and a Frigate box as the worked examples.









