Presets: a toolkit for agent-based inference optimization
Optimizing model inference is agent work now. Every inference provider does it inside its own process, on its own serving stack, with its own harness around the optimization agent. Despite the progress in open-source serving frameworks, what gets published is a benchmark, often without the workload, the concurrency, and the hardware behind it. The optimized deployment itself stays tied to the stack that produced it.
Today we're introducing a preview of presets: an open-source toolkit that streamlines inference optimization with agents, and a portable preset format.









