InferQoS documentation

InferQoS documentation

InferQoS makes finite AI inference capacity a shared, schedulable resource. It performs admission, fair scheduling, and capacity accounting; it is deliberately not a general AI gateway.

Start here

Operate and extend

The default configuration records no prompts or completions and sends no project analytics.