InferQoS documentation
InferQoS documentation
InferQoS makes finite AI inference capacity a shared, schedulable resource. It performs admission, fair scheduling, and capacity accounting; it is deliberately not a general AI gateway.
Start here
- Five-minute local demo: no cloud account or paid API key required
- Architecture and trust boundaries
- Scheduling, deadlines, and fairness
- Deployment guide
- PlugLayer deployment boundary
- Provider configuration
Operate and extend
- Identity: OIDC, mTLS, API keys, and trusted proxies
- Privacy and telemetry policy
- OTLP and Prometheus operations
- Embedded operations dashboard
- Secure request spooling and hot reload
- Durable queue adapters
- External provider gRPC protocol
- Provider SDK
- QoS headers
The default configuration records no prompts or completions and sends no project analytics.