AI
Môtô famintinana C mahomby ho an'ny modely LLaMA miaraka amin'ny mpizara HTTP
C++ implementation with SIMD acceleration (AVX2, AVX512, NEON) for exceptional CPU performance.
2-bit to 8-bit quantized models (GGUF) reducing memory footprint while maintaining quality.
HTTP server with /v1/chat/completions, /v1/completions, /v1/embeddings endpoints.
Compatible with LLaMA, Mistral, Mixtral, Yi, Phi, Falcon, StarCoder, and more.
Support for 4K to 32K+ tokens with efficient KV cache management.
Request queuing, concurrent inference, streaming, Prometheus metrics, health checks.
Manomboka ao anatin'ny minitra vitsy amin'ny alalan'ny fizotry ny fametrahana tsotra VPS
Mandefa amin'ny manodidina ny 2 minitra • Full root access • Fanohanana amin'ny alalan'ny tapakila sy mailaka