24/06/2026
Choosing between AI performance and strict data compliance? That’s all water under the bridge now. 😉 Exoscale Dedicated Inference is officially out of preview and live for production-grade workloads. 🚀
You can now deploy any Hugging Face model as a production-ready, OpenAI-compatible API endpoint, backed by the security of a fully sovereign European infrastructure AND powered by dedicated NVIDIA GPUs
What this means for your production environment:
𝗔𝗯𝘀𝗼𝗹𝘂𝘁𝗲 𝗜𝘀𝗼𝗹𝗮𝘁𝗶𝗼𝗻:
Dedicated instances mean zero resource sharing. Your proprietary data and prompts never leave your environment.
𝗙𝗿𝗶𝗰𝘁𝗶𝗼𝗻𝗹𝗲𝘀𝘀 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻:
Swap in your preferred models and start querying immediately through standard API frameworks.
𝗘𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗦𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆:
Production workloads are now fully covered by a 99.95% SLA, with operational and support processes fully integrated into the standard Exoscale service lifecycle.
𝗩𝗲𝗿𝘀𝗮𝘁𝗶𝗹𝗲 𝗔𝗜 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲:
Fully optimized not just for LLMs, but also for complex embeddings, RAG applications, AI agents and custom inference APIs.
Our documentation has been completely expanded with deployment scaling guidance, updated CLI references service boundaries and model compatibility requirements.
Learn more: https://changelog.exoscale.com/en/ai-dedicated-inference-is-now-generally-available