Endpoint Devstral-2-123B-Instruct-2512-int4-AutoRound
Devstral-2-123B-Instruct-2512-int4-AutoRound
Metadata
Log In To Use This Endpoint
This public page shows the published endpoint metadata and integration shape. Log in to get a tenant-scoped endpoint URL, inference API key, and the interactive playground. Log in
Integration
Start with the shared setup below, then choose the endpoint mode that fits your workload: token-based shared serving or a tenant-scoped dedicated GPU endpoint.
An inference API key is required for both endpoint modes. Until one is available, the fields below keep the correct request shape and use a placeholder token.
Token-Based
Use the shared OpenAI-compatible endpoint. The base URL is shared across published endpoints, and usage is billed per token.
curl https://model.inferx.net/endpoints/v1/chat/completions \
-H "Authorization: Bearer $INFERX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "Devstral-2-123B-Instruct-2512-int4-AutoRound", "messages": [{"role": "user", "content": "Hello"}]}'
Dedicated GPU
Use the tenant-scoped endpoint URL for this published endpoint when you want dedicated GPU serving instead of shared token-based routing.
curl -X POST https://model.inferx.net/funccall/<tenant>/endpoints/Devstral-2-123B-Instruct-2512-int4-AutoRound/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <INFERENCE_API_KEY>' \
-d '{"max_tokens": "1000", "model": "Intel/Devstral-2-123B-Instruct-2512-int4-AutoRound", "stream": "true", "temperature": "0", "messages": [{"role": "user", "content": "write a quick sort algorithm."}]}'