InferX Beta Serverless GPU Inference Platform, Built for Agent-Native Workloads

Endpoint Ornith-1.0-35B-FP8

Ornith-1.0-35B-FP8

Metadata

Name
Ornith-1.0-35B-FP8
Provider
protoLabsAI
Parameter Size
GPU Count
1
Context Length
262000
Concurrency
4.56x
Cold Start TTFT
Recommended Use Cases
Detailed Intro
The benchmark test result is much better than Qwen3.6-35B. Some score is even near Qwen3.5-397B. https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B

Log In To Use This Endpoint

This public page shows the published endpoint metadata and integration shape. Log in to get a tenant-scoped endpoint URL, inference API key, and the interactive playground. Log in

Integration

Start with the shared setup below, then choose the endpoint mode that fits your workload: token-based shared serving or a tenant-scoped dedicated GPU endpoint.

  1. Choose either the token-based shared endpoint or the dedicated GPU endpoint below.
  2. Copy that section's base URL into your client endpoint field.
  3. Copy the shared inference API key once and reuse it for either mode.
  4. Use the model name shown in the selected section.
<INFERENCE_API_KEY>

An inference API key is required for both endpoint modes. Until one is available, the fields below keep the correct request shape and use a placeholder token.

Token-Based

Use the shared OpenAI-compatible endpoint. The base URL is shared across published endpoints, and usage is billed per token.

https://model.inferx.net/endpoints/v1
Ornith-1.0-35B-FP8
Sample REST Call
curl https://model.inferx.net/endpoints/v1/chat/completions \
  -H "Authorization: Bearer $INFERX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "Ornith-1.0-35B-FP8", "messages": [{"role": "user", "content": "Hello"}]}'

Dedicated GPU

Model Spec