Start calling this model with endpoint-specific examples.
Headline benchmark standings and comparison context.
Key dates, capabilities, and model metadata.
Start calling this model with endpoint-specific examples.
Headline benchmark standings and comparison context.
Key dates, capabilities, and model metadata.
Start calling this model with endpoint-specific examples.
Headline benchmark standings and comparison context.
Key dates, capabilities, and model metadata.
11 Mar 2026
11 Mar 2026
License
Nvidia Open Model License
Input
Output
Core latency and throughput trends from recent traffic.
Headline benchmark standings and comparison context.
Top benchmark results for nvidia/nemotron-3-super-120b-a12b.
Detailed benchmark comparisons now live in the Compare tool.
Start calling this model with endpoint-specific examples.
Choose a supported endpoint, pick a main language, then select the example style you want to copy.
import AIStats from '@ai-stats/sdk';
const client = new AIStats({
apiKey: process.env.AI_STATS_API_KEY,
});
const response = await client.generateResponse({
"model": "nvidia/nemotron-3-super-120b-a12b",
"input": "Give me one fun fact about cURL.",
"service_tier": "standard"
});
const outputText = response.output
?.flatMap((item) => item.content ?? [])
.find((item) => item.type === "output_text")
?.text;
console.log(outputText ?? response);Parameters
Aggregated across active providers for the responses route.
Routing will select a compatible provider when a parameter narrows availability, so this list stays model-facing instead of provider-facing.
| Parameter | Description |
|---|---|
temperature | Controls how random token selection can be. |
top_p | Applies nucleus sampling by limiting candidates to a probability mass threshold. |
top_k | Restricts sampling to the top-k candidate tokens on providers that expose it. |
stop | Defines one or more sequences that terminate generation early. |
tool_choice | Controls which tool, if any, the model should call. |
tools | Defines callable tools or functions the model can invoke. |
response_format | Requests plain text, JSON, or schema-constrained output formats. |
structured_outputs | Capability signal for reliable schema-constrained output workflows. |
Select a route to update the request snippet and compatibility details.
/v1/responses/v1/chat/completions/v1/messagesLatency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime
Latency
Throughput
Uptime