DeepSeek V4 Pro - Benchmarks, Pricing & API Access

Performance

Core latency and throughput trends from recent traffic.

No gateway telemetry yet

This model hasn't processed any gateway traffic in the selected window. Live charts will appear as soon as requests arrive.

Quickstart

Start calling this model with endpoint-specific examples.

Step 1

Get an API key

Create an API key inSettingsKeysand store it asAI_STATS_API_KEY

Keep it server-side, never commit it, and rotate it immediately if exposed.

Step 2

Send the request

Choose a supported endpoint, pick a main language, then select the example style you want to copy.

Streaming

import AIStats from '@ai-stats/sdk';

const client = new AIStats({
  apiKey: process.env.AI_STATS_API_KEY,
});

const response = await client.generateResponse({
    "model": "deepseek/deepseek-v4-pro",
    "input": "Give me one fun fact about cURL.",
    "service_tier": "standard"
});

const outputText = response.output
  ?.flatMap((item) => item.content ?? [])
  .find((item) => item.type === "output_text")
  ?.text;

console.log(outputText ?? response);

Accepted IDsClick to use and copy

Parameters

Aggregated across active providers for the responses route.

Routing will select a compatible provider when a parameter narrows availability, so this list stays model-facing instead of provider-facing.

View all parameters

Parameter	Description
`temperature`	Controls how random token selection can be.
`top_p`	Applies nucleus sampling by limiting candidates to a probability mass threshold.
`max_tokens`	Caps output length on endpoints and providers that use the max_tokens field name.
`frequency_penalty`	Discourages repeated tokens in proportion to how often they already appeared.
`presence_penalty`	Encourages the model to explore new wording or topics after they first appear.
`stop`	Defines one or more sequences that terminate generation early.
`logprobs`	Requests token-level probability data in the response.
`tool_choice`	Controls which tool, if any, the model should call.
`tools`	Defines callable tools or functions the model can invoke.
`response_format`	Requests plain text, JSON, or schema-constrained output formats.
`reasoning`	Provider-specific reasoning configuration for reasoning-capable APIs.
`include_reasoning`	Requests reasoning content or reasoning summaries in responses where supported.
`top_logprobs`	Limits how many alternative token probabilities are returned per position.