NVIDIA runs a hosted inference API at build.nvidia.com under the NIM (NVIDIA Inference Microservices) brand. You can call it from any OpenAI-compatible client, and the free tier gives you enough credits to prototype without setting up a GPU. Here is what is actually available, what the limits are, and when the hosted NIM endpoint makes more sense than going through an aggregator.
For other free LLM API tiers side by side, see free LLM API credits. For a free-first comparison with an aggregator, see OpenRouter free models and rate limits.
What NIM is
NIM is NVIDIA's packaging format for inference: each model is bundled with a CUDA-optimized runtime and exposed as a containerized microservice. At build.nvidia.com, NVIDIA runs a hosted version of these microservices. You get API access without managing any GPU infrastructure yourself.
The catalog is not just Nemotron. As of October 2026 it includes DeepSeek-V4-Pro and DeepSeek-V4-Flash (both with 1-million-token context windows), Nemotron 3 (a hybrid Mamba-Transformer mixture-of-experts model, also 1M context), GLM-5.1, Llama 4 variants, Qwen 3.5, Kimi K2.5, Mistral, and Microsoft Phi. NVIDIA adds models frequently; the full current list is at build.nvidia.com.
Getting free access
Go to build.nvidia.com and create an NVIDIA Developer account. It is free and requires no credit card. Once you are in, click any model in the catalog and select "Get API Key." Your key is ready within a few seconds.
The free tier includes 1,000 credits per month. One credit corresponds roughly to one API call on a standard-size request, though NVIDIA does not publish a formal credit-to-token conversion. Use the usage panel in your account to track consumption.
The free tier is scoped to prototyping, development, and research. Production workloads require an NVIDIA AI Enterprise license.
The endpoint
The API base URL is:
https://integrate.api.nvidia.com/v1
It is fully OpenAI-compatible. Pass your key as a Bearer token and use the standard chat completions format:
from openai import OpenAI
client = OpenAI(
base_url="https://integrate.api.nvidia.com/v1",
api_key=os.environ["NVIDIA_API_KEY"]
)
response = client.chat.completions.create(
model="meta/llama-3.3-70b-instruct",
messages=[{"role": "user", "content": "Hello"}]
)
Model names follow the format owner/model-slug. Check build.nvidia.com for the current slug for each model; they change when NVIDIA updates versions.
Rate limits
NVIDIA does not publish universal rate limits for the hosted API. In practice, most accounts see around 40 requests per minute on the free tier. If you need more, NVIDIA accepts informal requests to raise the ceiling to 200 RPM via the developer support channel. There is no self-service toggle.
If you exceed your limit, you get HTTP 429. To check the specific ceiling for your account and a given model, sign into build.nvidia.com, open the model page, and read the rate-limit field in your account panel: it is account-specific and not shown in the general documentation.
When NIM is a better choice than OpenRouter
For low-volume and prototyping use, the build.nvidia.com hosted API is not cheaper than OpenRouter. OpenRouter has published per-token billing, predictable costs, and a wider choice of providers in one place.
NIM makes more sense in two specific cases.
First, if you need models that are not available anywhere else. Nemotron 3 with a 1-million-token context window and NVIDIA's multimodal builds are only accessible through NIM or self-hosting. OpenRouter does not carry them.
Second, if you are moving toward self-hosting at scale. NIM containers are downloadable and run on any NVIDIA GPU. At high, steady utilization on your own hardware, the economics shift: you are paying GPU time rather than per-token rates, and the per-million-token effective cost drops significantly compared to managed API providers. The build.nvidia.com hosted endpoint lets you prototype with the same API contract before committing to the infrastructure.
For anything else, prototyping with common models, low-volume scripts, or occasional batch work, OpenRouter or the Groq free tier will be simpler and will cost less.