GitHub Models is gone. On July 30, 2026, GitHub shut down the service completely. The playground, the model catalog, the inference API, and bring-your-own-key (BYOK) support all went offline that day. If you had workflows or scripts calling https://models.github.ai/inference with a GitHub token, they stopped working.
Here is what the service was, how the shutdown rolled out, and where developers are moving.
What GitHub Models was
GitHub launched GitHub Models in beta in August 2024. The premise was simple: use a GitHub personal access token (PAT) to call LLMs through an Azure AI Inference-compatible endpoint. No separate account, no billing setup. If you already had a GitHub token, you had model access.
The model catalog at the time of shutdown included GPT-4o, GPT-4o Mini, GPT-4.1, GPT-4.1 Mini, Phi-4, Phi-4 Mini, Llama-3.3-70B, Llama-4 Maverick 17B, DeepSeek-R1, DeepSeek-V3, Grok-3, and Grok-3 Mini.
Rate limits were tiered by model:
- Smaller models (GPT-4o Mini, Phi-4 Mini): 15 requests per minute, 150 requests per day
- Larger models (GPT-4o, Llama-3.3-70B): 10 requests per minute, 50 requests per day
- Reasoning models (DeepSeek-R1, Grok-3): 1 to 2 requests per minute, 8 to 15 requests per day
The service was most useful for prototyping and GitHub Actions workflows. Because GITHUB_TOKEN was already present in the Actions runner environment, scripts could call LLMs without any extra secrets. Simon Willison documented one concrete breakage after the shutdown: a GitHub Actions workflow that used GITHUB_TOKEN to summarize repository folders automatically, which had worked for two years and went silent on July 30.
The shutdown timeline
- June 16, 2026: Closed to new customers. Organizations and enterprises with no prior GitHub Models usage lost access on both free and paid plans.
- July 1, 2026: GitHub announced the full retirement, with a July 30 hard cutoff for all customers including existing users with active usage.
- July 16, 2026: First scheduled brownout, an advance warning for teams still depending on the service.
- July 23, 2026: Second brownout.
- July 30, 2026: Complete shutdown. No grace period, no maintenance mode.
Why it closed
The post-mortems from GitHub and outside observers point to the same driver: agentic coding. When developers started running AI coding agents that loop through dozens or hundreds of model calls automatically, the cost of free subsidized inference scaled faster than GitHub could absorb. The playground model held for manual prototyping and small scripts, but not for agent workloads.
GitHub's stated path forward is GitHub Copilot for AI work inside GitHub, and Azure AI Foundry at ai.azure.com for teams that need a broader model catalog with billing control.
Where to migrate
Most GitHub Models code used the OpenAI Python or JavaScript client pointed at the GitHub Models base URL. Each option below is a drop-in swap: change the base URL, swap the authentication token, keep the rest of the code. Application logic stays the same.
Gemini API via Google AI Studio
Google AI Studio at aistudio.google.com issues API keys for Gemini models with no credit card required. As of October 2026, the free tier on Flash-class models allows 15 requests per minute and 1,500 requests per day. Gemini 2.0 Flash allows 1,000,000 tokens per minute on the free tier. The Gemini API exposes an OpenAI-compatible endpoint at https://generativelanguage.googleapis.com/v1beta/openai/. This is the closest free replacement if your daily request volume fits within the ceiling.
Groq free tier
Groq at console.groq.com offers free inference with an OpenAI-compatible endpoint at https://api.groq.com/openai/v1. As of October 2026, the free tier allows 30 requests per minute and 14,400 requests per day for smaller models, and 30 requests per minute with 1,000 requests per day for larger models. Adding a credit card to the account unlocks up to 10 times the free tier limits without requiring any minimum spend. The free model catalog changes; check the Groq console for the current list.
OpenRouter
OpenRouter at openrouter.ai routes requests to models from OpenAI, Anthropic, Google, Meta, DeepSeek, and others through a single OpenAI-compatible endpoint. It has a set of free models and pay-per-token access to the rest. Useful if you were using GitHub Models partly to avoid committing to a single provider. See OpenRouter free models and rate limits for the current free roster and daily availability data.
Azure AI Foundry
GitHub's official recommendation. Azure AI Foundry uses the same Azure AI Inference SDK that GitHub Models was built on, so migration is minimal for code that used the Azure SDK path. Requires Azure billing setup.
What to do right now
If your project called GitHub Models, the endpoint has been dead since July 30. Pick one of the alternatives above based on your volume and whether you need a no-card option. For most projects the swap takes under an hour: update the base URL, replace GITHUB_TOKEN with the new provider key, run one test call.
If you had GitHub Actions workflows that relied on GITHUB_TOKEN granting model access automatically, you need to add a repository secret with the new API key and update the workflow to pass it explicitly.