The fastest free start #
If you have a Google account, this is the shortest way to a working model:
- Open aistudio.google.com/apikey, sign in with your Google account and create a key.
- In Runesmith Studio, open Thinking power and click Add thinking power. Under With a key, choose Google Gemini and paste the key.
- Runesmith asks Google which models it serves today and fills in a recommended one. At the time of checking that is
gemini-3.8-flash, withgemini-3.6-flashas the second choice. Click Save and test: one tiny call confirms that the model answers.
Free tiers change without notice and always have limits. Read the provider’s own page before you rely on a number, and give each role a second model from a different provider so that a spent allowance does not stop the work (see the tips at the end).
At a glance #
| Where | Free? | Get a key | Names to start with |
|---|---|---|---|
| Google AI Studio (Gemini) | Yes: the current flash models are listed as free of charge | aistudio.google.com/apikey | gemini-3.8-flash, then gemini-3.6-flash |
| NVIDIA | Yes, for many models: a free endpoint with rate limits | build.nvidia.com/models | nvidia/nemotron-3-super-120b-a12b, moonshotai/kimi-k3 |
| Groq | Yes: a free plan with per-model limits | console.groq.com/keys | openai/gpt-oss-20b, openai/gpt-oss-120b |
| OpenRouter | Yes, for models whose id ends in :free |
openrouter.ai/keys | any model id ending in :free |
| Mistral | A free plan with monthly API credits | console.mistral.ai/api-keys | codestral-latest, mistral-small-latest |
| Your own computer | Yes: no key, nothing leaves the machine | Ollama, LM Studio or llama.cpp | qwen2.5-coder:7b with Ollama |
| A chat window | Whatever chat you already use | none | none |
Google AI Studio (Gemini) #
Checked 5 October 2026
- Free today: yes. Google’s pricing page (marked “Last updated 2026-10-01 UTC”) shows
gemini-3.8-flashas “Free of charge” in its Free Tier column, for input and for output. It shows the same forgemini-3.7-flash,gemini-3.6-flash,gemini-3.5-flash,gemini-3.5-flash-liteandgemini-3.1-flash-lite. (Checked 2026-10-05.) - Which name: use
gemini-3.8-flashfirst, thengemini-3.6-flash. Model names go stale every few months. If Google answers that a name is “no longer available” and names a replacement, Runesmith shows that replacement as a button. It never switches by itself. The pricing page still shows Gemini 2.5 Pro and 2.5 Flash as free, but Google’s models page says access to the 2.5 models is limited to users who have used them before and recommends 3.8 Flash and 3.5 Flash-Lite for new projects, so a new key may not get 2.5. (Checked 2026-10-05.) - Your prompts: the pricing page marks “Used to improve our products” as Yes for the free tier and No for the paid tier. Read Google’s terms, linked from that page, before you point Runesmith at private code. (Checked 2026-10-05.)
- Getting a key: aistudio.google.com/apikey. Google’s API-key page (marked “Last updated 2026-09-25 UTC”) sends you there and says a new user gets a default Google Cloud project and key automatically after accepting the Terms of Service. It does not say whether a card is asked for. The rate-limits page lists the Free tier’s qualification as “Active project or free trial” and says that moving to a paid tier first requires setting up billing. When this guide was first checked by hand, on 28 September, no card was asked for at the key step; no Google page states that, so treat it as an observation, not a promise. (Checked 2026-10-05.)
- Limits: Google publishes no fixed free-tier numbers. Its rate-limits page (marked “Last updated 2026-09-02 UTC”) says limits depend on several factors, such as your usage tier, and can be viewed in Google AI Studio. Your own per-model figures for requests per minute, input tokens per minute and requests per day are on your own account. The same page says limits apply per project, not per API key, that daily request quotas reset at midnight Pacific time, and that the specified limits are not guaranteed. (Checked 2026-10-05.)
- Sources: pricing, models, API key, rate limits, all on ai.google.dev, checked 2026-10-05.
NVIDIA (build.nvidia.com) #
Checked 5 October 2026
- Free today: yes, for many models. The models page at
build.nvidia.com/modelsshows a Free Endpoint tag onmoonshotai/kimi-k3. The pages of both names Runesmith suggests,moonshotai/kimi-k3andnvidia/nemotron-3-super-120b-a12b, list “Free Endpoint” as available under Model Availability and say “Start building with a free API endpoint.” NVIDIA’s NIM documentation FAQ says members of the NVIDIA Developer Program have free access to NIM API endpoints for prototyping, and that the Developer Program is free to join throughbuild.nvidia.com. (Checked 2026-10-05.) - Limits: the data served with a model page includes the rate-limit text “Up to 40 rpm” and “10,000 requests per day”, with a note that limits may vary by model and that traffic from other users may cause throttling. That text is part of the page’s content but was not displayed anywhere on the signed-out page, so treat it as indicative and look for your own limit once you have a key. (Checked 2026-10-05.)
- Credits and sign-up: none of the NVIDIA pages checked says whether free credits are given, how many, or whether a card is asked for. The API description served with the Nemotron page lists a 402 “Payment Required” response with the example text “You have reached your limit of credits”, so a credit limit of some kind may exist; the pages do not say how large it is. This guide does not use figures about NVIDIA credits from other sites. (Checked 2026-10-05.)
- Read the terms: the endpoints are offered as a trial. The trial terms text on NVIDIA’s model page says that your input and output are recorded to provide the trial and to improve NVIDIA’s products and services, and it asks you not to upload confidential information or personal data. Think about that before you point Runesmith at private code. (Checked 2026-10-05.)
- Getting a key: open a model’s page on
build.nvidia.comand use Generate API Key. The model page’s code example callshttps://integrate.api.nvidia.com/v1/chat/completions, so the OpenAI-compatible base address ishttps://integrate.api.nvidia.com/v1, and the model id is the slug in the page’s address, for examplemoonshotai/kimi-k3. (Checked 2026-10-05.) - Sources: models, kimi-k3, nemotron-3-super-120b-a12b, NIM FAQ, all checked 2026-10-05.
Groq #
Checked 5 October 2026
- Free today: yes. Groq’s rate-limits page has a Free Plan Limits tab, next to a Developer Plan Limits tab. For
openai/gpt-oss-20bandopenai/gpt-oss-120b, the two names Runesmith suggests, the Free tab lists 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute and 200,000 tokens per day. (Checked 2026-10-05.) - What that means: 8,000 tokens a minute is small. Large requests need another model, so keep a second one in the lane. The page says limits apply to the organization, not to each user, that you can hit any one of them first, and that cached tokens do not count. Its table is a summary with possible exceptions; the exact numbers for your account are on the Limits page in your account settings. (Checked 2026-10-05.)
- Getting a key: console.groq.com/keys. Neither the rate-limits page nor Groq’s quickstart says what the free plan needs at sign-up, for example whether a card is asked for; the key page tells you when you make the key. (Checked 2026-10-05.)
- Sources: rate limits, quickstart, checked 2026-10-05.
OpenRouter #
Checked 5 October 2026
- Free today: yes, for models whose id ends in
:free. OpenRouter’s limits page says free models are capped at 20 requests per minute. The daily cap is 50 requests while the credits you have ever purchased total fewer than 10, and 1,000 requests once they total at least 10; the page says the higher ceiling starts one credit below that threshold and counts free-model requests per UTC day. It says “credits”, not a currency, so check what a credit costs on OpenRouter’s own page. It also says a negative credit balance can produce payment errors even for free models. (Checked 2026-10-05.) - What OpenRouter says about free models: its FAQ says they have low rate limits and are usually not suitable for production use, and it names a Free Models Router,
openrouter/free, that picks a free model for you. (Checked 2026-10-05.) - Getting a key: openrouter.ai/keys. The base address Runesmith uses is
https://openrouter.ai/api/v1, which matches OpenRouter’s quickstart. None of the pages checked says whether a card is asked for when you make a key; OpenRouter’s own sign-up page shows it. (Checked 2026-10-05.) - Sources: limits, FAQ, quickstart, checked 2026-10-05.
Mistral #
Checked 5 October 2026
- Free today: partly. The Free plan on Mistral’s pricing page includes, in the page’s own words, “$10 /mo in API credits”: a monthly credit allowance rather than an unlimited, rate-limited tier. Mistral’s quickstart says that in Free mode API access is on by default with no credit card required, and that usage and rate limits apply. (Checked 2026-10-05.)
- Getting a key: console.mistral.ai/api-keys. Mistral’s quickstart “Activate Studio and generate an API key” walks through it and says the full key is shown only once. (Checked 2026-10-05.)
- Limits: Mistral’s help centre says the API enforces three limits (requests per second, tokens per minute and tokens per month), that Free mode has the lowest limits, meant for evaluation and prototyping, and that you read yours in the Admin Panel under API, then Limits. It publishes no number for Free mode. Higher tiers come from pay-as-you-go billing, and adding credits does not raise your limits. (Checked 2026-10-05.)
- Names: Runesmith suggests
codestral-latestandmistral-small-latest. Once your key is in, Runesmith lists what Mistral serves today, so trust that list over this page. - Sources: pricing, quickstart, usage and limits, help centre, checked 2026-10-05.
Models on your own computer #
Checked 5 October 2026, on each program’s own pages; addresses and commands also against Runesmith 0.1.0
This is the one way with nothing to sign up for and nothing that leaves your computer. Runesmith supports three local server programs, each on its default address, and finds a running one for you: open Thinking power and look under On this computer.
| Program | Default address | Get it |
|---|---|---|
| Ollama | http://127.0.0.1:11434/v1 |
ollama.com/download (it lists macOS, Linux and Windows) |
| LM Studio | http://127.0.0.1:1234/v1 |
lmstudio.ai/download |
| llama.cpp server | http://127.0.0.1:8080/v1 |
github.com/ggml-org/llama.cpp |
- Ollama: install it, then pull a model from a terminal:
ollama pull qwen2.5-coder:7b, Runesmith’s suggested starting model. In Runesmith, click Use Ollama, then Save and test. Ollama’s pricing page lists a Free plan at $0 that includes “Run models locally” and says running models on your own hardware is always unlimited. Its pages now also advertise cloud models (a name ending in:cloud) that run on Ollama’s servers, not on your computer: only a model that runs locally keeps everything on your machine. Ollama’s OpenAI-compatibility page shows the addresshttp://localhost:11434/v1/, the same port as in the table. (Checked 2026-10-05.) - LM Studio: load a model with a context length of at least 16384 tokens, then start its local server (Developer: Start Server, in Runesmith’s setup hint). LM Studio’s home page now leads with a newer product, Bionic, so look for the LM Studio download link; the download page offers LM Studio for Windows (version 0.4.25 when checked), macOS and Linux, plus a headless daemon called llmster. Its pricing page lists a Free plan at $0 with the line “Run local LLMs on your machine”. Its developer documentation shows the OpenAI-compatible server on port 1234, matching the table. (Checked 2026-10-05.)
- llama.cpp: start its OpenAI-compatible server with a context of 16384 tokens or more. Runesmith’s hint is
llama-server -m model.gguf --ctx-size 16384; llama.cpp’s own README currently showsllama serve -hf ggml-org/Qwen3.5-0.8B-GGUF, so follow its README for the command your build uses. The project page shows an MIT license badge, and its server README gives 8080 as the default port. If your server uses another port, add it under Any OpenAI-compatible endpoint. (Checked 2026-10-05.) - Sources: Ollama download, Ollama pricing, Ollama OpenAI compatibility, LM Studio download, LM Studio pricing, LM Studio OpenAI compatibility, llama.cpp, llama.cpp server README, checked 2026-10-05.
- What to expect: a local model is free, private and entirely yours. Small models also fail more often, especially on larger projects. In our own sealed test with a very small model, Runesmith did not beat a plain fixed scaffold, and the paper reports that as a null result. Start with small, checked steps.
- Tested how: Runesmith talks to all three through their OpenAI-compatible endpoints. At the time of writing, its own journeys exercised local servers against stand-ins that follow their documentation, which is not the same as a long run on a real local model.
No key at all: a chat window #
Under No key needed, choose A chat window (copy and paste). Runesmith writes each request, you paste it into any chat you can already use, and you paste the reply back. It suits roles that need only a few calls, such as the Checker, the Improver and drafting a plan. Every call waits for a person, so it is human-assisted, not unattended work.
Tips that save the day #
- Give every role a second model. Free allowances run out. Under Thinking power, Who does what, the first model in a role is preferred and the others are fallbacks. When one is busy or at its limit, Runesmith moves on to the next without you noticing. A different provider is best, because quotas of one provider often run out together.
- Put your best model on the Checker. The Checker proposes the acceptance checks that decide what “done” means, so a few calls matter a lot there. A chat window works well for it, while a free API model builds.
- Keys stay on your computer. A key is saved in a file in the project’s
.runesmithfolder and is sent only to the provider it belongs to. The file is not encrypted: it relies on your account’s file permissions. - Your code goes to the provider you pick. With a hosted model, your goals and your project’s code are sent to that provider. Free tiers can come with different data terms from paid ones, so read them. A model on your own computer sends nothing anywhere.
The full walk-through, with every screen of the Studio, is in the guide.