Introducing Runesmith

Democratizing agentic coding.

Open source. Any model: free, paid or on your own servers. Your code, your models, your digital sovereignty.

Apache-2.0  ·  A research release  ·  Launchers for Windows, macOS and Linux

The Runesmith mark: a rune carved through a gold stone block, its leg breaking out of the corner.

01The film

What Runesmith is, from its maker.

Lars Horpestad introduces Runesmith: what it does, how it works with free, paid and local models, and where to start.

The film plays from YouTube, and only after you press play: until then this page sends nothing to YouTube. Watch it on YouTube.

Read the transcript

Hi, my name is Lars Horpestad. I'm the CEO at AI ThinkLab. With support of Innovation Norway, we are pleased to announce that we today are releasing our new model named Runesmith.

Runesmith is a non-traditional type of model that uses various sources of inference in order to build on a code basis, new or existing, improve on these, and operate entire digital systems.

It has been built together with an extensive paper named Beyond the Model that we are also releasing together with the open source project named Runesmith.

In that paper, we document how Runesmith can build, improve, and operate entire digital systems. We also document how it can improve on its own digital system.

So it is a nudge in the direction of RSI, or recursive self-improvement. It does take the inference away from the model, which is a little bit in the times these days, as there are so many different sources for inference,

both free from sources like Google AI Studio or NVIDIA, to more expensive such as Anthropic or OpenAI's models.

Many companies and institutions want to use models locally, on their own servers, in their own houses, and we have set this system up such that you can use APIs from, be that your own Ollama project or elsewhere.

Yes. So together with the model, which is open source, you can download it. We can't wait to see what you build.

There is also a guide, an extensive key by key guide that shows you how to use Runesmith.

And if you really want to get a head start using this model and building with it, we also have an extensive guide available on Amazon with a link below also for that.

Thank you for your time, and have a good day.

02Get Runesmith

Download it and start the Studio.

Runesmith is a non-traditional model: it has no weights of its own. It is a system that wraps whichever model you have and lets that model build software.

  1. Get it

    Open the repository on GitHub and choose Code, then Download ZIP, and unzip it. Or clone the repository with Git. You need Python 3.11 or newer and a web browser; nothing else is installed, because Runesmith uses only Python’s standard library.

  2. Start the Studio

    Runesmith Studio runs on your own computer and opens in your browser, through a private link.

    Windows

    Double-click
    Runesmith.cmd

    macOS

    In Terminal, in the Runesmith folder:
    sh Runesmith.command

    Linux

    In a terminal, in the Runesmith folder:
    sh runesmith.sh

  3. Give it a model

    Open Thinking power and add a model: a free key from Google AI Studio or NVIDIA, a paid one from Anthropic or OpenAI, a model on your own computer, or just a chat window you copy and paste into. The free inference guide walks through each, starting with the free ones.

  4. Plan and build

    Write a goal, draft a plan, and build it milestone by milestone. You read every draft before it touches your files. The guide goes from the first start to a finished project.

03What it does

Plan, build, check.

Around the model you already have, Runesmith plans the work, has the model write it, and checks every step.

Plan

Milestones from your goals

Write your goals and a brief in your own words. The Planner turns them into milestones, each with a plain “done when”. Edit anything.

Build

One milestone at a time

A model writes the code and you read the draft. Nothing is written to your files until you apply it, or until you allow automatic apply for checked builds in folders you name. Every write can be undone.

Check

Checks you can read

Acceptance checks are written in plain sentences, and you approve them before they count. Each draft is tried against them on a throwaway copy of your project.

Any model

Your code, your models

Runesmith keeps what it learns in records it owns, on your computer, not inside any one model. Use a free model, a paid one, or one you run yourself.

  • Free: Google AI Studio, NVIDIA
  • Paid: Anthropic, OpenAI
  • Your own: Ollama, LM Studio (tested with stand-ins so far)
  • Or a chat window
Improves itself

Within limits, and honestly

Runesmith can also work on its own system. It reads its own records, finds where it struggles, asks a model to rewrite the part responsible, and keeps the change only if it does better on work its author never saw. That is a nudge towards recursive self-improvement, and only a nudge: in two sealed tests, a change made this way solved more tasks than the version before it and then more than the repair step Runesmith shipped with, though a hand-written rule did better than the new version, and most of the other sealed tests did not show their effect. No study shows cumulative improvement. In the Studio, self-improvement stays off until you switch it on.

Read what the paper says, including what did not work.

04The record

The numbers, and what they are.

Every figure comes from the paper’s record or the project’s files. Studies that did not work are counted too.

33 days

From the research program’s first file, on 3 September 2026. Runesmith’s own repository begins later, on 25 September.

23 sealed or registered comparisons

The program has run 23, and the nulls and the failures count with the passes. Most did not show their effect. The first sealed self-improvement test still passes after correcting for all 23; the second, run on 6 October, also passed, and in it a hand-written rule did better than the self-improved step.

324 sealed sessions, released

Every session outcome of the first sealed self-improvement test, released with the script that recomputes its exact test. The second test’s sessions, seals and Bitcoin timestamp proofs are released too.

13 out-of-the-box journeys

Each run start to finish by an AI operator, with every finding recorded. This is not a user study.

1,924 automated tests

In the released repository’s test suite.

What these numbers do not show: no study compares a model working alone with the same model inside Runesmith, so the record does not show that Runesmith makes any model better than it is alone, and in the second sealed test a hand-written rule did better than the self-improved repair step. The paper says so too.

05Built with Runesmith

Made with Runesmith Motion

Runesmith Motion is a video program that Runesmith built under an AI operator’s direction, using free inference.

In practice, free-tier models were the last author of every applied change, through Runesmith, milestone by milestone. A human-directed AI operator planned and audited the work, and stepped in when it stalled. The paper says plainly where it fell short.

06The paper

Beyond the Model

Cover of the paper Beyond the Model, with the Runesmith mark and the words A paper from AI ThinkLab.

Beyond the Model: Runesmith, an Open Runtime That Improves Software and Itself, and the Instrument–Substrate Hypothesis

Lars O. Horpestad, AI ThinkLab. Preprint, not peer reviewed.

The paper asks where an AI agent’s competence lives: in the model it calls, or in tested, inspectable state that the system keeps for itself and can still use when the model is replaced. It presents Runesmith as the runtime built to test that idea and reports every sealed result, including the nulls and the failures; the central hypothesis has not yet been tested to its own definition. The paper is 17 pages. The technical report, 90 pages, gives the full method, the source of every number and the complete negative record.

The evidence, in graphs

Four of the paper’s figures and one from the technical report (the Runesmith Motion chart), each with the numbers it shows. Full captions and sources are in the two documents: the paper (PDF) and the technical report (PDF). Select a graph to open it full size.

In the figures, g0 is the shipped repair step, B the earlier generation, C7 the newer one, and SR7 and LOC1 the names of the two sealed tests.

Two-row dot chart of strict repair success in two sealed tests. First test, 54 fresh tasks: the earlier generation 35 of 162 sessions, the newer generation 57 of 162, exact p 0.000845. Second test, 56 fresh tasks from three large public repositories: the shipped repair step 17 of 168, the newer generation 42 of 168, exact p 0.00278, and, as a reference, a hand-written trace-aware rule 65 of 168.
The sealed two-step recordStrict repair success under two sealed, preregistered tests, each with one free repair model and one envelope for every arm. First, on 54 fresh tasks, the newer generation of the repair step repaired 57 of 162 sessions against 35 of 162 for the generation it replaced (exact one-sided p = 0.000845). Then, on 56 fresh tasks from three large public repositories, it repaired 42 of 168 against 17 of 168 for the repair step Runesmith shipped with (exact one-sided p = 0.00278, a pass at its declared level). Shown as it is: in that second test a hand-written rule that reads the failing test’s trace repaired 65 of 168, more than the newer generation. The rule was written after the newer generation existed, the test used one repair model and synthetic single-line regressions, and its replication on a second model finished one task and could detect no difference. Both tests were sealed before they ran, and the second one’s fingerprints were timestamped in the Bitcoin blockchain hours before its analysis.
Dot-and-interval chart of strict repair success of the predecessor and successor generations. Both repositories: 35 of 162 against 57 of 162 sessions. One repository: 28 of 150 against 48 of 150. The other: 7 of 12 against 9 of 12. Exact one-sided p over 54 tasks is 443/524288.
The sealed testThe first sealed test in detail. On 54 fresh tasks, the successor generation repaired 57 of 162 sessions against the generation it replaced: 35 of 162 (exact one-sided p = 0.000845). One repair model, an equal envelope; the intervals drawn are descriptive. The second test, above, compared the successor with the shipped version.
Dot-and-interval chart of first-round strict repair success on the same 17 tasks from two public repositories: the shipped repair organ 5 of 17, the first later generation 5 of 17, the second later generation 8 of 17, with each step's gain and interval. No step is significant.
Three generationsFirst-round strict repair success of the shipped repair organ and two later generations on the same 17 fresh, synthetic single-line regressions in two public Python repositories: 5, 5 and 8 of 17. Descriptive: no step is significant. Not drawn: with a simple fixed script in place of the repair organ, the same model repaired 9 of 17, against 8 of 17 for the second later generation (no detectable difference, the script ahead in direction).
Dot plot of SWE-bench Verified scores: Kimi K3 93.4, Gemini 3.7 Flash 80.8, 3.8 Flash 80.0, 3.6 Flash 79.6, 3.5 Flash 78.8, and Claude Opus 4.5 (Thinking), the paid model that authored the change tested in the sealed test, 76.4, the lowest of the six.
A public benchmarkOn SWE-bench Verified under one harness, the paid model that authored the sealed-test change (76.4%) was not stronger than free-tier Gemini Flash models (78.8% to 80.8%) or the open-weights Kimi K3 (93.4%). One public benchmark, not a measurement made by this project, and not the repair task of the sealed test.
Step chart of 54 milestones applied to Runesmith Motion from 28 September to 5 October 2026, with per-day counts (2, 27, 0, 13, 10, 0, 0, 2), two shaded unattended stretches, and a shaded interval from 2 October to the second stretch in which no milestone was completed.
Runesmith MotionMilestones applied to Runesmith Motion, 28 September to 5 October 2026: 54 at the read of 5 October, 11:20 UTC, with the two stretches that meet the lead investigator’s definition of unattended operation shaded. The owner in this project was an AI operator, and Runesmith has not completed the project on its own.

07Learn Runesmith

A free guide, and a book.

The book · paid

A head start, from Lars

Lars Horpestad’s own guide to building with Runesmith, for a head start. Coming to Amazon within a few days; the link will be here as soon as it is live.

08Coming soon

Planned for a later release. Not in this one.

Release 1 keeps everything Runesmith learns on your own computer. Nothing is shared between users, and nothing is uploaded. What follows is a plan, and it has open design questions (privacy, abuse, fair accounting), not a feature you can switch on today.

Planned · opt-in

Shared upgrades

Improvements that Runesmith wrote for itself could be offered between people through the open-source repository, only if you choose to take part. Each upgrade would have to pass the same evidence checks on the receiving machine before it could be used.

Planned · opt-in

A crowd-sourced, self-improving ecosystem

You would choose how much of your own inference goes into improving Runesmith itself. A scoreboard and legends would recognise those who give the most, and everyone who opts in would receive the upgrades that this shared effort earned.

The mission

Imagine being a kid with just a laptop and a Google account, and getting this gift.

Credits

AI ThinkLab

AI ThinkLabMakes and releases Runesmith.

Veristria

Empowered by VeristriaHosting and security partner.

Innovation Norway

With support from Innovation Norway

Contact: contact@aithinklab.com