A library that makes an agent better at its job, by itself.
OURO is an open source TypeScript SDK. You give it a goal in one sentence, a source of data, and a way to score results. It writes a set of strategies to chase that goal, runs them, watches how they do, works out why the weak ones lost, and writes new ones. It repeats this on a schedule. Over time the set of strategies gets better, and you can see it getting better on a chart.
You never write a strategy yourself. You can, if you want to start from something you already trust, but the default is that the SDK invents its own. That is the part that matters. Most "self improving" tools let a model tune a few numbers inside a plan a human wrote. OURO writes the plan, judges the plan, and replaces the plan.
It runs on your machine. There is no server, no account, no hosted anything. You bring your own LLM key from Anthropic, OpenAI, Google, or a local model through Ollama. Your data and your strategies stay in a folder next to your project.
It does not care what the goal is. The reference example is a perpetual futures trading agent on BTC and ETH because that is a domain where "better" is a number you can measure. The same loop works for a support bot that should resolve more tickets, a scraper that should break less often, or a pricing engine that should sell faster at margin. Anything with a stream of inputs, a decision, and a score.
And it does not care what chain or venue you are on. Market data and order execution are plugins. The repo ships one for Hyperliquid because it has free public candles and needs no key. A plugin for any other exchange, any EVM or Solana chain, or any non trading source is about sixty lines.
The difference between an AI product and SI technology.
People use "superintelligence" loosely. Underneath the hype there is a specific idea that every serious definition shares: a superintelligence is a system whose capability rises on its own, without a human editing it. The engine of that rise is called recursive self improvement. The system looks at its own failures, changes itself, checks whether the change helped, and keeps the change if it did. Then it does it again, starting from the improved version.
Today's AI agents do not have this. They are frozen the moment you deploy them. They only change when a developer opens the code. That is why they feel like tools, not minds.
OURO is recursive self improvement at the agent level. Not at the model level, we are not training neural networks, but at the level of what the agent actually does with the model. The strategies are the agent's behaviour, and the strategies rewrite themselves. That is the mechanism, packaged so you can install it.
Four properties show up in every account of what a superintelligent system needs. OURO ships each one as code you can read:
| Property | What it means in theory | What OURO ships |
|---|---|---|
| Self model | The system holds a readable description of itself. | The live population plus every strategy ever generated, with its parents and the reason it exists. The Critic and Generator read this. |
| Self critique | The system judges its own output. | The Critic. An LLM step that reads the agent's own worst episodes and writes down why they went wrong, before anything is changed. |
| Self modification | The system rewrites itself. | Seed, mutate, crossbreed, fresh write. New strategies enter and old ones leave without a human edit. |
| Corrigibility | The system can be bounded and switched off. | Sandbox, holdout data, risk caps, an approval switch, and one command rollback. |
The fourth one is why the tagline is "superintelligence you can switch off". The thing that makes self improvement scary is that you lose the ability to stop it. OURO is built so that you never do. Every guard is on by default and the switch is a single config flag.
The words you will see everywhere.
'paper' selects the built in paper executor. A venue, a contract, or a webhook when you go live.
One cycle, from the first episode to the next promotion.
A cycle runs every time each live strategy has collected a set number of new episodes, fifty by default (the perp example sets forty). Here is what happens inside one, in order.
Seed, once
On the very first start there is no population. The Generator reads your goal and the list of primitives you registered and writes eight strategies. Each one is compiled and run inside the sandbox before it is allowed to trade. If any fails, it is thrown out and another is requested. The seed generation becomes the baseline that every later Capability Index is measured against.
Collect
The loop pulls the most recent episodes for each strategy and splits them by time. The older seventy percent is the training slice. The newest thirty percent is holdout. The split is always by time, never random, because random splits leak the future into the past.
Rank
Every live strategy is replayed on the pooled holdout slice and scored. The bottom quarter is marked for retirement. Marked does not mean gone. A strategy only retires if something better replaces it, so the population never shrinks.
Diagnose
The Critic receives the ten worst episodes of each weak strategy and the five best of each strong one, summarised as compact rows: time, asset, side, result, how long it held, hour of day, volume ratio, and the three most relevant feature values. It writes back at most five patterns and a summary under sixty words. Something like "losses cluster between 00:00 and 04:00 UTC when volume is below average".
Generate
The Generator produces candidates in three ways. For each weak strategy it writes a mutation, which changes a parameter inside its declared range or swaps one feature. For the two strongest it writes a crossbreed, which takes the entry logic of one and the exit and filter logic of the other. And it writes one fresh strategy from the diagnosis that is structurally unlike anything alive. Every candidate goes through the sandbox and the guards before it is allowed anywhere near the data.
Trial
Candidates are replayed on the training slice. A candidate must beat the median training score of the live population by a margin, five percent by default. Anything that only ties is dropped.
Validate
Survivors are replayed on the holdout slice, which no LLM step has seen. A candidate that won on training and loses on holdout is curve fit. It learned the past instead of the pattern. It is rejected with the reason "holdout" and that rejection is written to history so you can see how often it happens.
Promote
Each survivor replaces the weakest remaining live strategy, weakest first. The population stays at eight. The new strategy's code, parameters, parents and rationale go into history along with the scores from both trials. The Capability Index is recomputed for every live strategy and the takeoff file is updated.
If nothing survives, the cycle records "no change" and the population is untouched. That is a normal and common outcome. The loop is allowed to find nothing. What it is not allowed to do is promote something that only looked good on data it had already seen.
Let it run free, or narrow it where money says you must.
Generate mode, the default
You pass a goal, primitives, a source, an executor, your assets and timeframe, and a scorer. Nothing else. The SDK invents everything. The only limits are the guards. This is the mode the SDK is built for, and the one you should start with on paper, because it finds things you would not have thought to try.
Seeded mode
Pass seed: ['./my-strategy.ts'] with one or more files that follow the strategy contract. They become the first members of the first generation, and the Generator invents the rest up to the population size. Useful when you already have something that works and want the loop to improve it rather than start over. The SDK treats your file like any other strategy. It can be mutated, crossbred and retired.
Constrained mode
Three optional knobs, each independent, each per key:
allowrestricts which feature keys the Generator may use. If you only want entries built from a handful of indicators, list them.boundscaps the range of a parameter. The Generator declares its own bounds for every parameter it invents; yours apply on top of them, so a value must satisfy both.freezenames parameter keys that may never change from the parent. A stop multiplier or lookback length is the usual one.
Use these only where safety or capital says you must. Every constraint removes a place the loop could have found an improvement.
Every brake, on by default.
Self improving code with no brakes is a liability. These are the ones OURO ships. The sandbox and the replay risk caps cannot be turned off; the others are config fields with safe defaults.
Sandbox
Every generated strategy runs inside an isolated virtual machine. It gets 64 MB of memory, 50 milliseconds per decision, 2 seconds to load, and nothing else. No network, no file system, no access to your process. No host object is in scope, only the language's own builtins such as Math, Number and JSON. Before a strategy even reaches the VM, its source is scanned and rejected if it contains import, require, fetch, process, globalThis, eval, Function, timers, sockets, constructor walks, or an unbounded loop. A model that produces bad code cannot escape this. It just gets rejected.
Two backends enforce the same limits. isolated-vm is used when it loads. If it does not, each strategy runs in its own worker thread under node:vm with the same memory cap and timeouts, so the guard holds on machines where the native module will not build.
Holdout
The newest slice of episodes is never shown to the Critic or the Generator. They can only learn from the older slice. Then the judge checks their work on the slice they never saw. This is the single most important guard against a loop that convinces itself it is improving.
Margin
A candidate has to beat the population by a set margin on training, and beat the specific strategy it would replace on holdout. Ties and near ties are not promotions.
Risk caps
If replaying a candidate shows a drawdown above maxDrawdownPct, or any position size above maxPositionPct, it is rejected before trial. The loop cannot promote its way into a bigger bet than you allowed.
Every rejection carries a reason and is kept in history. In the order they are checked: sandbox, bounds, freeze, allow, unknown feature, drawdown, size, then from the cycle itself train margin, no slot and holdout. If most of your candidates die on holdout, the goal or the scorer needs work, not the guards.
The switch
Set requireApproval: true and a cycle stops at the promotion step. It returns "pending" with the full diagnosis, the candidates and both sets of scores. Nothing changes until you call approve(cycle) or reject(cycle), from code or from the command line. The default is off. ouro run --live refuses to start unless it is on in the config, and there is no flag to override that.
Rollback
ouro rollback 7 restores the population exactly as it was after cycle 7. Later strategies are marked rolled back, not deleted. History is append only. You can always see what the loop did, and you can always undo it.
The one number that says whether it worked.
Capability Index, or CI, measures a strategy against the seed generation on data it has not learned from. The score used is the strategy's own holdout score: the mean score over the newest slice of its own recent episodes. The baseline is the seed generation's mean of that same number, stored once at the first cycle that has it and never changed. CI is the gap between the two, divided by the size of the baseline (or by the seed generation's mean absolute score when that is larger, so a baseline near zero does not blow the ratio up), so +0.20 means twenty percent better than the first generation. A strategy gets a CI once it has at least three of its own holdout episodes; until then it does not count toward the population average.
Why own episodes and not the pooled replay the promotion step uses? Pooled scores move with whatever the live population happens to be trading, so a ratio against them wobbles even when nothing changed. Own episode scores do not, and the takeoff curve stays readable.
Population CI is the average across the live strategies that have a CI. Plotted per cycle it is the takeoff curve. Velocity is how much population CI rose since the last cycle. When velocity stays under one percent for three cycles in a row the SDK reports a ceiling. That means the loop has squeezed what it can from the primitives it has. The fix is a human one: register another pack, loosen a constraint, or widen the goal. Then it climbs again.
All of this is written to .ouro/takeoff.json after every cycle and printed by ouro takeoff. It is the file you screenshot, and later the file a public registry can verify.
Ten minutes to a population.
npm i @ourointelligence/sdk @ourointelligence/source-hyperliquid
# pick one provider and set its key
export OURO_LLM=anthropic
export ANTHROPIC_API_KEY=sk-...
# or OURO_LLM=openai / gemini / ollama
cd my-project
npx ouro run --paper --assets BTC,ETH --tf 15m
Node 20 or newer. Keys can also live in a .env file in the project folder. The CLI reads ouro.config.ts from that folder, so write one first (next section). The run command seeds a population, pulls three hundred bars of history so the indicators have something to work with, trades through the last backfill bars on paper if you set that field, then subscribes to live candles and keeps trading on paper. The first cycle fires once every strategy has cycleEvery closed trades, fifty by default. Leave it running and watch ouro takeoff in another terminal.
ouro.config.ts, field by field.
import { primitives, type LoopConfig } from '@ourointelligence/sdk';
import { hyperliquid } from '@ourointelligence/source-hyperliquid';
export default {
goal: 'Maximise realised PnL after fees on BTC,ETH 15m, max drawdown 8%',
primitives: [primitives.ta, primitives.volume, primitives.time],
source: hyperliquid({ assets: ['BTC','ETH'], tf: '15m' }),
executor: 'paper',
score: (ep) => ep.outcome.pnl - ep.outcome.fees - 0.5 * ep.outcome.drawdown,
assets: ['BTC','ETH'], tf: '15m',
population: 8, cycleEvery: 40, holdout: 0.3, margin: 0.05,
guards: { maxDrawdownPct: 8, maxPositionPct: 10, requireApproval: false },
ensemble: 'weighted',
} satisfies LoopConfig;
| Field | Default | What it does |
|---|---|---|
goal | required | One sentence. The Generator reads it verbatim. |
primitives | required | Packs the Generator may build from. Built in: ta, volume, time. orderbook and onchain ship as interface stubs. |
source | required | Where inputs come from. Any plugin that implements Source. |
executor | required | Where decisions go. 'paper' selects the built in paper executor, which fills at next bar open with a fee and slippage model. ouro run without --live forces 'paper'. |
score | required | Episode to number. Higher is better. The only thing the loop optimises. |
assets | required | Assets the source streams, for example ['BTC','ETH']. --assets overrides it. |
tf | required | Bar timeframe, for example '15m'. --tf overrides it. |
population | 8 | How many strategies stay alive. |
cycleEvery | 50 | New episodes per strategy before a cycle fires. |
holdout | 0.3 | Share of episodes kept hidden for validation. |
margin | 0.05 | How much a candidate must beat the population median by. |
guards | see text | maxDrawdownPct, maxPositionPct, requireApproval, maxProposalsPerCycle. |
seed | none | Paths to your own strategies to use as the first generation. |
allow, bounds, freeze | none | Constrained mode knobs. See section 05. |
ensemble | 'weighted' | How eight decisions become one order. Weighted sums each side by CI plus one and acts if a side holds more than 55 percent. 'majority' or 'none' also available. |
dispatch | 'per-strategy' | One order per strategy decision, or 'ensemble' for a single order carrying the combined decision. Use ensemble when going live. |
backfill | 0 | Bars of recent history to trade through on paper before subscribing live, so the first cycles happen in minutes instead of days. |
autoCycle | true | Fire a cycle automatically when every live strategy has cycleEvery new episodes. Turn off to cycle only on a timer or by hand. |
retireShare | 0.25 | Share of the population marked weak each cycle, at least one. |
llm | from env | Override the adapter chosen by OURO_LLM. |
dir | .ouro | Where everything is stored. |
warmupBars | 300 | History pulled before the first live bar so indicators are warm. |
What the SDK writes, and what you can hand it.
Every strategy, generated or yours, is a TypeScript module with exactly four exports. No imports are allowed. The function must be pure: same input and parameters, same decision, every time. That is what makes replay honest.
// .ouro/population/s-0412.ts (crossbreed, cycle 7, parents s-0203 x s-0311)
export const params = { hullLen: 34, atrStop: 1.8, rr: 2.2, minVolRatio: 1.5 };
export const bounds = {
hullLen: { min: 9, max: 55, step: 1 },
atrStop: { min: 1, max: 3, step: 0.1 },
rr: { min: 1, max: 4, step: 0.1 },
minVolRatio: { min: 0.5, max: 3, step: 0.1 },
};
export function decide(x, p) {
if (x.features['time.hour'] >= 0 && x.features['time.hour'] < 4) return null;
if (x.features['volume.volRatio'] < p.minVolRatio) return null;
if (x.features['ta.hull34.crossUp'] && x.features['ta.qqe14.value'] > 0)
return { side: 'long', size: 0.05, stop: p.atrStop * x.features['ta.atr14'], tp: p.rr };
return null;
}
export const describe = 'Hull 34 cross up with QQE confirm, skips 00-04 UTC, needs 1.5x volume';
Feature keys are flat strings the primitive packs expose: ta.hull21.crossUp, ta.atr14, ta.rsi14, volume.volRatio, time.hour, time.isWeekend and so on. Each pack's describe() lists every key with a sentence of documentation, and that list is exactly what the Generator is shown. A strategy may only read keys a registered pack produces; anything else is rejected as an unknown feature.
A decision is null for do nothing, or an object with side (long, short or flat), size as a fraction of equity, and optional stop and tp. Outside trading these map to whatever your executor understands; the support bot example uses long for resolve and short for escalate.
You can open any generated file and read it. The loop runs the code recorded in history.json, so editing the file in place changes nothing. To hand it an edited version, pass the file in seed when you start a fresh population; it enters with origin "user" and then has to earn its place like everything else.
Everything from the terminal.
The CLI reads ouro.config.ts from the current folder.
| Command | What it does |
|---|---|
ouro run [--paper|--live] [--assets] [--tf] | Seeds if needed, warms up, backfills, subscribes, trades, cycles automatically. --live refuses to start unless the executor is a plugin and requireApproval is on. |
ouro cycle | Runs one cycle now on whatever episodes exist. |
ouro start --every 1h | Same as run, and also cycles on a timer as well as by episode count. |
ouro population | Table of live strategies: id, origin, cycle born, holdout score, CI, description. |
ouro history [--json] | Every strategy ever generated and every cycle result. |
ouro explain s-0412 | A paragraph in plain words: what it does, where it came from, what the Critic said that led to it. |
ouro takeoff | The curve as text: cycle, population CI, best CI, velocity, and a ceiling flag. |
ouro rollback 7 | Restore the population as it was after cycle 7. |
ouro approve 8 / ouro reject 8 | Finish a pending cycle when requireApproval is on. |
ouro export | Write ouro.json with population, history, takeoff and goal, for dashboards or a registry. |
Any venue, any chain, any domain, in about sixty lines.
Four interfaces. Implement one, pass it to the loop, done.
interface Source {
name: string;
subscribe(o: { assets: string[]; tf: string }): AsyncIterable<Bar>;
history(o: { assets: string[]; tf: string; bars: number }): Promise<Bar[]>;
}
interface Executor {
name: string;
place(d: NonNullable<Decision>, x: Input): Promise<{ orderId: string }>;
onClose(cb: (strategyId: string, outcome: Outcome) => void): void;
}
interface PrimitivePack {
name: string;
compute(bars: Bar[], i: number): Record<string, number | boolean | null>;
describe(): Array<{ key: string; doc: string }>;
}
interface LLM {
name: string;
complete(req: { system: string; user: string; json: true; maxTokens?: number }): Promise<string>;
}
A Source turns anything into a stream of bars. For trading that is candles. For a support bot it is tickets with a timestamp. A Bar is just a time, an asset, a timeframe, and open, high, low, close and volume, and your primitive pack decides what features to derive from it.
An Executor takes a decision and reports back when the episode closes. Two optional hooks make real fills possible: onBar(bar) is called with every new bar before any decision is made on it, which is where you simulate fills, track stops and mark positions, and stop() is called on shutdown so you can flatten. The loop tags every order with the strategy that made it in x.meta.strategyId; put the asset in outcome.raw.asset when you report a close. The paper executor is in the SDK. A live one for a venue is where you put the API calls, the signing, and the position tracking.
A PrimitivePack is how you teach the Generator about your domain. The describe() output is pasted into its prompt, so write those docs the way you would explain the field to a smart colleague who has never seen your data.
An LLM adapter is forty lines. Four ship with the SDK, and any server that speaks the OpenAI chat format works through OPENAI_BASE_URL with no code at all. If a provider returns JSON from a system and user prompt, it can be an adapter. Pick the model with OURO_MODEL.
The full guide with one worked example per interface is in docs/plugins.md in the repo at github.com/ourointelligence.
Two in the repo. One is trading, one is not.
examples/perp
The reference. Goal is realised PnL after fees with a drawdown penalty on BTC and ETH 15 minute perps. No strategy supplied. Source is Hyperliquid public candles, executor is paper, primitives are the three built in packs. It seeds eight, trades through a backfill of recent history so the first cycles land in minutes, then goes live on the websocket and cycles every forty closed trades. The paper executor fills at the next bar's open with 3.5 basis points of fee and 2 of slippage, and reports pnl, fees and drawdown as a percent of equity. The takeoff.json committed next to the example says in its README exactly which run produced it and whether a model or a stand in answered the prompts.
examples/support-bot
About a hundred lines. A synthetic source yields tickets as bars, a tiny text primitive pack exposes length, a sentiment stub, a category, urgency and prior contacts, the decision is resolve or escalate, and the scorer rewards resolving without escalation and penalises long holds. It exists to prove the loop is not a trading bot in disguise.
Paper first. Then the switch stays on.
- Run on paper until population CI has risen for at least three cycles and stayed up. Read the rejections too. If most candidates die on holdout, the goal or the scorer needs work.
- Set
maxDrawdownPctandmaxPositionPctto numbers you can live with. The loop cannot exceed them. - Turn
requireApprovalon. Read each pending cycle's diagnosis and rationale before approving. It takes a minute and it is the whole point. - Pass
--livewith a real executor plugin anddispatch: 'ensemble', so one combined order goes out instead of eight. The CLI will not start live withoutrequireApprovalon. - Keep
ouro rollbackin your head. It restores the previous population; a running loop picks it up on its next start.
Everything is a plain file you can read or commit.
ouro export. Population, history, takeoff and goal in one file for anything downstream.
Cents per cycle, zero per month.
Episodes are summarised before they reach any prompt, so the Critic prompt stays under about four thousand tokens, and each Generator call (two mutations, one crossbreed and one fresh strategy per cycle at the default population, capped at six) is of a similar size. On a mid tier model that is a few cents. The SDK itself is free, open source under MIT, and runs on your hardware. There is nothing to subscribe to.
The only recurring cost is your own LLM key, and you control the cadence. Forty episodes per cycle on a fifteen minute chart is a handful of cycles a day.
The SDK has no token. Say that plainly.
OURO the library is free, tokenless, and works on every chain or no chain. Nothing in it checks a balance, gates a feature, or pays a fee. If a token exists around the project, it is a separate community asset tied to the superintelligence thesis, not a key to the product.
Things a token could touch later, all optional and all chain agnostic: a public Capability Registry where agents publish their takeoff curves so anyone can verify a track record, fees for hosted loops if those ever exist, and votes on which primitive packs enter the public library. None of that is in v0.1 and the SDK will keep working without it.
The ones that come up.
Does it train a model?
No. It uses a model you already have access to, through its normal API, to write and critique code. The thing that improves is the agent's behaviour, not the weights.
Can it invent a strategy that loses everything?
It can invent one. It cannot promote one. Risk caps reject anything whose replay breaches your drawdown or position limits, and the sandbox means the code cannot reach your money directly. The executor is the only thing that touches a venue, and you wrote or chose it.
Why eight strategies and not one?
One strategy that improves might be lucky. Eight that keep replacing each other on data none of them has seen is a pattern. Crossbreeding also needs parents.
Why does it reject so much?
Because most ideas do not survive new data. A healthy run rejects more than it promotes. If it promoted everything that looked good on training, it would be a curve fitting machine.
Will it work for my non trading problem?
If you can describe the goal in a sentence, turn your inputs into a stream, and score an outcome with a number, yes. Write a source and a primitive pack and the rest is the same loop.
Is it chain specific?
No. The SDK never talks to a chain. Sources and executors do, and those are plugins. Hyperliquid is the shipped example because it is free and keyless, not because anything depends on it.
What happens when it stops improving?
It tells you. Velocity under one percent for three cycles raises a ceiling flag. Add a primitive pack, loosen a constraint, or widen the goal, and it climbs again.