Documentation

How Genesis is designed and measured

A reference for the project's scope, method, instrumentation and safeguards. These documents are the canonical description of how the study is run: what data the model may use, how its judgements are recorded and scored, how the wallets track progress once live capital begins, and what protections bound the experiment at every stage.

01

Overview

Genesis studies a single question: can an AI model act as an effective long-term trading assistant across both crypto and equity markets?

The project treats the model as an analyst rather than an execution engine. It reads live market data, states what it believes is happening, and commits that view to a permanent log before outcomes are known. Performance is then judged against those recorded views instead of against a narrative written afterwards. This distinction matters more than any model choice, because it is the only part of the design that cannot be quietly revised when results disappoint.

The study is deliberately plain. There is no proprietary data feed, no exotic strategy, and no claim that the model possesses insight unavailable to a careful human. The hypothesis under test is narrower and more useful: that a well-instrumented AI system, constrained to public information and forced to commit its reasoning to writing, can provide durable decision support to a human operator over a multi-year horizon.

Everything on this page is the reference version of the project's rules. Where the site and these documents disagree, these documents govern.

02

Scope

Two market classes are in scope. Crypto covers Solana, Bitcoin and Ethereum, chosen to represent a high-throughput execution environment, a slow long-horizon reference asset, and a rich on-chain settlement ecosystem respectively. Equities cover large, liquid listed names and broad index exposure, chosen so that liquidity and information quality are never the explanation for a result.

Out of scope: leverage, derivatives, market making, high-frequency strategies, short-dated event trading, and anything requiring privileged or non-public data. Each exclusion exists for a reason. Leverage and derivatives would let luck masquerade as skill for longer. High-frequency approaches would measure infrastructure rather than judgement. Privileged data would make the study impossible to replicate or audit.

Genesis is deliberately slow. Decisions are considered on horizons of weeks to quarters, and the system's written assessments reflect that cadence. A correct call made for the wrong stated reason counts against the model, not for it, because reasoning quality is the quantity under study.

03

Method

Inputs are limited to publicly available market data and public disclosures: prices, volumes, on-chain metrics where relevant, filings, and scheduled macroeconomic releases. No alternative data, no order-flow feeds, no sentiment products whose provenance cannot be inspected.

Each cycle produces a written assessment with a fixed structure: what changed since the last assessment, what the model expects over the stated horizon, how confident it is, what evidence would change its mind, and what would prove it wrong. The fixed structure exists so that assessments can be compared to each other and scored mechanically rather than charitably.

Assessments are immutable once written. Revisions are appended as new entries that reference the entries they supersede, so the full sequence of reasoning, including its corrections and hesitations, stays visible. This is the core record of the study. Every later claim about performance, calibration or usefulness traces back to it.

Scoring happens on a lag. An assessment is only judged once its stated horizon has fully elapsed, and interim movements, however dramatic, are not treated as evidence in either direction. This removes the temptation to grade the model on the most flattering available slice of time.

04

Wallet instrumentation

The project maintains three wallets: one on Solana, one holding Bitcoin, and one on Ethereum. They exist so that behaviour, transaction cost and risk can be attributed per chain rather than blended into a single number that hides where things went right or wrong. Each chain stresses something different: Solana stresses execution speed and fee sensitivity, Bitcoin stresses patience across long quiet stretches, and Ethereum stresses cost-aware rebalancing in an ecosystem where every action has a visible price.

The wallets also serve a transparency function. Once the study reaches the constrained live capital phase, the on-chain record of each wallet is publicly inspectable by anyone, without needing to trust any report this site publishes. The addresses and a running summary of activity are maintained here in the documentation, so interested readers can track progress independently and verify that reported results match what actually happened on-chain.

Until that phase begins, the wallets are read-only observation points with no published balances, because there is nothing meaningful to report. When live allocations start, each wallet carries its own hard position and loss limits, and every transaction is recorded alongside the written assessment that motivated it. A transaction without a motivating assessment is treated as a protocol violation and disclosed as one.

Solana
teTUaGfTMoej1rcJ78mkwiA7RF6LAVFdoG1kjbKwU8p
Bitcoin
bc1qmsk0enm39t702n6px6fmknrn7rmxwfw8fm7ky5
Ethereum
0x2d450e0ce94c1451d7f41c9fe53147cfe4fa4f7c

05

Safeguards

No autonomous capital deployment takes place before the paper-intent phase has produced a record long enough to review honestly. A good quarter is not enough; the paper record must survive different market conditions before live capital is considered at all.

When live capital is enabled, it is bounded in three independent ways: per-wallet caps that limit total exposure, per-period loss limits that pause activity when breached, and a manual stop that halts all activity immediately. The three mechanisms are independent so that no single failure can disable all of them at once.

Custody, keys and limits stay under human control at all times. The model can propose; it cannot widen its own constraints, move its own limits, or obscure its own record. Any change to the limits is made by the operator, in writing, before it takes effect, and the change itself becomes part of the published record.

06

Evaluation criteria

Genesis is judged on four things. Calibration: does stated confidence match outcomes, so that a view expressed at high confidence is right more often than one expressed tentatively. Consistency: does the reasoning hold across regimes rather than only in the conditions it was tuned on. Cost awareness: does the system respect fees, spread, slippage and taxes rather than ignoring them in its arithmetic. Usefulness: does a human operator make measurably better decisions with the system than without it.

Risk-adjusted return matters, but it is not the headline metric, because it is the easiest number to get lucky on. A system that outperforms while being unexplainable is treated as a failed result for this project's purposes: its operator cannot know when to trust it, when to doubt it, or when to turn it off, which is precisely the failure mode a long-term assistant must not have.

Each evaluation criterion has a written scoring rubric defined before the relevant phase begins. Rubrics are not adjusted once scoring is underway. If a rubric proves to be badly designed, that is reported as a finding about experimental design rather than repaired mid-stream.

07

Publication

Findings, including negative ones, are published as the phases complete. Interim updates are posted by the author on X, and substantive write-ups appear with the full supporting record: the assessments, the scores, and the wallet history where applicable.

The standard for publication is completeness, not favourability. A phase that produced ambiguous or disappointing results is published with the same structure and the same detail as a successful one. Silence after a phase begins is itself a reportable anomaly, and the publication schedule is stated in advance so that absence of news cannot be used to hide results.

Nothing published here is investment advice, a solicitation, or an offer of any financial service. The project is a research study, and its outputs are descriptions of an experiment, not recommendations to act.