Project retrospective at the 30 July 2026 evidence cutoff. AI OS began with a familiar ambition: collect promising trading ideas, translate them into code, optimize them, and find a system worth operating. The research did produce engines, interfaces, datasets and millions of evaluations. It did not produce the shortcut the original ambition implied.
That failure changed the project in a useful way. The central question stopped being “Which bot should we run?” and became “What evidence would make any strategy safe enough to trust—and can we produce the same verdict again?”
This is not a claim that every component of AI OS is finished, that every tested idea is worthless, or that every broker, bot author or strategy vendor is unsafe. It is a narrower and more defensible conclusion: the current evidence does not support an executable endorsement, while the machinery built to reach that conclusion is itself becoming a product.
One project, four product layers
The names around the project describe different jobs. Treating them as four separate startups would create confusion and duplicate infrastructure. The coherent architecture is one evidence system with four visible layers:
| Layer | Role | What it must never imply |
|---|---|---|
| Tvijo AIOS | Research and orchestration: ingestion, registries, datasets, simulations, optimizers, evidence objects, memory and policy gates. | That an AI-generated answer is evidence or approval. |
| FinEdg | Discovery and screening: finding candidate edges and explaining the methodology used to challenge them. | That a screen, ranking or attractive chart is a tradable verdict. |
| BibaMoney | The public trust layer: reports, methodology, corrections, source boundaries, product offers and evidence passports. | Custody, guaranteed performance or a hidden affiliate endorsement. |
| GOGA | The operator interface: research workspaces, paper observation and, only after hard approval gates, carefully bounded actions. | That a visible button means live risk is currently authorized. |
In plain language: FinEdg discovers, AI OS tests, BibaMoney explains and publishes, and GOGA lets a human operate the approved part of the workflow. They share one proof contract.
What was actually built
The foundation is broader than a backtester. The completed foundation phase established configuration registries as the source of truth, loaders that expose those registries to applications, generation scripts for derived evidence bundles, and admin views that make system state inspectable. On top of that base, the trading programme assembled several practical rails:
- Research engines: causal signal construction, next-bar and resting-order execution, fee and slippage scenarios, walk-forward folds, held-out periods, passive and mechanism-matched baselines, and family-wise statistical correction.
- Evidence records: named campaign specifications, manifests, dataset and result hashes, explicit verdicts, failure reasons and dated reports.
- Product surfaces: a research optimizer, a public methodology, the research funnel, Lab case studies, an honest storefront and a sample strategy audit.
- Governance: paper-before-live posture, human approval for publication, no custody on public pages, preserved negative results and a machine-readable STOP decision for the live lanes.
Those are real capabilities. But they are not the same as an industrial evidence platform. A unified strategy ontology, end-to-end lineage, safe isolation for untrusted uploads, complete broker/entity passports, customer billing ledgers and repeatable delivery economics are still work to finish. The project is beyond a prototype collection, but it should not pretend to be a completed institution.
The campaigns that changed the thesis
The project tested unlike ideas under different contracts, so their raw returns do not belong in one performance leaderboard. What can be compared is the disposition of the claim: did it survive the proof gate, and what action did that authorize?
| Campaign | Load-bearing result | Project lesson |
|---|---|---|
| FVB | 47 / 47 stored composite runs are NOT READY; the DXY layer was present in 0 / 47. The latest gated BTC OOS arm remained −14.97% on 17 trades. | Verifying a component or improving a loss does not prove the whole architecture. |
| MoonBot | Three of four economic replicas were net-negative; one +$0.54 result remained paper-only. In a frozen eight-trader cohort, 2 / 8 held and 6 / 8 failed. | Vendor software, a coarse replica and copy-trader telemetry are three different claims. |
| PHOENIX/ZigZag | A look-ahead repair collapsed 119 candidates to 1. In the final 57-hypothesis campaign, 0 / 57 survived correction at either 5 or 10 basis points per side. | An optimizer is useful when it attacks the result instead of decorating it. |
| MRS | The optimized search produced 17 candidates where 16.95 were expected by chance, despite a high median win rate. | Win rate and selected winners are not evidence of an edge without a matched null. |
| ZZBoba | Attractive historical performance depended on choosing the correct regime with hindsight; no sealed regime classifier was demonstrated. | A strategy that needs future regime knowledge is still an untested classifier. |
| NASAlgo | 10 / 60 assets were positive in both development and OOS; average net performance across the fixed-rule universe was −4.4%. | A few attractive assets inside negative breadth form a research queue, not a family claim. |
| HiDeep | Across 487 symbols, 9 candidates appeared against roughly 5.55 expected by chance; the family verdict was NO EDGE. | Scaling a campaign also scales the number of accidental winners. |
| ai4trade | A forward set of 33,393 external signals produced a negative 24-hour net result and negative alpha versus hold. | Large forward evidence can close a copying thesis more decisively than another backtest. |
The detailed numbers, scope limits and broker boundary are documented in the cross-campaign strategy review. The broader experiment scoreboard explains how individual negative results improved the test bench.
The bugs were not side notes; they were the turning point
Several of the most valuable outputs were defects caught before money moved. A trailing simulator used the current bar's high before checking that bar's low. A highly active portfolio changed from an attractive positive result under a light cost assumption to a severe loss under a plausible spot-cost scenario. A disabled indicator silently passed through one experiment. A direction benchmark was wrong. A strategy with a striking win rate selected almost exactly as many winners as chance predicted.
Each repair made the headline worse and the system better. That is the key cultural shift. In a bot-marketing project, a failed result is an inconvenience to hide or retune. In an evidence project, finding the failure mode is part of the deliverable.
What the project became
The clearest name for the product is a Strategy Evidence Platform. Its public research branch can also operate as a Trading Claims Observatory. The observatory does not publish one opaque star score and does not merge strategy failure with allegations about a person, company, broker or exchange.
- Artifact and provenance: what exact code, rule, document or commercial claim is being reviewed; who may provide and use it; which version was frozen.
- Executable specification: signals, timing, fills, costs, position sizing, exits, instruments, period and rejection criteria written before selection.
- Reproduction: deterministic simulation, sanity tests, benchmarks, nulls, OOS, sensitivity and correction for the number of attempts.
- Verdict: PROVEN, CANDIDATE, NOT REPRODUCED, INSUFFICIENT EVIDENCE, NOT READY or another scoped label—with reasons and expiry.
- Strategy Passport: hashes, engine and data versions, limitations, risk, execution assumptions, reviewer approval and a repeatable report.
- Publication and correction: public evidence, clear conflicts, right of reply, version history and human approval before a consequential claim appears.
- Action gate: paper observation first. Live action remains a separate permissioned decision and is currently closed for the reviewed crypto families.
An Entity passport belongs beside—not inside—the strategy verdict. It can record the legal entity, domains, jurisdiction, regulator sources, fee tier, custody model, permissions, incident history, evidence freshness and affiliate conflict. A safe venue cannot rescue a losing strategy; a failed strategy test cannot prove a venue unsafe.
What is live, what is next, and what is blocked
| State | Capability | Honest boundary |
|---|---|---|
| Live now | Public methodology, research funnel, optimizer interface, evidence articles, sample audit, read-only and paper-oriented product surfaces. | These show the method and current evidence; they do not authorize capital. |
| In progress | A verdict hub and separate broker, exchange, bot and entity passports. | Every row needs exact scope, source, date, freshness and disclosure. |
| Validate next | Concierge strategy audits using the existing report and payment rails. | Demand, delivery time, refunds, compute, manual cost and margin must be measured before automation. |
| Gated later | Quarantined uploads, isolated execution, a customer billing ledger, authorized source catalog and rating integration. | No untrusted customer code runs on the main host; missing rights or unsupported formats fail closed. |
| Blocked | Live orders, automatic signal copying, deposits for these campaigns, blanket source scraping, magical rankings and a strategy marketplace. | These do not reopen because another optimized cell looks attractive. |
A business model that does not require a profit promise
The first commercial hypothesis is deliberately service-led: a free precheck, a small fixed-price AI-assisted review, a verified audit from $149 and bespoke sweep or porting work from $500. These are pilot price hypotheses, not proof of established revenue, guaranteed turnaround or universal availability. The fixed quote must be agreed before sensitive material is submitted, and an audit can end with a negative or inconclusive verdict.
What the customer buys is avoided uncertainty: a reproducible description of what was tested, what broke, what survived, and what decision is justified. The business should not depend on the reviewed strategy winning. It should not attach an undisclosed referral to a negative report. The reviewer is paid for the protocol and the report, not for manufacturing a positive outcome.
The product gate is operational rather than rhetorical: at least 10 qualified submissions, at least 3 paid audits or an explicit human override, median manual delivery of no more than three hours, zero source-permission failures, and measured conversion, refunds, compute cost, manual cost and margin. Until that evidence exists, building a large automated intake and marketplace would be premature.
Why the broker and exchange question remains separate
The project has already seen exchange feeds, monitoring dashboards, fee tiers and affiliate links mixed into strategy narratives. The resulting temptation is to issue one total verdict—good broker, bad bot, safe exchange, verified trader. That collapses different evidence units.
A broker or exchange passport should answer its own dated questions: which legal entity and domain, which jurisdiction and regulator source, which fee and account tier, which custody and withdrawal permissions, which execution evidence, which incidents, and whether the publisher benefits from a referral. Strategy portability then becomes a separate measured question: does the same locked rule survive the actual costs, liquidity and order semantics of that venue?
Until those passports are current, the project should neither endorse nor accuse a venue based on a strategy backtest. Referral links may support the publication layer only when visibly disclosed and kept outside the evidence verdict.
What still has to be proved
- One canonical ontology: strategy, version, author, artifact, dataset, test, claim, entity, verdict and permitted action must resolve to stable objects across services.
- End-to-end lineage: every public number must trace to a source capture, dataset, code version, engine, cost contract, result hash and reviewer decision.
- Safe customer intake: rights checks, quarantine, isolation, resource limits and deterministic outputs must exist before arbitrary uploads are accepted.
- Entity evidence: broker and vendor facts need jurisdiction-specific sources, observed dates, freshness rules, conflict disclosure and a correction path.
- Forward proof: promising candidates need sealed paper ledgers long enough to test the claim they actually make.
- Product evidence: real qualified submissions and paid delivery must demonstrate that users value the report enough to support the work.
Success is not the number of strategies ingested or charts published. It is the number of current, reproducible Strategy Passports with provenance, hashes, hard gates, OOS evidence, uncertainty, failure modes and a report that can be regenerated without relying on an AI's opinion.
The culmination
AI OS did not culminate in a bot. It culminated in a decision system for resisting one.
The research campaigns supplied the painful raw material: non-causal simulations, cost-sensitive curves, over-selected candidates, impressive win rates without alpha, regimes identified after the fact, copy statistics without a capital ledger, and external signals that failed forward. The platform turned those failures into tests, gates, reports and public corrections.
That is a smaller claim than “we solved trading,” but a much more valuable foundation. There are endless places to discover a strategy and very few places designed to tell its author, buyer or operator—reproducibly—why it should not yet touch capital.
Editorial policy and right of reply
This retrospective is a project and evidence statement, not investment advice, a scientific paper, a custody offer or an invitation to trade. It contains no affiliate link to a reviewed strategy, bot, broker or exchange. Strategy authors, operators and venues may submit dated primary evidence, corrections or a response. Material changes should update the evidence object, report and public verdict rather than silently replace the past.
Start with the sample strategy audit, inspect the live research funnel, read the FinEdg methodology, or compare the architecture in the operating-system design note.