Project retrospective at the 30 July 2026 evidence cutoff. AI OS began with a familiar ambition: collect promising trading ideas, translate them into code, optimize them, and find a system worth operating. The research did produce engines, interfaces, datasets and millions of evaluations. It did not produce the shortcut the original ambition implied.

That failure changed the project in a useful way. The central question stopped being “Which bot should we run?” and became “What evidence would make any strategy safe enough to trust—and can we produce the same verdict again?”

This is not a claim that every component of AI OS is finished, that every tested idea is worthless, or that every broker, bot author or strategy vendor is unsafe. It is a narrower and more defensible conclusion: the current evidence does not support an executable endorsement, while the machinery built to reach that conclusion is itself becoming a product.

One project, four product layers

The names around the project describe different jobs. Treating them as four separate startups would create confusion and duplicate infrastructure. The coherent architecture is one evidence system with four visible layers:

Layer役職What it must never imply
Tvijo AIOSResearch and orchestration: ingestion, registries, datasets, simulations, optimizers, evidence objects, memory and policy gates.That an AI-generated answer is evidence or approval.
FinEdgDiscovery and screening: finding candidate edges and explaining the methodology used to challenge them.That a screen, ranking or attractive chart is a tradable verdict.
BibaMoneyThe public trust layer: reports, methodology, corrections, source boundaries, product offers and evidence passports.Custody, guaranteed performance or a hidden affiliate endorsement.
GOGAThe operator interface: research workspaces, paper observation and, only after hard approval gates, carefully bounded actions.That a visible button means live risk is currently authorized.

In plain language: FinEdg discovers, AI OS tests, BibaMoney explains and publishes, and GOGA lets a human operate the approved part of the workflow. They share one proof contract.

What was actually built

The foundation is broader than a backtester. The completed foundation phase established configuration registries as the source of truth, loaders that expose those registries to applications, generation scripts for derived evidence bundles, and admin views that make system state inspectable. On top of that base, the trading programme assembled several practical rails:

  • Research engines: causal signal construction, next-bar and resting-order execution, fee and slippage scenarios, walk-forward folds, held-out periods, passive and mechanism-matched baselines, and family-wise statistical correction.
  • Evidence records: named campaign specifications, manifests, dataset and result hashes, explicit verdicts, failure reasons and dated reports.
  • Product surfaces: a research optimizer, a public methodology, the research funnel, Lab case studies, an honest storefront and a サンプル戦略監査.
  • Governance: paper-before-live posture, human approval for publication, no custody on public pages, preserved negative results and a machine-readable STOP decision for the live lanes.

Those are real capabilities. But they are not the same as an industrial evidence platform. A unified strategy ontology, end-to-end lineage, safe isolation for untrusted uploads, complete broker/entity passports, customer billing ledgers and repeatable delivery economics are still work to finish. The project is beyond a prototype collection, but it should not pretend to be a completed institution.

The campaigns that changed the thesis

The project tested unlike ideas under different contracts, so their raw returns do not belong in one performance leaderboard. What can be compared is the disposition of the claim: did it survive the proof gate, and what action did that authorize?

CampaignLoad-bearing resultProject lesson
FVB47 / 47 stored composite runs are NOT READY; the DXY layer was present in 0 / 47. The latest gated BTC OOS arm remained −14.97% on 17 trades.Verifying a component or improving a loss does not prove the whole architecture.
MoonBotThree of four economic replicas were net-negative; one +$0.54 result remained paper-only. In a frozen eight-trader cohort, 2 / 8 held and 6 / 8 failed.Vendor software, a coarse replica and copy-trader telemetry are three different claims.
PHOENIX/ZigZagA look-ahead repair collapsed 119 candidates to 1. In the final 57-hypothesis campaign, 0 / 57 survived correction at either 5 or 10 basis points per side.An optimizer is useful when it attacks the result instead of decorating it.
MRSThe optimized search produced 17 candidates where 16.95 were expected by chance, despite a high median win rate.Win rate and selected winners are not evidence of an edge without a matched null.
ZZBobaAttractive historical performance depended on choosing the correct regime with hindsight; no sealed regime classifier was demonstrated.A strategy that needs future regime knowledge is still an untested classifier.
NASAlgo10 / 60 assets were positive in both development and OOS; average net performance across the fixed-rule universe was −4.4%.A few attractive assets inside negative breadth form a research queue, not a family claim.
HiDeepAcross 487 symbols, 9 candidates appeared against roughly 5.55 expected by chance; the family verdict was NO EDGE.Scaling a campaign also scales the number of accidental winners.
ai4tradeA forward set of 33,393 external signals produced a negative 24-hour net result and negative alpha versus hold.Large forward evidence can close a copying thesis more decisively than another backtest.

The detailed numbers, scope limits and broker boundary are documented in the cross-campaign strategy review. The broader experiment scoreboard explains how individual negative results improved the test bench.

The bugs were not side notes; they were the turning point

Several of the most valuable outputs were defects caught before money moved. A trailing simulator used the current bar's high before checking that bar's low. A highly active portfolio changed from an attractive positive result under a light cost assumption to a severe loss under a plausible spot-cost scenario. A disabled indicator silently passed through one experiment. A direction benchmark was wrong. A strategy with a striking win rate selected almost exactly as many winners as chance predicted.

Each repair made the headline worse and the system better. That is the key cultural shift. In a bot-marketing project, a failed result is an inconvenience to hide or retune. In an evidence project, finding the failure mode is part of the deliverable.

What the project became

The clearest name for the product is a Strategy Evidence Platform. Its public research branch can also operate as a Trading Claims Observatory. The observatory does not publish one opaque star score and does not merge strategy failure with allegations about a person, company, broker or exchange.

  1. Artifact and provenance: what exact code, rule, document or commercial claim is being reviewed; who may provide and use it; which version was frozen.
  2. Executable specification: signals, timing, fills, costs, position sizing, exits, instruments, period and rejection criteria written before selection.
  3. Reproduction: deterministic simulation, sanity tests, benchmarks, nulls, OOS, sensitivity and correction for the number of attempts.
  4. 判定: PROVEN, CANDIDATE, NOT REPRODUCED, INSUFFICIENT EVIDENCE, NOT READY or another scoped label—with reasons and expiry.
  5. Strategy Passport: ハッシュ、エンジンとデータのバージョン、制限、リスク、実行の前提、レビュー担当者の承認、再現可能なレポート。
  6. 公開と訂正: 公開された証拠、明確な利害の衝突、反論の権利、バージョン履歴、および結果に影響を与える主張が現れる前の人間による承認。
  7. アクションゲート: まず紙面上の観察。ライブアクションは別途許可された決定であり、現在はレビュー対象の暗号通貨ファミリーに対して閉じられています。

エンティティパスポート は、戦略の判定の隣に属し、その中には属しません。それは、法的エンティティ、ドメイン、管轄区域、規制当局の情報源、手数料レベル、カストディモデル、権限、インシデント履歴、証拠の鮮度、および関連会社の衝突を記録できます。安全な会場は負け戦略を救うことはできません。失敗した戦略テストは会場が安全でないことを証明できません。 ライブで何が、次に何が、何がブロックされているか

正直な境界

状態能力現在ライブ
公開方法論、リサーチファネル、オプティマイザーインターフェース、証拠記事、サンプル監査、読み取り専用およびペーパー指向の製品サーフェス。これらは方法と現在の証拠を示します。それらは資本を許可しません。進行中
判定ハブと、別のブローカー、取引所、ボット、エンティティパスポート。すべての行には、正確な範囲、情報源、日付、鮮度、および開示が必要です。次に検証
既存のレポートと支払いレールを使用したコンシェルジュ戦略監査。需要、納期、返金、計算、手動コスト、およびマージンを自動化前に測定する必要があります。後でゲート
隔離されたアップロード、分離された実行、顧客請求台帳、承認されたソースカタログ、および評価統合。信頼されていない顧客コードはメインホストで実行されません。権限がないか、サポートされていない形式はフェイルクローズします。ライブ注文、自動シグナルコピー、これらのキャンペーンの預金、包括的なソーススクレイピング、魔法のようなランキング、および戦略マーケットプレイス。
ブロック中これらは、別の最適化されたセルが魅力的に見えるからといって再開されません。利益の約束を必要としないビジネスモデル

最初の商業仮説は意図的にサービス主導です:無料の事前チェック、小さな固定価格のAI支援レビュー、149ドルからの検証済み監査、および500ドルからの特注スイープまたは移植作業。これらは

パイロット価格仮説 であり、確立された収益、保証された納期、または普遍的な可用性の証明ではありません。固定見積もりは、機密資料が提出される前に合意される必要があり、監査は否定的または結論が出ない判定で終了する可能性があります。顧客が購入するのは、回避された不確実性です:何がテストされ、何が壊れ、何が生き残り、どの決定が正当化されるかの再現可能な説明。ビジネスはレビューされた戦略が勝つことに依存すべきではありません。否定的なレポートに未開示の紹介を添付すべきではありません。レビュー担当者はプロトコルとレポートに対して支払われ、肯定的な結果を製造するためではありません。

製品ゲートは修辞的ではなく運用上です:少なくとも10件の適格な提出、少なくとも3件の有料監査または明示的な人間によるオーバーライド、手動納品の中央値が3時間以下、ソース権限の失敗ゼロ、および測定されたコンバージョン、返金、計算コスト、手動コスト、マージン。その証拠が存在するまで、大規模な自動インテークとマーケットプレイスの構築は時期尚早です。

ブローカーと取引所の質問が別のままである理由

プロジェクトはすでに、取引所フィード、監視ダッシュボード、手数料レベル、およびアフィリエイトリンクが戦略の物語に混在しているのを見てきました。結果として生じる誘惑は、1つの総合的な判定を発行することです—良いブローカー、悪いボット、安全な取引所、検証済みトレーダー。それは異なる証拠単位を崩壊させます。

ブローカーまたは取引所パスポートは、独自の日付の質問に答える必要があります:どの法的エンティティとドメイン、どの管轄区域と規制当局の情報源、どの手数料とアカウントレベル、どのカストディと引き出し権限、どの実行証拠、どのインシデント、および発行者が紹介から利益を得るかどうか。戦略の移植性は、別の測定された質問になります:同じロックされたルールが、その会場の実際のコスト、流動性、および注文セマンティクスを生き残るかどうか。

これらのパスポートが最新になるまで、プロジェクトは戦略のバックテストに基づいて会場を承認または非難すべきではありません。紹介リンクは、公開レイヤーをサポートする場合にのみ、目に見えて開示され、証拠の判定の外に保たれる場合に限ります。

まだ証明されなければならないこと

1つの正規のオントロジー:

  • 。重要な変更は、過去を静かに置き換えるのではなく、証拠オブジェクト、レポート、および公開判定を更新する必要があります。 strategy、version、author、artifact、dataset、test、claim、entity、verdict、permitted actionは、サービス間で安定したオブジェクトに解決されなければなりません。
  • エンドツーエンドの系統性: すべての公開数値は、ソースキャプチャ、データセット、コードバージョン、エンジン、コスト契約、結果ハッシュ、レビュー担当者の決定に遡れる必要があります。
  • 安全な顧客受け入れ: 任意のアップロードを受け入れる前に、権利チェック、隔離、分離、リソース制限、決定論的な出力が存在しなければなりません。
  • エンティティの証拠: ブローカーとベンダーの事実には、管轄固有のソース、観測日、鮮度ルール、利益相反の開示、修正パスが必要です。
  • 前方証明: 有望な候補は、実際に行う主張をテストするのに十分な期間、封印された紙の台帳が必要です。
  • 製品の証拠: 実際の有資格の提出物と有料配信は、ユーザーがレポートを評価して作業を支援するのに十分であることを示さなければなりません。

成功は、取り込まれた戦略の数や公開されたチャートの数ではありません。それは、 現在の再現可能なStrategy Passports の数であり、来歴、ハッシュ、ハードゲート、OOS証拠、不確実性、障害モード、およびAIの意見に依存せずに再生成できるレポートを備えています。

集大成

AI OS ボットに至ることはなかった。それは、ボットに抵抗するための意思決定システムに至ったのだ。

研究キャンペーンは、苦痛を伴う生の材料を供給した:非因果的シミュレーション、コスト感応曲線、過剰選択された候補、アルファのない印象的な勝率、事後的に特定されたレジーム、資本元帳のないコピー統計、そして前方で失敗した外部シグナル。プラットフォームはそれらの失敗をテスト、ゲート、レポート、公的訂正に変えた。

それは「トレーディングを解決した」という主張より小さいが、はるかに価値のある基盤である。戦略を発見する場所は無限にあるが、その著者、購入者、運用者に、再現可能な形で、なぜまだ資本に触れるべきでないかを伝えるために設計された場所はほとんどない。

編集方針と反論権

この回顧録はプロジェクトおよび証拠の表明であり、投資アドバイス、科学論文、保管オファー、または取引への招待ではない。レビューされた戦略、ボット、ブローカー、取引所へのアフィリエイトリンクは含まれていない。戦略の著者、運用者、会場は 日付入りの一次証拠、訂正、または回答を提出できるから始めて

、ライブ サンプル戦略監査を検査するか、 research funnelを読み、 FinEdg手法のアーキテクチャを比較してください オペレーティングシステム設計ノート.