AI to Bank: The October Ledger
The state of AI trading in October 2026: crypto agents, three months of research, and the monitoring discipline that connects a prediction to an order, a fill, and an accountable outcome.

In The House Learns an Orbit, we drew the future season by season: a system supporting its compute, its roof, its soil, and eventually the village table. October is where we open the first door and inspect its hinges. The first interface is financial trading, initially in crypto. The question is how an evolving intelligence becomes a dependable participant in that environment.
The research of the past three months suggests a sharper task than choosing the most persuasive trading agent. We need to observe an entire chain: what information existed, what the model decided, what the controller allowed, what reached the exchange, what filled, and what remained after costs. Each link needs its own evidence.
The first form of autonomy is an account of what happened.
This is the opening issue of Pandemonium’s Algo trading pillar and its AI to bank research ledger. The review covers July 11–October 11, 2026, with earlier work for context. It is a curated primary-source review, not an exhaustive survey. The monitoring contract below is our proposed implementation direction. No private trading account or local runtime was inspected for this issue; the numerical example is constructed.
More capable systems.
More exact questions.
Publication dates describe when evidence entered this review. The underlying experiments often occurred earlier. Paper results remain attributed to their authors and settings.
Can Agentic Trading Systems Pay for Their Own Intelligence?
TradeLens reconstructs trading trajectories from records, runtime traces, and configuration to study whether the extra decisions made by an agent repay their costs. It examines variation across models, capital, frequency, and architecture.
- What we take forward
- Account for inference and tool costs beside incremental trading value. A profitable portfolio alone does not establish that the intelligence added value.
- Evidence boundary
- Evidence is specific to the evaluated systems and runs; the result is not a universal break-even rule.
Fin-Analyst at FinMMEval 2026
A hybrid system used eight LLM specialists for Tesla and a rule-based vote for Bitcoin. Its reported asset ranking changed between interim and final evaluation; the BTC strategy ended flat while the underlying benchmark fell.
- What we take forward
- An equity result cannot stand in for a crypto result. Publish the asset, architecture, dates, and benchmark together.
- Evidence boundary
- A short live window and different pipelines across assets prevent a clean claim that one model or architecture is generally superior.
FORESIGHT-9
Across nine counterfactual market worlds and 36 runs, a fixed equal-weight policy beat 31 runs. In one strong-return run, the factor library failed while holdings converged to a fallback, although decision records still described an active ensemble.
- What we take forward
- Monitor whether the named strategy is the one actually operating. Returns can conceal a broken process.
- Evidence boundary
- These are generated stress trajectories branching from a July information boundary, not nine observed live markets.
What LLM Trading Agents Actually Do in Production
A report on two related crypto fleets describes real ETH vault activity and perpetual-futures agents. The authors found no directional edge in either fleet; the DXAP fleet was unprofitable, and many positions surrendered favorable excursions before closing.
- What we take forward
- Separate signal direction, sizing, exits, execution, and realized capture. A move seen on a chart is not a gain banked.
- Evidence boundary
- An observational report from one design lineage is valuable field evidence, not a controlled verdict on all AI traders.
Can AI Make Money in Crypto?
The study compares ML, RL, LLM, and agent approaches through historical testing, prospective exchange paper trading, and a short real-money stage. Performance and family rankings change between stages.
- What we take forward
- Record temporal realism and execution realism separately. Preserve the selected configuration and its selection date when advancing a strategy.
- Evidence boundary
- Methods use differing action formulations, and only selected contenders advance. The short live stage cannot establish durable superiority.
Building crypto portfolios with agentic AI
A multi-agent allocation study compares static and rolling optimization using daily data for ten cryptocurrencies from 2020–2025 and reports better risk-adjusted results for its dynamic strategy.
- What we take forward
- Agent orchestration is also useful for research and portfolio construction, beyond direct buy/sell prompting.
- Evidence boundary
- Historical out-of-sample results are not live trading. The paper contains a cost-adjusted metric description but explicitly lists ignored spreads and fees as a limitation; treat execution-cost coverage as unresolved.
From learning a policy to operating a desk.
Earlier frameworks supply different pieces of the machinery. Reinforcement learning made the environment and reward explicit; memory and tools expanded what agents could use; specialist teams made analysis modular. Recent evaluation work asks whether those parts produce stable behavior and incremental value.
FinRL ↗
Reusable reinforcement-learning environments, trading constraints, and a pipeline from data through training and evaluation. It provides infrastructure for experiments, not a transferable return guarantee.
FinMem ↗
Layered memory, profiling, and reflection bring past observations into an LLM decision process. The engineering question is whether memory remains timely and useful as the regime changes.
FinAgent ↗
Multimodal inputs and tools combine news, prices, and charts. Richer inputs make timestamp discipline and leakage controls more consequential.
TradingAgents ↗
Specialist analysts, opposing researchers, traders, and risk roles model a collaborative trading desk. More roles also create more interfaces whose contribution needs measurement.
AlphaForgeBench ↗
This benchmark treats LLMs as strategy researchers producing executable factors, with deterministic evaluation. It responds to unstable sequential action choices by separating research from execution.
Our October inference is to give each component a measurable job: an LLM can extract an event, propose a factor, or explain an anomaly; a predictive model can estimate a defined outcome; a controller can apply exposure and execution rules. Competing architectures should be judged under the same contract, including their cost and latency. Calling the future lineage “super intelligence” describes our horizon of ambition. The published studies here evaluate specific present-day systems.
Follow one decision all the way through.
A healthy process and a completed model response are useful observations. Delivery requires an exact link to the intended source and a result within its declared deadline.
- 01
Source
Immutable source ID, canonical content hash, market timestamps, completion time, and data version.
Watch for: Stale, incomplete, duplicated, or future-dated inputs. - 02
Decision
Exact source join; model snapshot, prompt/schema versions, decision time, action, forecast target, and any abstention reason.
Watch for: Missing join, invalid schema, deadline missed, or changed forecast meaning. - 03
Controller
Deterministic eligibility and risk checks, strategy version, target exposure, accepted/rejected result, and reason.
Watch for: Decision differs from permitted action; limits or configuration drift. - 04
Exchange
Client order ID, venue order ID, submission and acknowledgement times, state transitions, and query reconciliation.
Watch for: Unknown order state, rejected order, duplicate retry, or stream gap. - 05
Fill
Trade IDs, quantity, price, commission and commission asset, partial-fill status, and reconciled balances.
Watch for: Order acknowledgement mistaken for a fill; missing fees or wallet mismatch. - 06
Outcome
Maturity cutoff, label version, realized net P&L, open mark-to-market, benchmark, and operating costs.
Watch for: Premature labels, overlapping cohorts, unresolved observations counted as negatives.
Count delivery from completed source artifacts, not only from the responses that survived. Report the fraction joined to a valid decision within the chosen service window, plus median, p95, and p99 latency and the late/missing split. Freeze the source completion boundary and observation cutoff so unfinished work is not scored prematurely. Superseded sources need an explicit policy and their own count.
At the exchange, acknowledgement and execution are different states. Binance’s official user-stream documentation describes account updates and order execution reports. The proposed monitor joins these to order queries, trade IDs, commissions, and balances. It records partial fills, cancellations, rejections, and unknown states. A later HOLD decision cannot establish what an earlier order did.

Define “good” before the market answers.
A confusion matrix is meaningful only after its population, prediction, label, and clock are fixed. A direction call, a profitable round trip, a target touched, and an order filled answer different questions. We will give each question its own ledger.
Constructed example / BTC–USDT spot / one venue
- Unit and cohort
- One scheduled daily opportunity; one frozen model, prompt, controller, data source, venue, and label version. Fixed $100 shadow size; no leverage or shorting.
- Prediction and deadline
- At source cutoff t, predict the probability that this fixed long action clears the label. Positive if p ≥ 0.65, negative below it; explicit abstentions stay separate. A valid response must arrive by t + 60 seconds.
- Shadow action and outcome
- Use archived executable ask depth for entry at t + 60 seconds and bid depth for exit exactly 24 hours later. Quantity respects the $100 budget including entry fees. Net return includes both sides’ fees; depth prices incorporate spread and size-dependent slippage.
- Positive label
- Y = 1 only when the shadow round trip earns more than +10 basis points net. Y = 0 otherwise, including break-even and gains up to the buffer. The illustrative fee assumption is 10 bps each side, not a current venue quote.
- Maturity and unknowns
- Score only after exit data and the declared ingestion grace are complete. Missing executable depth is censored. A pending observation, abstention, failed delivery, or unknown outcome is never silently turned into TN.
The shadow policy gives the same hypothetical action an outcome even when the model said no. That is how FN and TN become observable in this example. Their labels describe missed or avoided benchmark opportunities, not actual trades. If only executed fills are available, publish the trade ledger and its selection bias; a complete opportunity matrix is unavailable.
| Model prediction | Y = 1 Cleared the net buffer | Y = 0 Did not clear it |
|---|---|---|
| Positive / p ≥ 0.65 | TP 180Called yes; opportunity cleared | FP 120Called yes; opportunity failed |
| Negative / p < 0.65 | FN 70Called no; opportunity cleared | TN 630Called no; opportunity failed |
Precision is 180 / (180 + 120); recall is 180 / (180 + 70); accuracy is (180 + 630) / 1,000. The positive prediction rate is 300 / 1,000, or 30%, within this scored population. None of those fractions establishes profitability. If the 180 successful selected cases average +20 bps and the 120 failures average −40 bps after trading costs, the mean selected return is −4 bps. Compute and other operating costs reduce it further.
Give the surrounding funnel equal prominence. In a separate constructed delivery example, 1,200 completed sources produce 1,140 timely, valid, exactly joined decisions: 95% delivery. Of those decisions, 40 abstain, 100 await outcome maturity, and 1,000 enter the matrix above. The remaining 60 sources are late or missing. The matrix is an assessment of those 1,000 cases; it cannot certify the other 200.
For probabilistic outputs, report calibration and Brier or log loss against the same event definition. A categorical confidence badge is not automatically a probability. Add uncertainty intervals and block-aware resampling because overlapping horizons and repeated market regimes reduce the effective sample size. The scikit-learn evaluation reference supplies the metric definitions; the trading contract supplies their meaning.
A clock for the work.
A ledger for the exceptions.
This cadence defines what future implementation and reporting should collect. It is a protocol published with this issue; it does not imply a connected account or an unattended trading monitor is running.
Every event
Validate lineage and freshness, record state changes, deduplicate exchange events, and enforce the declared deadline and exposure limits. An unresolved order enters reconciliation before any retry.
Daily / UTC close
Close a maturity-bounded ledger: delivery denominator and latency, orders and fills, balance reconciliation, pending/censored/abstained counts, costs, and operational incidents. Show the exact as-of time and missing evidence.
Weekly
Compare frozen cohorts against cash, buy-and-hold, and a simple deterministic strategy under the same capital and cost rules. Review calibration, precision/recall, coverage, tail losses, drawdown, turnover, and errors grouped by regime.
Monthly / research issue
Freeze a reproducible snapshot; document changes in models, prompts, thresholds, data, venues, and execution. Compare to the prior issue only where the contract matches. Log what improved, what failed, and the next falsifiable experiment.
Scoping is part of the measurement. Keep spot and perpetuals, asset pairs, venues, model snapshots, prompt revisions, controllers, horizons, leverage, capital, fee tiers, and backtest/paper/live stages explicit. When one changes, open a new cohort or run a matched comparison. Do not attribute a changed execution policy’s result to a model upgrade.
Economic monitoring has two views. The trade view uses actual fills, actual fee assets, funding or borrow where applicable, and reconciled inventory. The system view adds inference, market data, hosting, network, maintenance, and allocated hardware costs. Report realized P&L and open mark-to-market separately, neutralize deposits and withdrawals in return calculations, and compare to passive exposure. Capital losses cannot be counted as a paid household bill.
The operating contract also declares when to pause new exposure: stale inputs, unresolved order state, unreconciled balances, breached limits, or sustained drift. Existing exposure still needs its prescribed reconciliation and exit handling. An alert is evidence to investigate; it does not authorize a new strategy or larger risk budget.

Earn the next interface.
Begin with a reproducible historical harness and frozen source/decision contracts. Advance to prospective shadow or exchange paper trading with the same versions and costs recorded. Establish delivery, labels, reconciliation, and matched baselines before considering a separately bounded live trial. Capital increases require evidence across enough regimes and enough independent observations to support the decision.
Our first research experiments are concrete: compare an LLM event extractor plus deterministic controller to that controller alone; compare a learned return model to simple momentum and cash under matched turnover; measure whether a richer reasoning pass adds value after its latency and compute cost; and test whether an exit change captures more of an observed move without increasing tail loss. Each experiment gets a frozen hypothesis, an untouched evaluation period, and a record of all attempts, including the ones that fail.
That gives the seasonal atlas a way to meet time. Future issues can say which capability arrived, which interface became reliable, and which projected milestone remains open. The moon can wax in the illustrations while the ledger changes at its own pace. The bank-to-farm-to-table story begins with an ordinary accomplishment: knowing what our system did, what it cost, and what we learned.
Harmony begins when the account balances with the world.