Every honest description of a trading idea ends the same way: here is how it could fail. This chapter is that ending. No hedging language, no "but we're confident" — just the plain list of things that could go wrong, followed by a glossary so you never have to guess what a word meant.
By now you've seen the whole machine. A crowd on a prediction market tends to overpay for surprises — for the low-probability weather events that feel dramatic but rarely happen. We estimate that overpayment at roughly 1.3x: the crowd charges about 30% more than the fair price for those long-shot outcomes. We forecast the weather ourselves with a large ensemble of models, compare our number to the market's number, and bet only when the gap is big enough to be worth it. That's the edge. This chapter is about all the ways that story could be wrong, incomplete, or simply unlucky.
The mindset to carry through this chapter: a good trading idea and a good story about a trading idea look identical from the outside — both are confident, both cite numbers. The only thing that separates them is whether the world actually pays out the way the story predicts, over enough trials that luck washes out. We have not run that test yet with real money.
Where things actually stand
Let's start with the single most important fact, stated bluntly.
polyAether has never made a real bet with real money. Everything so far is paper trading — running the full system live, but recording imaginary trades instead of placing them. There is no track record. The edge is a hypothesis supported by historical analysis, not a proven fact confirmed by profit.
"Paper trading" means the software watches real markets in real time and decides what it would do — but no money changes hands. It's a flight simulator, not a flight. Simulators are genuinely useful; they catch bugs and let you measure behaviour safely. But a simulator has never crashed a real plane, and it has never landed one either. Keep that framing for everything below.
What has changed recently is worth stating precisely, because it's easy to over-read. The paper-trading loop now runs a full profit-and-loss lifecycle end to end: when a genuine edge clears every safety gate, the system opens a position; it holds that position across many decision cycles; when the real weather observation lands, the position settles against the actual reading (the same official METAR data a real market would use); and the resulting win or loss flows into the shown balance and equity. The imaginary account now behaves like a real one — trades open, live, close, and book their result. That is a meaningful engineering milestone. It is not a track record. A paper ledger that moves correctly is evidence the plumbing is sound, not evidence the edge is real — two entirely different claims this chapter is careful never to conflate.
A worked example makes the distinction concrete. Say the system opens a paper position on a "high temperature 88–90°F" bucket at 22 cents, believing the true probability is 35%. Days later the airport reports 89°F; the bucket resolves YES; the position pays out at 100 cents, and roughly +78 cents per share of realized profit lands in the paper equity. The arithmetic and the settlement source are real — but a single win tells you almost nothing. Flip the weather two degrees and the identical, equally-well-reasoned trade is a total loss. The ledger records both faithfully; only after hundreds of them does the average start to mean something.
The risks, plainly
The edge might not be real
The whole thesis rests on that 1.3x overpricing. We measured it in historical data — but the past is not a promise. Markets learn. If enough traders notice the same pattern, they stop overpaying, and the gap we're betting on quietly closes. It's also possible the pattern was partly an illusion in the data — a coincidence that looked like a rule. We won't know for sure until real money has been at stake for a long time.
It's worth being exact about what "the edge" even is, because it's narrower and more boring than the exciting version people imagine. It is not that we forecast the weather better than the professionals — we mostly use the same public models. It is not that we win some nanosecond speed race. The claimed edge is one specific, testable thing: our probabilities are better calibrated than the crowd's. The crowd, drawn to dramatic outcomes, tends to say "35%" when the honest number is closer to 27%; that 1.3x wedge is the entire game. If our confident "35%" also happens only 27% of the time, there is no edge at all — just extra typing. We check this with Brier and PIT scoring, which only earn trust as they accumulate over hundreds of resolved markets. Until then, "the edge is real" is a hypothesis wearing a lab coat.
The classic trap: a pattern that fits yesterday's data perfectly but has no predictive power tomorrow — like a "system" for picking lottery numbers that only works on last week's draw.
Thin liquidity
Liquidity is how much you can buy or sell without moving the price against yourself, and this is the risk where the honest reality is most surprising. The naive worry is "the markets are empty." That's not what we see. Almost every weather market has real size resting in it: tens to hundreds of thousands of dollars of standing asks (offers to sell). On paper that looks deep. The problem is the shape of that depth — it is overwhelmingly one-sided: mostly market-maker sell walls sitting above very thin bids. If we want to buy there's plenty to buy from; but for someone to willingly take the other side of a fair-priced bet, the crowd is wafer-thin. And where prices are genuinely alive — a bucket trading between roughly 10 and 92 cents — the market is usually already efficiently priced. The obvious mispricings we'd love to pounce on tend to be against 0.1-cent "dust" on far-tail buckets that will almost certainly resolve NO. Buying dust isn't an edge; it's a rounding error dressed up as one, and our minimum-price gate correctly refuses to touch it.
Run the model across a set of markets and it might flag ~18 apparent edges. After the gates strip out the dust-tail illusions and one-sided books, only about 2 tend to survive as genuinely tradeable — and on many days, none do. That is the system working, not failing. The tradeable window is narrow: often just a day before resolution, when two-sided interest appears, before the market collapses toward near-certainty. Expect few trades; frequency is not the goal.
Competition
We are not the only clever people looking at these markets. If a better-funded, faster competitor is doing something similar, they can take the good prices before we do, leaving us the scraps. Edges are shared until they're gone. This connects to the liquidity picture: the reason so many live-priced buckets are already efficient is that other sharp participants have already done the obvious work. Our hope isn't to out-muscle them but to be well-calibrated in the specific corner — crowd overpricing of long-shot weather — where the sharp money isn't fighting over pennies.
Model risk
Our forecast comes from weather models, and weather models are wrong sometimes. A physics assumption breaks, a data feed goes stale, our probability estimate is miscalibrated — meaning when we say "20% chance," it doesn't actually happen 20% of the time. If our number is off, our "edge" is just confident nonsense, and confident nonsense loses money faster than honest uncertainty.
The sneaky version of model risk isn't the model being wildly wrong — that's easy to catch. It's the model being subtly overconfident: our ensemble reports a tighter spread than reality warrants, saying "35%" when the honest number is "30%." Every forecast looks reasonable, nothing throws an alarm, and yet we're manufacturing a fake edge out of false precision. This is exactly what the PIT diagnostic is built to catch — it asks, over many past forecasts, whether reality landed inside our stated ranges as often as it should have. A lopsided PIT is the early-warning light for "our confidence is a lie," which is why calibration, not raw forecasting skill, is the thing we obsess over.
Venue risk
The venue is the platform where the market lives (for us, Polymarket). Venue risk is everything outside our control there: the site goes down, rules change, a market settles (pays out) based on a weather reading we didn't expect, funds get frozen, or the legal picture shifts. You can be completely right about the weather and still lose because the venue did something surprising. Chapter 7 covers why settlement is the sneakiest part of all this.
A worked example of how venue risk bites: we forecast a city's high perfectly, our bucket is dead right, and we still lose — because the market settled on a different station's reading than we assumed, or rounded 89.5°F up to 90 (the resolution source rounds, it does not truncate), tipping the outcome into the neighbouring bucket. Nothing about our forecast was wrong; the venue's definition of "the answer" simply wasn't the one in our head. That's why we curate which station settles each market and read the resolution rules literally.
Weather is genuinely uncertain
This one isn't a bug — it's the nature of the thing. The atmosphere is chaotic. Even a perfect model can only give probabilities, never certainties. We are betting on being right on average over many bets, not on any single forecast. Over a small number of trades, plain bad luck can look exactly like a broken strategy. Only volume tells the difference, and volume takes time.
Return to that 88–90°F trade. We bought at 22 cents believing the true probability was 35% — a genuine, well-reasoned edge. Play it out over a hundred similar setups: if our 35% is honest, we win roughly 35 and lose 65. Each win pays about +78 cents; each loss costs the 22 cents we paid. The math is positive on average — that's the whole point — but notice the texture: we lose almost twice as often as we win. A run of five, eight, even ten losses in a row is completely ordinary and tells you nothing about whether the edge is real. Anyone watching a short slice could reasonably conclude the strategy is broken. It might be. It might just be Tuesday. Only sample size can tell those apart, and sample size is slow to arrive — especially when, as the liquidity section explained, the disciplined move is often to place no trade at all. Patience here isn't a virtue we're advertising; it's a structural requirement.
The edge is plausible and carefully measured, but unproven. Even if it's real, thin liquidity, competition, model errors, and venue surprises can eat it. And weather itself guarantees that being right "on average" still means losing plenty of individual bets. This is a research project, not a sure thing.
The guardrails
We can't remove these risks, but we can refuse to let any one of them sink us. Chapter 8 is the full story; here's the short version. We bet only a fraction of what the math says is optimal (fractional Kelly), we cap how much rides on any single market, on the total book, and on any one day, we limit exposure to bets that would all lose together (a correlation cap), and there's a kill switch — a single control that halts everything instantly if the system misbehaves. None of this makes the edge real. It just makes sure that if we're wrong, we live to learn from it.
Two of those guardrails do quiet, load-bearing work against the risks above. The minimum-price gate turns the ugly liquidity picture into a feature: by refusing to trade against 0.1-cent tail dust, it converts most of the model's fake edges into a clean no — why a run yields two survivors or zero rather than eighteen bad fills. And calibration scoring (Brier and PIT accumulating over time) is the guardrail against model risk: a slow, honest audit of whether our probabilities have earned their confidence. Speed plays a supporting role, but not the heroic one people expect. The system reacts to order-book changes in milliseconds — not to win a nanosecond race, but so we're first to genuinely new information (for weather, that arrives in seconds-to-minutes, as fresh model runs and hourly observations land) and hard to pick off. Speed protects the edge; it is not the edge.
Every term, defined
If a word in this course ever felt like a bluff, here's the plain meaning.
- Ensemble — Running the weather model many times with slightly different starting conditions and getting a spread of answers instead of one. The spread is the forecast's uncertainty. polyAether uses a ~122-member ensemble built from three model families (GFS, ICON, ECMWF).
- Brier score — A grade for probability forecasts. Lower is better. It rewards being both confident and right, and punishes confident-but-wrong hard. It's how we check whether our forecasts are actually any good.
- Kelly — A formula for bet size that grows your money fastest over the long run given your edge. Full Kelly is aggressive and swingy, so we use a fraction of it to stay calmer and safer.
- Edge — The gap between our estimated probability and the market's price. If we think an event is 40% likely and the market prices it at 25%, the edge is that 15-point difference — the reason to bet at all.
- Calibration — Whether your probabilities match reality. If everything you call "30%" happens 30% of the time, you're well calibrated. Miscalibration turns a real edge into a fake one.
- Liquidity — How much you can trade without moving the price against yourself. High liquidity = deep market, easy to get in and out. Low liquidity = thin market, your own order shifts the price.
- Bucket — A range a market question carves the weather into, e.g. "high temperature 70–72°F." Each bucket is a separate yes/no bet.
- Station — The specific weather-observation site whose reading officially decides a market. polyAether tracks ~80 curated stations. Which station settles a market matters enormously (see Chapter 7).
- PIT — Probability Integral Transform. A diagnostic that checks whether our ensemble's uncertainty is honest — not too confident, not too timid. If the PIT looks lopsided, our spread is off and needs fixing.
- Settlement — How a market pays out: which real-world measurement, from which station, at which moment, decides who was right. For us the reading comes from official METAR observations, and it rounds (89.4°F → 89, 89.5°F → 90) rather than truncating — a small rule with big consequences at bucket edges.
- Ask / Bid — An ask is a standing offer to sell at a price; a bid is a standing offer to buy. A healthy market has both stacked deep. Weather markets tend to have deep asks (market-maker sell walls) sitting above thin bids — plenty to buy from, few willing to take the other side.
- Minimum-price gate — A rule that refuses to trade against near-zero "dust" (e.g. 0.1-cent offers on far-tail buckets). It's what turns most of the model's apparent edges into a disciplined no-trade rather than a bad fill.
- Realized P&L — Profit or loss that is locked in because a position has actually settled, as opposed to paper gains on a position still open. In our system this now flows all the way into the shown paper balance and equity — using imaginary money.
- Paper trading — Running the whole system live but with imaginary money. The full open-hold-settle-and-book-the-result lifecycle now works end to end; it's still imaginary money, and still where we are right now.
That's the honest picture, and that's the whole vocabulary. If you've read this far, you understand polyAether as well as we do — including, importantly, everything we don't yet know.