polyAether
Textbook · Chapter 6
Chapter 6

From a forecast to a bet

~11 min read

By now you know the pieces. A prediction market lets people buy shares that pay $1 if something happens (Chapter 2). Weather is forecastable but never certain (Chapter 3). A well-calibrated probability plus a price gap gives us an edge (Chapter 4). And the crowd systematically overpays when the outcome feels uncertain (Chapter 5). This chapter connects those pieces into a single, honest pipeline — the exact chain of steps that turns a weather forecast into a decision to buy, skip, or pass.

Think of it as a factory line. Raw material goes in one end — the atmosphere, and a market quoting a price. A finished decision comes out the other — buy this many shares, or nothing at all. In between are five stations, and a share only becomes a bet if it survives every one of them. The important thing to notice up front: most raw material never makes it to the end of the line, and that is by design. On a typical scan the machine looks at a wide slate of live markets and buys in almost none of them. A rejection at any station is not a missed opportunity — it is the pipeline doing exactly its job. Let's walk the line.

We'll follow one concrete market the whole way down so the abstractions stay attached to real numbers: "Chicago high above 90°F tomorrow." Watch it enter at the top as raw atmosphere, and watch what has to be true at each station for it to survive as an actual position.

01 — Forecast

Start with a probability, not a guess

The first station produces a number: the chance that some specific weather event happens. Not "it'll probably be hot" — an actual percentage, like "there is a 71% chance the high in Chicago tomorrow is above 90°F."

Where does that number come from? We don't run one forecast; we run a crowd of them. An ensemble is a large batch of forecasts, each started from a slightly different guess about the current state of the atmosphere, because we never measure the present perfectly. polyAether blends roughly a 122-member ensemble drawn from three of the world's major weather models — the American GFS, the German ICON, and the European ECMWF. (A "model" here is just a giant physics simulation of the atmosphere; each institution runs its own.)

Why a crowd? Because the spread of the ensemble is itself the forecast of uncertainty. If 87 of 122 members land above 90°F, that's a 71% probability — and the fact that 35 disagreed tells us how firm the call is. One forecast gives you a guess. A hundred forecasts give you a probability and an honest measure of your own confidence.

It helps to picture the raw output. Each of the 122 members hands back a single number: its predicted high for Chicago tomorrow. Line those 122 numbers up and you get a little histogram — a hill of predictions. Maybe it peaks around 91°F, with a tail reaching down to 87° and up to 94°. To turn that hill into the probability of "above 90°," we simply count: how many members landed on the far side of the 90° line? Eighty-seven did; thirty-five did not. Eighty-seven divided by one hundred twenty-two is 0.713 — call it 71%. That is the entire trick. We are not asking any single model to be a genius; we are asking a whole committee to vote, and reading the vote share as a probability.

The shape of that hill matters as much as its peak. A tall, narrow hill — nearly all 122 members clustered at 91° or 92° — means the models agree, and a call of "above 90°" is firm. A low, wide hill straddling the 90° line means the members genuinely disagree, and even a probability of 71% is a shaky 71% that later stations will treat with suspicion. The width of the hill is the confidence meter, and we carry it downstream rather than throwing it away.

Key idea

A single forecast is a guess. A large ensemble — many forecasts from many models — turns that guess into a probability, and the disagreement among members is a built-in confidence meter. We read the probability as a vote share and carry the spread downstream.

02 — Calibrate

Pin the forecast to the exact station

Here's a subtlety that quietly sinks amateurs. A weather model doesn't report the temperature at the exact rooftop thermometer the market cares about. It reports the temperature for a grid box — a chunk of sky maybe several miles across, averaged and smoothed. But the market settles on one specific, official instrument: a particular airport weather station.

The model's grid box and that one thermometer are not the same thing. The station might sit in a valley that runs a couple of degrees cooler, or on a sun-baked tarmac that runs warmer. If you bet the raw model number, you're betting on the wrong place.

Calibration fixes this. We take the model's history and the station's history and learn the persistent offset between them: "at this station, the model tends to read 1.4°F too warm, and it's a little overconfident near the threshold." Then we correct the live forecast the same way. polyAether does this for roughly 80 curated stations — a deliberately small, well-understood set, not the whole map — because a probability you can trust at one station beats a vague one everywhere.

Follow our Chicago example through this station. The raw hill peaked near 91°, but suppose history says the models run about 1.4°F warm at this particular airport. We slide the whole hill left by 1.4°. Now some members that had squeaked over the 90° line fall back under it. The count that were above 90° drops — say from 87 to 78 — and our probability comes down from 71% to something nearer 64%. That correction is not pessimism; it is accuracy. The market pays out on the airport thermometer, not on the model's smeared grid box, so the only probability worth carrying forward is the one measured at the airport thermometer.

There is a second, subtler correction hiding in the same step. Ensembles are often overconfident — their hills are a touch too narrow, so they claim more certainty than the atmosphere delivers. Calibration widens the hill back out to match how the station has actually behaved over hundreds of past days. A well-calibrated 64% means that across all the days we said "64%," the event really happened about 64 times in a hundred. Chapter 4 is where we prove this with Brier and PIT scoring; here it's enough to know that the number leaving this station is one whose track record we can check, not a vibe.

Key idea

Markets settle on one exact thermometer, not a smeared model grid box. Calibration learns the persistent difference between the two — the temperature offset and the ensemble's overconfidence — so our probability describes the place the bet actually pays on, and means what it says.

03 — Compare

Our probability versus the market price

Now we have a trustworthy number — the calibrated probability leaving station 2. The market is quoting a price for the same event. Remember from Chapter 2 that in a prediction market the price is a probability: a share that pays $1 if the event happens, trading at 58¢, means the crowd is pricing the event at 58%. Price and probability are the same currency here, which is what lets us compare them at all.

So we lay the two side by side. Take a clean case: our calibrated model says 71%, the crowd is charging 58%. That gap — 13 percentage points — is our raw edge: the amount by which we believe the market is mispriced. If our number and the market's number match, there's no edge and no reason to trade. The whole business lives in the gap.

But direction matters, not just size. A gap only helps if it points the buyable way. Because our number (71%) is above the price (58%), the "Yes" share looks cheap — we'd buy Yes at 58¢, expecting it to be worth 71¢ on average. Had our number come out below the price — say we thought 40% and the crowd charged 58% — the very same 13-point-ish gap would tell us the "No" side is the cheap one instead. The edge is a signed quantity: its size says how much, and its sign says which side of the market to stand on.

0% 50% 100% 71% Our model 58% Market price edge = 13 pts
Our calibrated probability What the crowd is charging
We say 71%, the market charges 58%. The 13-point gap is the edge — the reason to look closer. It is not yet a reason to buy.
04 — Clear the costs

An edge is not free money

A 13-point gap looks like an easy win. It isn't — not yet. Every real trade drags along costs that eat into that gap, and a bet only makes sense if the edge is bigger than all of them combined.

What eats the edge?

So the rule is blunt: trade only where the edge clears every cost with room to spare. Let's run the arithmetic on our clean 13-point case. Start with the raw gap of 13 points. The best ask sits a couple of cents above the midpoint we quoted, so crossing the spread costs about 2 points. Fees on the fill and eventual payout shave off roughly another point. We're buying more than a handful of shares, so slippage up the order book costs perhaps 1 more point. And because we are honest that our 71% is itself uncertain, we haircut the edge by a couple of points as a safety margin. Tally the drags — 2 + 1 + 1 + 2 = 6 points — and subtract: 13 − 6 leaves a net edge of about 7 points. That clears, comfortably, so this one becomes a candidate to size. Change one input — say the raw gap had been 4 points instead of 13 — and the same 6 points of cost turn a tempting-looking opportunity into a guaranteed slow loss. Skip it.

Here is where the honest reality of these markets bites, and it is worth stating plainly. These weather markets are not empty — most have tens to hundreds of thousands of dollars in resting orders. But that depth is overwhelmingly one-sided: it is market-maker sell walls — big stacks of shares offered for sale — sitting above thin, sparse bids. And wherever the price is genuinely live — anywhere in the meaty 10¢-to-92¢ range — the market is already efficiently priced, so our number and the crowd's number tend to agree and there's no gap to trade. Most of the "edges" our model flags are against tail buckets priced at a fraction of a cent — dust on outcomes that almost certainly won't happen — and a minimum-price gate correctly throws those out, because a 0.1¢ "bargain" that will almost surely expire worthless is not a bargain. On a scan of roughly eighteen real, non-dust edges, often only about two survive all the gates. That is not the pipeline malfunctioning. That is calibration doing its job in a market that is mostly fairly priced.

When the window is actually open. The genuinely two-sided, tradeable moment for these markets is roughly a day before resolution — recent enough that a fresh forecast run carries real information, but before the outcome hardens into near-certainty. As resolution approaches, the market collapses toward 0¢ or 100¢ and the gap vanishes. So expect few trades. On many honest days the correct output of this whole pipeline is zero bets — not a bug, but discipline.

Key idea

The edge is the gap minus the spread, fees, slippage, and a safety margin for our own error. We only bet when what's left is clearly positive. Passing on thin edges — and on dust-priced tails a min-price gate rejects — is the feature, not a failure. Zero trades is often the right answer.

05 — Size it

Decide how much to risk

The final station answers a different question. The first four decided whether to bet. The last decides how big. This matters enormously: a genuine edge can still ruin you if you bet too much on any one thing and a run of bad luck wipes out your bankroll before the math has time to work.

The size scales with the edge and with our confidence — bigger, surer gaps get more money; thin, shaky ones get a token stake or nothing. Return to our survivor, the 7-point net edge. A full-Kelly formula might suggest staking a meaningful slice of the bankroll on a gap that size; we deliberately bet a fraction of that — a quarter, say — because full Kelly is brutally volatile and assumes our probability is exactly right, which we never fully believe. Then the hard caps clamp down on top: even if the fractional-Kelly number came out large, a per-market limit might cap this single bet at a small percentage of the bankroll, so one bad settlement can't dent us. The order that finally reaches the market is the smallest of what Kelly wants and what every cap allows.

polyAether does this with that shrunk-down version of the classic Kelly criterion (Chapter 8 covers it in full), plus a stack of hard caps: a limit per market, a limit on total exposure, a daily loss limit, a cap on making too many bets that would all win or lose together (many "different" weather markets on the same hot day are really the same bet in disguise), and a kill switch that halts everything if things go wrong. Sizing is where an edge turns into a survivable strategy — so it gets its own chapter.

A word on speed, because people assume trading edges come from being fastest. For weather, they don't. Our edge lives in seconds-to-minutes — the moments right after a new forecast run publishes or a fresh hourly observation lands — not in nanoseconds. Underneath, polyAether does react fast: a persistent connection watches the order book and responds to changes in milliseconds rather than slowly re-checking on a timer, and the server sits only a couple of milliseconds from the exchange. But that speed is defensive. It keeps us from being picked off by a stale quote and lets us be first to act on genuinely new information — it is not the source of the edge. The edge is calibration; speed just protects it.

1
Forecast
122-member GFS + ICON + ECMWF ensemble produces a probability with a built-in confidence meter.
2
Calibrate
Correct the forecast to the exact settling station — one of ~80 curated thermometers.
3
Compare
Lay our probability next to the market price; the gap is the raw edge.
4
Clear costs
Subtract spread, fees, slippage, and a safety margin. Reject dust-priced tails. Few edges survive — often zero, by design.
5
Size
Fractional Kelly plus hard caps decide how many shares — or none.

Why the whole chain matters

Any single station done well is worthless if another is done badly. A brilliant forecast pointed at the wrong thermometer is wrong. A real edge sized recklessly is a blow-up waiting to happen. A perfectly sized bet on a gap that fees erase is a slow bleed. The chain is also multiplicative in a quieter way: our Chicago market started at a raw 71%, got pulled to 64% by calibration, was measured against price, then had its edge whittled by costs and its stake shrunk by Kelly and clamped by caps. At every station the market lost candidates or gave up size. That relentless narrowing is the point. The value isn't in any one clever step — it's in refusing to skip any of them. Discipline, not genius, is the product.

And the line doesn't stop at "buy." A position that clears every station is then held across cycles and eventually settles on the real, official weather observation, at which point its profit or loss is realized and flows back into the running balance. That full loop — open on a cleared edge, hold, settle on the actual outcome, book the P&L — now works end to end.

One honest caveat, repeated because it's true: polyAether is strictly paper-trading right now — every bet in this pipeline is simulated, with no real money and no proven track record. What does work today is the full simulated lifecycle just described: positions open on a cleared edge, are carried across cycles, settle on the real resolution, and their realized gains and losses move the shown balance. The chain is built, wired together, and tested; it has not yet earned its keep with real capital. Chapter 7 tackles the sneakiest station of all — how a market actually settles, which decides whether your winning bet actually pays.