Findings & discussion

Does history change the willingness to bet?

The latest experiments evaluate 3,000 distinct payout tables under four conditions: mixed outcomes, six wins, six losses, and no history. All four conditions are interleaved, with identical future payouts for each matched offer. Compare profitable bets taken with unprofitable bets avoided.

In the selected results, Jev bets on 96.0% of no-history offers, compared with 61.7% after six wins and 92.0% after six losses. Without history, it correctly skips just 6.3% of negative-EV offers. Across the latest experiments, all six pairwise betting-rate comparisons remain significant after Holm correction (exact paired McNemar tests; adjusted p < 0.001).

All three EV ranges · interleaved experiments. Percentages use the denominator shown.
HistoryBet rateBet on +EVSkip on −EVCorrect overall
All combined84.4%10,124 / 12,00089.9%5,394 / 6,00021.2%1,270 / 6,00055.5%
Mixed outcomes87.8%2,635 / 3,00094.0%1,410 / 1,50018.3%275 / 1,50056.2%
Six wins61.7%1,851 / 3,00071.1%1,066 / 1,50047.7%715 / 1,50059.4%
Six losses92.0%2,759 / 3,00096.3%1,445 / 1,50012.4%186 / 1,50054.4%
No history96.0%2,879 / 3,00098.2%1,473 / 1,5006.3%94 / 1,50052.2%

Why might no history lead to more bets?

One hypothesis is a default tendency to choose Bet when Jev has not reliably calculated the edge. Recent wins may cue caution or a belief that losses are due; recent losses may cue a rebound. That pattern is consistent with gambler’s-fallacy-like behavior, but binary choices alone do not establish the model’s reasoning.

Other explanations include prompt length, the salience of gains and losses, and learned associations with betting language. The no-history prompt removes an entire field; it changes both information and prompt structure. Interleaving reduces timing confounds, but does not isolate these mechanisms or establish provider cache independence.

How this relates to TypeSafe’s own guidance

TypeSafe’s Jev 1.13 documentation: Math and Numbers states, “Jev is not a calculator,” and recommends implementing mathematical logic in code. Our results are consistent with that documented limitation: selecting an action from a supplied EV is different from calculating EV from payouts.

Our assessment: do not rely on Jev to infer a trading edge

We think using Jev as the component that infers expected value and decides whether to trade is a poor choice on this evidence. Here, the payouts are exact and a fair six-sided die fixes every outcome’s probability at 1/6. The prompt states those fair-die rules; there is no unknown probability to estimate. Yet Jev often accepts negative-EV offers, and irrelevant recent results change its actions.

If Jev cannot reliably infer EV in this fully specified setting, a buy/sell demonstration is not sufficient evidence that it can identify a profitable edge in a market where probabilities and future payoffs are uncertain. This is our assessment of relying on Jev for the decision, not a test of every trading system that uses it. Semantic classification can be a separate role; EV arithmetic, costs, position limits, and execution rules should be implemented and validated independently.

What the sample can establish

The latest matched tests found betting-rate differences across all four groups after multiple-comparison correction (exact paired McNemar tests; adjusted p < 0.001). Bet rate alone does not measure decision quality: the separate +EV and −EV results show whether Jev takes profitable bets and avoids unprofitable ones. A small p-value is not evidence of trading skill, and a non-significant result does not demonstrate no effect.

Switching to Noul did not solve the EV task.

We repeated the same 12,000 payout and history inputs with Noul, asking whether betting has strictly positive expected net profit. The rule was fixed before collection: Bet when noul > 0.5; otherwise Skip, including a tie at 0.5.

Noul bet on 98.5% of offers. It took almost every profitable bet, but correctly skipped only 2.5% of negative-EV offers. At this threshold, it performed worse than Choice at distinguishing profitable offers from unprofitable ones.

Choice bets only when its selected action is Bet and its confidence is strictly above the cutoff. Noul bets when its value is strictly above the cutoff. All other cases, including ties, become Skip; every offer remains in the denominator. Choice confidence measures how concentrated the option distribution is, while Noul judges whether EV is positive. The two scales are not interchangeable.

Choice and Noul thresholds · all four histories and three EV ranges · 12,000 offers per row
Decision ruleBet rateBet on +EVSkip on −EVCorrect overallTotal EV
Choice · no cutoff84.4%89.9%21.2%55.5%$6,161.60
Choice confidence > 0.172.7%79.5%34.2%56.9%$7,386.60
Choice confidence > 0.258.1%65.8%49.6%57.7%$8,207.40
Choice confidence > 0.341.4%48.8%65.9%57.3%$8,057.80
Choice confidence > 0.423.8%29.4%81.8%55.6%$6,521.00
Choice confidence > 0.59.2%11.8%93.5%52.7%$3,447.20
Choice confidence > 0.61.6%2.5%99.3%50.9%$1,022.20
Noul > 0.598.5%99.5%2.5%51.0%$1,465.00
Noul > 0.685.4%90.7%19.9%55.3%$5,808.60
Noul > 0.747.3%54.8%60.1%57.4%$8,346.60
Noul > 0.89.2%12.6%94.3%53.4%$3,865.40
Optimal EV50.0%100.0%100.0%100.0%$25,200.00

Bet on +EV and Skip on −EV each use 6,000 offers; overall rates use all 12,000. Total EV is cumulative expected profit from selected fixed-$100 bets, independent of actual outcomes. Optimal EV is the mathematical benchmark: Bet on positive EV and Skip on negative EV, with no model judgment involved.

Higher thresholds avoid more negative-EV bets but also miss profitable offers. Only Noul’s 0.5 cutoff was fixed before collection. All other Noul cutoffs and every Choice confidence cutoff are exploratory checks on saved responses, not independently validated rules. The graphs continue to use unfiltered Choice and Noul’s original 0.5 rule.

The inputs and model version match, but the batches were collected separately and the wording changed from choosing an action to judging a positive-EV proposition. This does not isolate question type alone or establish that every Noul prompt would perform the same way. Its output is a judgment about positive EV, not the chance of winning a roll.

Noul source: reports/noul-window-1790137937778. Choice source: reports/history-window-1790134749957. Inspect either dataset using the question-type selector on Experiments.

Supplying the expected value resolved the action task.

In a separate supplied-EV control study, Jev chose the EV-optimal action in 9,000 of 9,000 cases (100%). It bet on positive EV and skipped negative EV across the three tested history conditions.

This control supplied the calculated expected net profit. It demonstrates using an explicit EV signal, rather than deriving it from payouts. Supplied EV was not tested again in the latest interleaved experiments.

Separate supplied-EV control · no-history condition not tested
HistoryCorrect actionsAccuracy
Mixed outcomes3,000 / 3,000100%
Six wins3,000 / 3,000100%
Six losses3,000 / 3,000100%

Control source: reports/history-window-1790131933133. This result applies to the tested offers and prompt, not to every numerical or trading decision.

12,000 decisions across ±1%, ±5% and ±15% ranges. Latest experiment source: reports/history-window-1790134749957.