What the external critique changes
On this page
The goal remains an exact integer solution of x³ + y³ + z³ = 114. The public product must make the scope of unsuccessful work inspectable, without treating activity as evidence that a discovery is close. This review records decisions after an independent critique on 10 September 2026.
Use measured units, not an invented conversion
The leaderboard's inputs are coefficient-generator positions from independently replayed fixed tasks. It does not reward a client for reporting more seconds or for splitting a task into smaller pieces. The current tasks have fixed bounds, but some regions reject inputs much earlier than others, so this is also not a perfect measure of useful mathematical exposure.
Converting those inputs into reference zcubes core-days requires a paired, coverage-matched calibration. The Python norm sampler, browser BigInt sampler, native C/PARI campaign and Booker–Sutherland divisor enumeration search different domains with different support. A conjectured 20–50× speed ratio is not that calibration. We retain the actual units and publish counters and replay timings separately. The conditional 1-in-1.7-million illustration is a reference-engine frontier calculation; it must not be used as a browser or whole-project win probability.
The portable runner is a convenience and reproducibility client. It is not advertised as a native-speed zcubes wrapper. The upstream implementation needs GMP and primesieve and enumerates a different, explicitly bounded divisor/progression domain. A native campaign expansion needs its own task format, replay or audit contract, platform builds, benchmarks, and coverage proof before entering the same ledger.
Preserve the meaning of verified
Checking an emitted triple is cheap. Certifying a negative task is harder. A client can report plausible prime or progression counts without actually testing every candidate; a hash or analytic expectation does not prevent that. Random audits can make fraud risky under stated assumptions, but a 2–5% replay rate is not a proof that every unaudited task was completed.
The current public label therefore continues to mean full independent bounded replay. This duplicates work and limits cluster scaling; that cost is visible. A future sampled channel must use a distinct audited/reported label, independent sampling, reputation rules and explicit detection limits. It must not silently enlarge the verified exclusion index.
Booker–Sutherland's Remark 5.1 itself cautions against unconditional coverage claims despite its enumeration checks and reports separately rerunning erroneous jobs. That is a reason to keep the distinction, not evidence that expected counts prove honest execution.
Show the lanes rather than imply a new bound
The 81 contexts are three class multipliers × three coefficient shapes × three divisor shells × three quotient bands. The approach page lists their exact bounds. Completed task IDs exclude those finite coefficient blocks; different generators can represent the same root, and unselected roots remain outside the claim.
Our baseline is D₀ = floor(10¹⁹/54). The value 2⁵⁸ is larger than D₀, so a lane below 2⁵⁸ is not automatically inside the historical rectangular search. Historical completion uncertainty and selective support also prevent claiming that every checked candidate is globally new. We do not claim a certified complete bound d ≤ D, |z| ≤ RD, or a completed mathematical paper.
Keep discovery learning as an experiment
The deployed model estimates execution efficiency and retains at least 40% uniform exploration. The held-out discovery experiment failed its promotion gate. Learning cost from verified observations is useful, but the Heath-Brown density heuristic does not assign a known hit probability to each of these norm-generator contexts.
The proposed p = 17 Huisman bias is not a result established in our records. Before computing a significance claim, the statistic, target selection, signs/permutations, height split and null distribution need to be specified. Modular constraints already induce nonuniformity, and a catalogue of known solutions is not an unbiased sample of all solutions. Whole-target and higher-height holdouts, correction for inspected primes, and an actual search-yield comparison would be needed before promotion. No p-value or frontier enrichment is invented here.
Concrete changes in this release
Recorded random seeds, exact published-task exclusion, versioned runner downloads, optional GitHub login and automatic banking, prompt receipt processing, model histories, readable mathematical notes and explicit lane boundaries make the campaign more reproducible. They are engineering improvements. They do not establish a new asymptotic algorithm, global optimality, a short discovery ETA, or a journal-level exclusion bound. The experiment remains moneyless.