dyor.wick.pics/logs

Log

Every update and decision for DYOR and the evolve arena, newest first. Times are UTC. Tap a line to read all of it. DECISION UPDATE

2026-10-07

  1. UGave the other AI its owed items: 102 audit rows (all reproduce within 1e-6) and the Callers no-print lines, both posted.
  2. DChef chose to climb the 60 s+ entry-delay region first, highest first: our best maxima enter in 0-1 s, but live seats trade 60 s+.
  3. UExplorer's trimmed plot rows went live: the page and its data match the build, with the approved note on the Blood grid.
  4. UExplorer now proposes farther-out tests and draws only tested rows; one chart shrank from 31x74 to 28x11 (Chef approved).

2026-10-05

  1. UA 12-hour guard now pauses helper backtests whenever the edge hunter is running a batch, then resumes them afterwards.
  2. UScript seats now trade the way the backtest assumes (Chef's word, all four parts): a buy goes out at the signal plus its entry delay, the referee prices an order the moment it arrives instead of after valuing every run, take-profit fires in real time on the backtest's own test (the price after our own selling impact, before the sell fee), and the backtest's drain exit is added. Replayed on 11-25 Sep (days already searched), the new seat code matched the backtest on all 3,235 signals of 9 setups; the one difference is that it sells at the take-profit print, which the backtest books 60 s later. Seats stay paused until a short paper run passes the timing check.
  3. UBLD-41v2 -5.3%, BLD-55 -12.4%, BLD-56 -15.1% at the 5 h close; all cash spent in the first 1-2 h, no take-profit hit. Script seats paused for now on Chef's word.
  4. UHome FAQ now describes today's seats (5 hours, $500, 10% a buy, live checks); the findings headline now shows 5 Oct numbers and the luck checks.
  5. DThe scripts page no longer shows a win rate (Chef's word): these early seats are not dialled in yet, and it only counted take-profit sells.
  6. UFixed: the take-profit check time went to the first seat every pass, so BLD-56 and BLD-41v2 re-priced their holdings only every 1-3 h after 13:56; the stalest seat now goes first.
  7. UNew seats made 9-10 buys each in hour one, no sells yet, at -0.9%, -3.0%, -3.1%; 10% buys run out of cash near 10 positions.
  8. UCorrected a depth figure: traded pools hold about 6 ETH a side, so a $50 buy moves price ~0.3%; price impact isn't the limit.
  9. UAfter the two speed fixes, a pass over every script seat takes 20-40 seconds instead of 70 seconds to 5.5 minutes.
  10. UScript seats now act on a new dip within about 30 seconds of their entry delay, instead of up to 10 minutes: buys are placed before any take-profit check, take-profit checks share a 15-second budget each pass, and the seat no longer waits while its holdings are valued.
  11. ULag cause found: the seat loop priced every held coin one by one, so buys waited 5-11 minutes with ~25 held vs 55 s with none.
  12. UMeasured over 17 script-seat buys: from the dip printing on chain to the paper fill takes about 3.4 minutes (median), not the 11-63 seconds quoted earlier. Seeing the dip takes ~10 s, the seat placing its order ~2 min, the paper fill ~1 min.
  13. DThree new script seats opened on the standard (5 hours, no trade limit, $500, 10% of the balance a buy), picked to trade often, at the backtest's 60-second entry delay (the closest to live): BLD-55, a 30% dip with 4+ buyers and 5%+ volatility (+8.2% a trade, 95% range +3.5 to +12.9); BLD-56, the same with coin age 1 h+ (+5.9%, +1.5 to +8.5); BLD-41 again, a 25% dip on coins active 48 of the last 72 hours (+5.0%, +1.8 to +8.2). Backtest days already searched; a paper test, not a proven edge.
  14. UBLD-53 and BLD-54 were ended early on Chef's word with no trades: no coin made a 35% dip after 09:03.
  15. DTwo new script seats opened on the same standard (5 hours, no trade limit, $500, 10% of the balance a buy), at the next UNTESTED points of the explorer's one-seat profit-per-day view (Chef's word). Points that differ from BLD-52 only in ways that cannot show live were skipped: entry delay 0 or 5 s (live entries land 11-63 s after the dip) and age or volume floors that never bind. BLD-53 is BLD-52 with take profit +50% instead of +40% (the highest lower bound left: +44.8% a day, 95% range +21.6 to +65.6). BLD-54 is BLD-52 without the volatility filter (the highest dot left: +50.6% a day, 95% range +21.2 to +80.0). Backtest days already searched; a paper test, not a proven edge. [R-0128: wrong, the measured dip-to-fill median is ~3.4 min]
  16. UBLD-51 and BLD-52 finished their 5 hours in profit: BLD-51 (dip on all pools merged) +7.95%, BLD-52 (dip on the main pool) +6.81%.
  17. UExplorer's LIVE button now says what the hunter is doing, such as testing setups or scoring results, and for how long.
  18. UHunter round 105 ran 1,000 setups in about 1h54m; best was +42% a day for one seat, no new record.
  19. UScript seats no longer show ticks: they run live (a check every 5 seconds), so their page now reads 'Live · checks every 5 s' and counts updates. AI runs keep their tick countdown.
  20. UBoth new script seats were down about 15 minutes in (-1.67% and -0.42%), valued at exit prices; none had closed, two buys each.
  21. UEntry delay measured: from dip to order took 11-63 s, to fill 32-76 s. One coin filled at about +93% above the signal price. [R-0128: wrong, the measured dip-to-fill median is ~3.4 min]
  22. UExplorer's DEFAULT button now opens the chart through the highest white dot, the best script-seat result: +8.18% a position.
  23. DTwo new script seats opened, 5 hours each, no trade limit, $500, 10% of the balance a buy (now the standard, Chef's word): BLD-51 is the explorer's highest one-seat profit-per-day dot, which is also its highest lower bound (+52.2% a day, 95% range +22.9 to +81.4), dip measured on all pools merged; BLD-52 is the next distinct setup, the same rule with the dip on the main pool (+51.3% a day). Both: 35% dip, 4+ buyers in the last hour, 2+ recovered dips, volatility 5%+ an hour, take profit +40%, 8-hour hold.
  24. UThe live signal feed learned two backtest filters it did not have: buyers in the last hour, and a dip measured across all of a coin's pools.
  25. UFixed an explorer bug: opening a Callers link showed empty charts because it landed on a view Callers has no values for.
  26. UCorrected an earlier overfit-check post: the splits that fail are quiet vs busy periods, not those holding out 21-22 Sep.

2026-10-04

  1. UCheck on the unused 11-16 Sep days: of 2,592 strong setups, 92% were also positive there, at about 58% of the size.
  2. UCleared up the data window: 8-26 Sep, with 11-16 Sep used to fit, 17-25 Sep to judge and selection, and 27 Sep onward sealed.
  3. UThe results page got a 'Latest, 4 Oct' box and a Blood scoreboard cell: above zero on backtests, exam pending.
  4. UExplorer's DEFAULT button now picks the chart showing the most kinds of dot; a test pick showed 5 kinds in about 6 seconds.
  5. UExplorer's bottom buttons were renamed: DEFAULT, MAX DOTS, MAX GAIN and MAX LOW-BOUND. The five navigation buttons were hidden.
  6. UThe one-seat profit-per-day script seat closed at -0.95% over 4 fills, worst drawdown -2.49%.
  7. UHunter round 83 tested 1,000 setups in about 12 minutes.
  8. DOne new script seat opened on its own: the edge hunter's best one-seat profit-per-day setup (dip 30%, coin older than 48 h, 2+ recovered dips, volatility 1%+ an hour, take profit +40%, 12 h hold). 5 hours, $500, 10% of the balance per buy, no trade limit.
  9. UThe four seats from 2026-10-03 were ended early: +0.11%, +0.46%, -3.04% and +1.58%. Their signal feed had stopped about 9 hours earlier.
  10. DThe edge hunter now never waits: no machine-load pauses, and the explorer publishes after every round (Chef's word).
  11. UFixed the status page saying LIVE with nothing running; it now checks which processes are actually running.
  12. DAll DYOR scripts and AI runs were paused so the edge hunter gets the box (Chef's word).

2026-10-03

  1. UThe six 5-hour profit-per-day seats closed: one finished up 0.24%, five down, from -0.54% to -10.0% (3 to 12 fills each).
  2. UThe live feed was restarted for the six new seats and now holds about 1.5-2.4 GB against a 2.8 GB limit, so it is being watched.
  3. USix more 5-hour script seats opened (BLD-44 to BLD-49), no trade cap: the next two Blood setups by profit per day, the top two by profit per day for one seat, and the top two by the lower bound of profit per day.
  4. DChef chose six more seats: next 2 by profit per day, top 2 by per-day-for-one-seat, top 2 by lower bound; 5 hours each, no limit on trades. Setups needing a filter the live feed lacks pass to the next one.
  5. UThe profit-per-day seats made their first buys: a coin printed 27% under its 1-hour high.
  6. DThe profit-per-day seats got a 36-hour limit instead of 5 hours, because 24-hour holds would be cut short. Chef was told.
  7. UFour new script seats opened (BLD-40 to BLD-43), the top Blood setups by profit per day: $500 each, 10% of the balance per buy, 10 trades or 36 hours.
  8. UA first opening of the same four seats was voided before any trade, because a field was missing from their rules file; they reopened at once with the corrected file.
  9. DOn Chef's word, the Blood seats are re-picked by profit per day (return per trade x trades per day): the top 2 setups plus the top 2 lower bounds, all with at least 30 backtest trades.
  10. UWhy the last seven seats never bought: they needed a 30-35% dip on coins that traded every hour for 3 days, and in 5 hours of a quiet market no coin got closer than 13% to that.
  11. UThe seven 5-hour script seats (BLD-33 to BLD-39) closed with no trades at all, each flat at $500 (0.00%) after 14 checks.
  12. UWhy the new seats had not traded: not a fault. The live feed was running with 2-3 s lag but no setup had hit its trigger yet.
  13. UClosest coins were 5.7% short of a trigger; expected rate is only about 1 signal an hour across all seven seats.
  14. UOlder (V2/V3) pools are now read live from the chain too, so every Blood signal reaches the seats about 7 seconds after its block. Replayed over 6,000 past blocks, 6,115 rows matched the old file on token, side and pool type, with nothing missed.
  15. UThe seven 5-hour seats were restarted on the live feed (Chef's word); the first batch had not traded. They now close about 12:00 UTC.
  16. UThe Blood feed now reads new-pool swaps straight from the chain every 2 seconds: a dip reaches the seats about 6 seconds after its block, not 20-30 minutes. Chain rows matched the old trade file exactly (1,732 of 1,732). Older (V2/V3) pools still arrive with the hourly file.
  17. DScript seats are now live on Chef's word: they check every 5 seconds (was once a minute) and re-quote take-profit every 30 seconds (was 2 minutes). Signals still reach them about 20-30 minutes after the dip, because our trade files run behind the chain.
  18. USeven new 5-hour script seats opened (BLD-33 to BLD-39), one for each distinct rule among the backtest points above +20% per trade after real costs: $500 each, 10% of the balance per buy, no trade cap.
  19. DOn Chef's word, every running script seat was ended early and scored at a live liquidation quote, to make room for seats built from every backtest point above +20% per trade after real costs (5 h each, no trade cap).
  20. UThe 14 seats that were ended: 2 were up (best +0.30%), 9 were down (worst -6.1%), and 3 opened today had not yet traded. Every result is kept with its run.
  21. ULow-fee take-profit +50 seats closed: plain script -0.12% over 20 fills, its AI-veto twin +0.79% over 20 fills.
  22. UA new 7-day low-fee script seat opened (take-profit +40, 12-hour hold) with about $500 of paper ETH.
  23. UThe Docs menu now links to the findings report (/findings), on Chef's word.
  24. UA low-fee take-profit +40 seat that needs 2+ recovered dips closed: +0.25% over 20 fills, worst drawdown -0.52%.
  25. UAnother low-fee take-profit +40 script seat closed: -0.08% over 21 fills on simulated routes (+0.06% counting any route).

2026-10-02

  1. DChef said yes: end the all-pools rule seat early and shut the old blood3 feed, which frees memory for the hunter.
  2. UAll-pools rule seat closed early: -19.5% over 154 fills, worst drawdown -43.6%.
  3. DChef: the shared edge hunt with the partner AI centres on three things: DYOR, the explorer and the GUI.
  4. DChef: partner AI gets data only to 26 Sep and no updater; we run each sealed exam first and release that block only afterwards.
  5. UExplorer got a LIVE button: it cycles through the setups under test, marks them on the 3D plot and refreshes every 20 seconds.
  6. DChef: give the partner AI a full kit (tape, frozen engine, cost model, scam protections) to run edge backtests itself.
  7. UExplorer no longer shows the Blood 2.5% dip setting; a before/after check confirmed every other point is unchanged.
  8. UOn Chef's yes, three helper agents started: Spot signals, faster Callers, and LP costs. Results go on the explorer when they land.
  9. UReport 01 corrected on Chef's word: Callers, LP and Spot are under-explored, not closed. Blood has had 80% of all testing, so their next step is the same full build-out Blood had.
  10. UReport 01 on /findings rewritten for the mission (Chef): it opens with the headline points, brings in every strategy from the edge explorer (2,435 measured setups), and gives each finding a chart plus what we found, what it means, what we do and how high its priority is.
  11. UNew page /findings: Report 01, a short report with charts on what the DYOR tests show so far (best results with their uncertainty, live seats, AI vs scripts vs random). A new report about once a week; older ones stay one click away.
  12. UNew script seats now start with $500 of paper ETH, and each buy is 10% of the current balance, so wins make the buys bigger and losses make them smaller (about $50 a buy at the start). Since 29 September they had started with $5,000 and fixed $50 buys (and $1,000 before that). That was our own setting, not an instruction. Seats already running keep their old setup to the end.
  13. DChef: script seats start with $500 and every buy is 10% of the balance, so wins scale the buys up. Runs already going are not changed.
  14. UA new round of four script seats opened: the edge explorer's two highest results per trade and its two highest lower bounds (BLD-29 to BLD-32, low-fee 25-30% dips, take-profit +50 or +30, 12-hour hold). Each starts with $500, buys 10% of its balance, and ends at 10 trades or 72 hours. The limit is 72 hours, not 4, because these seats make about 4-6 buys a day and hold up to 12 hours.
  15. UTold Chef: 6 of 7 traded hunt seats are up 0.3-1.0%, but their 43 buys span only 8 coins, so it is one bet seen 7 times.
  16. DChef: run a round now, taking the two highest dots and the two highest lower bounds from the edge plots into script seats, ending at 10 trades or 4 hours, longer if there is evidence for it.
  17. UFour new low-fee seats opened (take-profit +40 and +50 with a 12-hour hold, each a plain script and an AI-veto twin) that only buy coins with 2 or more recovered dips: in the 7 days before the signal, the coin's main pool fell at least 20% under its own 1-hour high and then traded back at least 20% above that low within 24 hours, at least twice. Checked against the backtest first: on the coins the live feed covers, the new rule fires exactly the backtest's signals, and the older seats' signals did not change.
  18. UFour new script seats opened (BLD-25 to BLD-28) that need 2+ recovered dips; on 1,245 old signals the new filter matched exactly.
  19. USeat rounds now chain on their own: when a 10-trade round closes, the next opens, with 2 more rounds for each of four setups.
  20. UThe two low-fee seats that are holding up live (take-profit +40 and +50 with a 12-hour hold, each with its AI-veto twin) now chain: when a 10-trade round closes, the next one opens on its own, two more rounds each.
  21. DChef: run more rounds of the low-fee take-profit seats, and add a seat for the new backtest lead (2 or more recovered dips). No more seats for the plain all-pools rule, which is down 12.3% over 296 fills live.
  22. UFixed three things an outside reviewer found: the FAQ now describes the scripted seats that run today (20-minute checks, up to 7 days, AI veto on some), not the old one-hour AI runs; the big score now quotes the start in dollars at today's ETH price, so the dollars and the percent can no longer point opposite ways; and the privacy note now says Cloudflare may add its own cookie-free visitor counter.

2026-10-01

  1. UV4 refit re-score: no K0-K8 or hunt-leader cell passes the multiple-test bar; K3 and K4 flip negative, the rest weaken.
  2. UBlood fill-in batch landed: 175 cells added to the explorer; none beats the best two, only dip size carries across weeks.
  3. UAsked outside reviewers on the forum to check the DYOR pages and the explorer, listing their known faults (Chef's ask).
  4. USecond 2 h double closed: plain script -0.13%, script with Space Bunny veto -0.13%, the same 3 buys (Bunny said go to all 3). No trade reached +30% in 2 h.
  5. UTwo more script vs Space Bunny pairs opened as 10-trade seats (Chef): low-fee 25% dip TP +40 12 h (BLD-21 vs SV-BLD022-ME) and low-fee WETH-only 25% dip TP +50 12 h (BLD-23 vs SV-BLD024-ME), each pair on the identical signal.
  6. UShort codes: each edge-cell script now has its own 3-2 code, BLD-01 to BLD-20 in the order they first ran (Chef). The script + Space Bunny veto seat shows in the AI row as SV-SPAC00-ME; the Bunny momentum run reads A6M instead of H?.
  7. USecond 2 h double opened: plain script (BLD-19) and the same script with Space Bunny allowed to pause or veto each action (SV-SPAC00-ME), on the same signals.
  8. UHarness check closed: same 4 signals and 4 buys in all three 2 h seats. Plain script -0.18%, script with Space Bunny veto -0.14% (Bunny said go to all 4), solo script -2.06% on the headline only because one coin had no simulated sell route at its close (-0.18% on the any-route ledger). No trade reached +30% in 2 h; the result is the closing sale.
  9. UHarness check running (2 h each, ends ~04:31): the 25% dip, every-hour setting as a plain script (r20261001-023053) and the same script with Space Bunny allowed to pause or veto each buy and sell (r20261001-023102), on the same signal; plus a solo 2 h script seat (r20261001-022314).
  10. UBuilt AI pause/veto seats: before each buy or sell the script decides, the AI answers go, skip or wait, with a note. It cannot add trades or change sizes; if it errors, the script's action goes ahead. 11 tests; 4 of 4 deliberate breaks caught.
  11. DChef: run the best current point as a 2 h script and, in parallel, a 2 h Space Bunny run of the same script where the AI may pause or veto each action. The point is to check the harnesses work.
  12. DChef: once the Blood fill batch lands, 2 top-mean and 2 top-lower-bar parameter sets go into script seats at 10 trades each.

2026-09-30

  1. UFixed Scout rows showing initials: public channel names now map to their photos; 3 new fetched, 13 of 13 Scout rows show one.
  2. DChef: research must stay generic, so no backtest engine gets our own fee added; that fee term was withdrawn from the spec.
  3. U4,165 real trades re-scored: after real costs no slice is positive, and none overlaps the backtest's positive cells.
  4. UReferee: when a sell looks far under market and a rival quote was better but unverified, it now asks all sources once more and uses the answer only if it is simulated and near market. Shipped for all runs, open ones included (Chef's word). 101 tests pass.
  5. UCallers fill batch 1: 652 new cells tested, none with a positive lower bar; 15 and 60 min delays are clearly worse.
  6. UCorrection: the closing-sell fault we called unfixed was ~95% fixed by last night's rival-quote change; re-scoring the 7 runs that closed before it improves each by 0.2–1.8 points, none turns positive.
  7. UOpened 2 more trade-count seats (10 trades each) on the screen's two most robust candidates: 27.5% and 25% dip, coin traded every hour of 72 h, low-fee pools, take-profit +40%, max hold 8 h.
  8. DChef said yes to two more 10-trade seats on the two candidates that held without their top coin and best day.
  9. UShipped the tape carry-over for late-resolved trades: a replayed capped hour went from 65.5% to 99.9% complete. Backfill started.
  10. DRule for promoting a strategy: only what two independent sources agree on. First two steps await Chef's yes.
  11. UProposed replaying ~80 live seat trades through the backtest engine, line by line, to find where it is too optimistic.
  12. UOpened 2 trade-count seats (10 trades each, 7-day backstop): 30% dip TP +50 and 27.5% dip TP +40, low-fee pools, only dips that took 20+ min after the high (the slow-dip filter).
  13. UToken-filter backtest: neither filter (slow dip, many sellers) meets its rule, each passes 1 of 3 cells; one coin is still the top coin in every cell. Nothing here is a find.
  14. UThree 12 h low-fee script runs closed slightly down: -0.34%, -0.66%, -0.45% in ETH (5, 7 and 5 buys). 14 of 14 script runs negative.
  15. UBuilt trade-count seats: a seat stops buying at its target, and a checker every 20 min ends the run once all of them are sold; the referee scores it as usual. A 7-day backstop closes a stuck seat.
  16. UDesigned a nightly edge screen; hand round 1: 103 of 903 points look positive, none pass; too few independent trades.
  17. DChef: judge script seats by completed trades, not hours: 10 trades per seat, checker every 20 min.
  18. UExplorer opens without the 5 best coins; DYOR runs added as white dots: -24 to -60% per trade vs backtest -5 to +7.
  19. UAudit of 14 top explorer cells: one thin-pool coin is in all 14, 12-51% of totals; no cell is robust.
  20. UAutopsy of live losses: mostly referee and routing (side-pool fills 32%, bad closing quotes 23%); real price moves 17%.
  21. ULow-fee hunt re-scored with the hop cost and loaded into the explorer; froze the plan for re-testing on the V4 backfill.

2026-09-29

  1. UShipped two referee scoring fixes for simulated holdings; tests pass and open runs were kept.
  2. UCorrected an error: the V4 backfill adds no fresh days; only the sealed later blocks are truly unseen.
  3. UHunt 2's best-lower-bar run closed: -1.9% in ETH (-$95), 7 buys, 1 take-profit (+50%). 11 of 11 script runs negative today; no Space Bunny run.
  4. UExplorer loaded the third hunt: two new dials (main pool quote, capital cap) and 2 and 10 min entry delays, 32 points.
  5. UThree 12 h low-fee hunt runs closed, all down: -1.2%, -2.5%, -0.9% in ETH (two scored one unpriced coin as $0; priced normally -1.8%, -0.2%). 10 of 10 script runs negative today; no Space Bunny run.
  6. UThree 12 h low-fee runs closed, all slightly down: -1.2%, -0.3%, -1.7% in ETH. 13 buys, none sold. None positive, so no Space Bunny run.
  7. UFour 12 h top-dot runs closed, all down: -3.1%, -7.5%, -3.4%, -4.2% ($192-$406 on $5,000). 50 buys, 2 sold at take-profit.
  8. UThose finals still carry both known referee faults (unshipped fixes): 7 holdings scored $0; priced normally the losses are 2.3-6.6%.
  9. UOpened a 12 h paper run: 25% dip, TP +50, low-fee pools, signals only from WETH-quoted pools (best lower bar +6.4, 11 coins).
  10. DChef said yes to a fifth script seat restricted to WETH-quoted pools, from the third backtest hunt's best lower bar.
  11. UThird backtest hunt: 5-10 min entry delay costs ~1-4 pts/trade; best WETH-only cell +19.6 (lower +6.4); none passes correction.
  12. UOpened a 12 h paper run: 27.5% dip, TP +40, low-fee pools (backtest +13.7 per trade, lower bar +4.1, 46 trades).
  13. UOpened a 12 h paper run: 30% dip, TP +75, low-fee pools (backtest +16.8 per trade, lower bar -4.5, 34 trades).
  14. UOpened a 12 h script for the hunt-2 best lower bar: 25% dip, traded every hour, TP +40%, 12 h hold, low-fee pools.
  15. UBacktest hunt 2 (160 cells, stock hop now costed): best lower bar +4.4; this morning's leaders were 1-2 pts too high.
  16. UEarnBench /entrance now lists models by highest gain first, across every run length (Chef).
  17. UChef asked for a verdict on the backtesting: parts are sound, but the search is flawed (~2,930 tests on the same 9 days).
  18. URead the outside reviewers' long critique of the edge work and posted a reply accepting three of the explorer principles.
  19. ULive vs backtest for rule #3: live lost 10.9 pts/trade (28 trades), backtest at the same delay +4.9. Late entries and wrong venue.
  20. UOpened 3 more 12 h low-fee scripts from the backtest hunt: 30%/TP+50/12 h, 25%/TP+30/12 h, 30%/TP+30/12 h.
  21. ULow-fee backtest hunt (108 cells): best +19.1%/trade, best lower bar +2.2; none passes the multiple-testing check.
  22. UThe chart no longer draws a reading with no price at all as a crash (a coin the referee could not price counted as zero).
  23. UChart spikes were mostly measurement: 33 of 35 big jumps had an unpriced holding counted as zero; 13 of 108 old finals affected.
  24. UOpened 3 x 12 h low-fee script runs (pools taking <=1% a side): 40%/48 h, 30%/every hour, 25%/every hour.
  25. UEdge feed gained low-fee cells; replay matched the backtest on 193 of 203 signals (the 10 missed are dynamic-fee pools).
  26. DChef: promising backtests run here as scripts early; any positive script gets promoted to a Space Bunny AI run.
  27. UOpened 4 x 12 h runs of the explorer's top dots: rules only, buying all window, $5,000 stakes, $50 clips.
  28. DChef: run each strategy's own rules the whole 12 h and read the balance at the end; no extra time rules.
  29. UReplaying Sep 8-26 with the chain clock kept every V4-triggered signal; the 24-35% lost per cell were all V2/V3-triggered ones.
  30. UEdge feed now reads each block's time from the chain itself: signals arrive ~2-3 min after the dip, not ~50.
  31. DChef chose A: a chain clock for the edge feed; the next top-4 runs will be 12 h, not 28 h.
  32. URead the 4 explorer top-dot runs: 33 signals in the buying window, 21 too late, 11 refused, 1 trade (-1.2%).
  33. UWhy late: V4 trades get a time only once the hourly V2/V3 tape passes their block, so signals land ~50 min late.
  34. DChef set the edge explorer's uncertainty display to Off by default; it's live.
  35. UThe four 28 h top-dot runs closed: one made a round trip and lost 1.2% (about $977); the other three never traded and stayed flat.

2026-09-28

  1. UA scripted Blood venue v1 leverage test closed by liquidation, down about 21.3% (about $77.61).
  2. UA scripted BloodSelect venue v1 leverage test closed by liquidation, down about 19.1% (about $79.84).
  3. UFixed the referee refusing buys of V4-only coins: gas is now priced at the buy's own quoted rate (about 0.02% of a $50 trade). 33 of 37 earlier refusals would have filled.
  4. DChef approved the referee gas fix (option A).
  5. UA scripted BloodSelect v2 leverage test closed by liquidation, down about 17.4% (about $82.26).
  6. UA scripted BloodSelect v1 leverage test closed by liquidation, down about 9.8% (about $89.78).

2026-09-27

  1. UOpened four 28 h scripted Blood runs, one per top dot on the edge explorer (30%, 25%, 40% dips on coins that traded every hour; 40% dips on coins that traded 48+ of 72 h). Buys happen in the first 4 h only.
  2. UThe new four-rule signal feed fired exactly the backtest's signals on the backtest days: 524 of 524, same coins, same moments, same pools.
  3. DChef asked for the edge explorer's four highest dots to run on DYOR as scripts for 28 hours.
  4. UThe V4 swap collector now writes every 5 minutes, not hourly (Chef). Rule #3 signals are still up to about an hour late: the V2/V3 swap tape, which also sets the feed's clock, is still written once an hour. The 11:0x line blamed only the V4 collector, which was incomplete.
  5. DChef approved adding 2,713 stock-token-priced pools to the backtest universe, and approved speeding up the V4 swap collector.
  6. UA scripted Liquid Blood v1 leverage test that started Sep 25 closed by liquidation, up 27% (about $127).
  7. UA follow-up post said real trading costs turn the Blood edge to -2.9%, and its best zone wasn't repeated elsewhere.
  8. UA memory-heavy background job was stopped so the new run could start; it's blocked from running again for the rest of the day.
  9. UBlood rule #3 is running in public: run r20260927-110350, paper, $1,000, 7 days. Signals arrive up to about an hour late for now, because the V4 swap collector writes to disk once an hour.
  10. UA second copy of that run was opened by mistake 7 seconds later. It made no trades and was voided.
  11. DRule #3 will run in public (Chef). Its signals overlap part of rule #2's sealed test, and he chose to accept that. It starts when the box has memory free.
  12. UBuilt a scripted bot for Blood rule #3, exactly as frozen (Chef): paper only, $50 a trade, not opened yet.
  13. UIts signal code fires the same 203 signals as the backtest over Sep 8-26; fed the live-style files, it caught 18 of 20.
  14. UNot started yet: our V4 swap collector writes once an hour, so signals would be late; its timing also needs a ruling.
  15. DDecided to test trading edges in two steps: a frozen script alone, then paired with an AI that can only pause or veto trades.
  16. DLocked the final sealed-exam rule's exact settings and file hashes; it runs once, no earlier than October 7.
  17. UA deeper test found the dip-buying edge held up against every check; it's a bounce effect, not smart crash-avoidance.
  18. UStarted a backtest testing dip depth from 10% to 40% against trading history length, to check an explorer's guesses.
  19. UCompressed nine days of swap data losslessly, shrinking it from 1.8 to 1.2 gigabytes.
  20. UPosted the edge research and results to the public forum, along with four questions for outside review.
  21. UDYOR's automatic bug-catching tests still catch every real bug, 156 of 156; 5 misses were for limits already removed.
  22. UA review of the four 2-hour Space Bunny tests found it lost to its control every time; momentum fell 44.5% from overtrading.
  23. UProposed charging real round-trip trading costs and capping trades per hour, after overtrading sank the momentum test.
  24. UEarnBench /entrance is now one table: a Harness column, a Run length column and a Result column, one row per model, harness and run length.
  25. DFrom now on, new runs get rival sell prices that are actually simulated: the engine lends the paper account the coin inside the check, so a Nordstern-only coin is priced for real instead of counted as zero. Runs already open keep the old pricing to the end.
  26. UFound DYOR's mutation tests had a broken control, so earlier claims that every code break gets caught were never verified.
  27. UThe leverage 2-hour test finished: Space Bunny down about 0.36%, its random-trader control down about 1.7%.
  28. UThe momentum 2-hour test finished: Space Bunny down about 44.5% from overtrading, its random control down about 1.5%.
  29. UThe Blood 2-hour test finished: Space Bunny down about 17.9%, the scripted control bot up about 4.7%.
  30. UThe LP 2-hour test finished: Space Bunny down about 4.1%, the scripted control down about 1.8%.
  31. UThe LP round's random-trader control looked down 35% by quick pricing, but the full ledger showed it actually up about 2.2%.
  32. UA 2-hour scripted Blood-v2 test that started Sep 25 finished, down about 10.8%.
  33. UFixed: hover tooltips on some hollow dots in the /stats chart weren't clickable; the whole dot area now responds.
  34. UA statistics check found Space Bunny's edge over random trading isn't proven yet; one slice hinted at a possible +1.7% edge.
  35. DPast runs that counted an unpriceable coin as zero are rescored at the quote that existed for it, and starred; the strict figure is kept. A coin nothing would buy still counts as zero.
  36. UEarnBench /entrance now groups scores by window and chain only. The starting ETH amount had split almost every run into its own group, so most effort levels showed no result.
  37. ULeverage run opened (r20260927-022408): its first turn went 3x long ETH, liquidation price $1,841; beside a random trader.
  38. UMomentum run opened (r20260927-022512) WITHOUT a plan: its research plan ran over the 6,000-character cap; beside a random trader.
  39. UAn order whose database write hits 'database is locked' is now retried up to 3 times instead of being refused.
  40. UAn answer the harness cannot read now gets one re-ask in the same turn (format only, no new data), recorded as format_retry.
  41. UHoneypot 'unknown' now gets one re-check 8 s later before a buy is refused, and the refusal says the checker's own reason.
  42. UForgotten coins found in 5 scripted runs: BLD-SV1 2, BLD-V1 2, BloodSelect v2 1, v1 1, Liquid Blood v2 4; Liquid Blood v1 none.
  43. UFixed: scripted Blood bots forgot a coin when its sell was refused; a refused sell is now retried until the coin is gone.
  44. UScripted script-gz bot's zero trades explained: noxa sent no Golden-Zone signal between Sep 24 21:00 and Sep 26 21:00.
  45. UMomentum and leverage runs started their research turns at 02:23; each opens beside a random trader when its plan is ready.
  46. UNew paper leverage: an ETH perpetual priced from Hyperliquid's public data, up to 5x, fees and funding charged, wicks between turns liquidate.
  47. UNew momentum feed: liquid coins that print 30%+ above their 1-hour low, rebuilt every minute; the momentum run trades only from it.
  48. UBlood run opened (r20260927-021417) beside a 2-hour scripted Liquid Blood bot on the same dip feed; first try hit a quote-engine restart.
  49. ULP run opened (r20260927-021013) after a 4-min research turn, beside the scripted LP bot and the random trader.
  50. DChef: four 2-hour Space Bunny runs at once, one per instrument (leverage, LP, momentum, Blood), each beside a control bot.
  51. UChef was asked to decide whether the results score should flag incomplete valuations with a range, or keep showing zero.
  52. UA review of 70 AI runs found the score can zero a coin's value when its last price is missing, hiding up to 26.9 points of gains.
  53. UThe review found the four space-bunny copies from the Sep 26 2-hour test actually all beat random once every real trade counts.
  54. UThe review found the safety check wrongly refuses trades when a coin's honeypot status is unknown, blocking 14 of 186 buy orders.
  55. UThe review found a Blood test called 'no signals' actually tried 7 buys wrongly blocked for two hours; the coin trades fine today
  56. UThe review found 34 of 746 AI trading turns failed to answer properly with no retry, including 18 of 356 for Space Bunny.
  57. UThe review also flagged a script bot's forgotten-sell bug and a mismatched test set for the random trader.
  58. UOpened a 2-hour test pitting Space Bunny (high effort) against a Blood script bot, both trading the same coin swings.
  59. UOpened a 2-hour test pitting Space Bunny (high effort) against a script bot and a random trader, in a liquidity-pool setup.

2026-09-26

  1. UThe random trader in the 2-hour noise-floor test finished, down about 2.9%.
  2. UThe minimal-effort copy in the 2-hour noise-floor test finished, down about 17.2%.
  3. UThe xhigh-effort copy in the 2-hour noise-floor test finished, down about 12.4%.
  4. UThe medium-effort copy in the 2-hour noise-floor test finished, up about 3.9%.
  5. UThe max-effort copy in the 2-hour noise-floor test finished, up about 2.2%.
  6. UOpened a 2-hour noise-floor test: four Space Bunny copies (max, medium, xhigh, minimal effort) plus a random trader.
  7. UThe 5-hour re-run of the +3.37% winning Space Bunny test finished, down about 2.3%.
  8. UChef tested a blood-script idea of waiting longer than one minute to buy — result: no better than the 1-minute wait.
  9. UThe 2-hour best-chance Space Bunny test finished, down about 7.3%.
  10. UThe re-run 12-hour max-effort Space Bunny test finished, down about 10.2%.
  11. UThe fresh 12-hour random trader finished down about 8.6%.
  12. UA copy-mirror bot run finished down about 16.9%, cut short by a coin it could no longer price.
  13. UThe Sniper script's run finished exactly flat with zero trades, matching its quiet signal feed.
  14. UThe Gold Zone script's run finished exactly flat with zero trades, matching its quiet signal feed.
  15. UA copy-bot run finished up about 0.3%, close to flat.
  16. UBLD-B1's run finished with zero trades, exactly flat.
  17. ULIQ-1's run finished, down about 9.3%.
  18. UThe headline no longer says Unpriced when one coin cannot be priced: it shows the last good market-price estimate and when it is from.
  19. UOpened two Space Bunny runs for Chef: a 2-hour best-chance test (high effort, 5-minute ticks) and a 5-hour re-run of the +3.37% winner (high effort, 10-minute ticks).
  20. UFixed bags showing Unpriced: the referee now has its own price-check allowance on the engine and re-asks when the engine is busy or answers empty.
  21. UOpened two new Blood script bots: BLD-V1 (liquid Blood that skips mass-launched V3 1% coins, buys 1-3 minutes after the dip, takes profit from the fill) and BLD-SV1 (the history-gated BLD-S2 with the same filter).
  22. UThe re-run 9-hour max-effort Space Bunny test finished, down about 8.4%.
  23. UOur swap engine can now buy and sell coins whose only pool is Uniswap V2, so arena bots no longer have V2-only dips refused; the referee's receipt check also stopped giving up on tokens like SHRUB.
  24. DChef chose to add a V2 route to the swap engine before opening the new venue-filtered Blood bots.
  25. UBuilt two new Blood bots, BLD-V1 and BLD-SV1 (Blood skipping mass-launchpad coins, take-profit measured from the real fill), but held them unopened: most of their backtest trades are on coins whose only pool DYOR still cannot fill.
  26. UNew page /scripts: every scripted bot's P&L over time, where each run stands, and a collapsed table of all runs, read live from the arena.
  27. DChef decided the new script-bot results page will show only graphs, and moved it to DYOR's own site instead of EarnBench's.
  28. UChef named picking the coin and dip-size-plus-timing as Blood bots' two hardest settings, guiding the ongoing audit.
  29. UChef asked for a check that every profitable Blood script bot's paper results are real, starting a deeper audit of its rules.
  30. UThe re-run 6-hour max-effort Space Bunny test finished, down about 7.3%.
  31. UThe re-run 3-hour max-effort Space Bunny test finished, down about 4.3%.
  32. UA price audit's final report found stale frozen prices could mistrigger stops, though no live trades are affected now.
  33. UA price-accuracy audit found DYOR's trade prices match the real blockchain almost exactly, for both buying and selling.
  34. UEarnBench Stats plots finished scripted bots beside the AI runs and the random trader, as violet hollow squares; they stay out of the model funnel (Chef).
  35. UScripted bot pills read TYPE-VARIANT and version (Chef): BLD-L = LiqBlood, BLD-S = BloodSelect, BLD-B = Blood, LIQ- = LP, CPY-C = copy every call, CPY-M = copy mirror, GLD- = Gold Zone, SNP- = Sniper.
  36. UFixed the safety check that wrongly flagged coins like astro as scams — it missed one pool type, causing 65 of 77 wrong refusals.
  37. UFound the copy-bot page falsely says live trades are paper-only (they're real); a display bug showed a huge wrong number.
  38. UThe random trader that ran beside the first max-effort test batch finished its 12 hours up about 10.3%.
  39. UBLOODSELECT1 and 2 opened: LIQBLOOD plus a history gate (only coins that bounced from at least half of 2+ earlier dips), selling at +30 and +50; 48 h, entries in the first 24 h (Chef).
  40. DRun the blood-select research rule as paper scripts (Chef): the first Blood variant that passed on data it was not built on.
  41. UDYOR reviewed 64 finished runs: no AI model beat the random trader on average, and trading more often did worse (Chef).
  42. UThe 3/6/9/12-hour max-effort Space Bunny tests re-opened as a clean re-run, alongside a fresh 12-hour random trader (Chef).
  43. UThe re-opened managed-Blood 2-hour Space Bunny test made no trades at all and finished exactly flat.
  44. UThe random trader running beside the four 2-hour Space Bunny tests finished down about 3.0%.
  45. UThe 2-hour Space Bunny managed-LP test finished up about 2.3%.
  46. UThe 2-hour Space Bunny web-only test finished up about 8.1%.
  47. UThe 2-hour Space Bunny control test finished up about 3.4%.
  48. UResearch found the live price feed only tracks already-popular coins, making past Blood tests look too good.

2026-09-25

  1. UEarnBench Stats: colouring by Model keeps each lab's colour and tells models apart by shape — the lab's biggest model gets the biggest-area shape: diamond, square, circle, hex, cross, triangle, star (Chef).
  2. UEarnBench Stats labels stealth models' lab as Stealth, not Other (Chef).
  3. UEarnBench Stats added Fable as a new top Anthropic model tier, ranked above Opus (Chef).
  4. UEarnBench Stats: each model's dot also grows in size by its rank within the lab, biggest model gets the biggest dot (Chef).
  5. UBlood scripts now show what they are fishing for: each tick names the coins watched and how far each is from its buy trigger, and the Bags tab lists them under Waiting to buy (Chef).
  6. UFound the Gold Zone and Sniper scripted bots have made no trades since Sept 24 evening — their signal feeds went quiet.
  7. UVoided runs now read VOID on the Past runs list and can never show as a model's best run; a run whose final mark could not price a holding shows a *.
  8. UThe managed-Blood run was voided after 4 minutes (our prompt let it open an LP position when the feed was empty) and re-opened with the fix.
  9. UFour 2-hour Space Bunny tests opened, 10-minute turns, beside a random trader: standard (control), web only, managed Blood, and liquidity only.
  10. DThe 3/6/9/12-hour max-effort Space Bunny runs were deleted: our own request cap stopped their turns after ~1 h 40 min. A clean re-run is queued (Chef).
  11. DNo request caps of our own any more (Chef): the provider's reply is the only limit. Space Bunny was never in OpenRouter's free allowance.
  12. UFixed a bug that silently blocked every Space Bunny round since noon: scripted bots wrongly counted as busy (Chef).
  13. ULIQBLOOD2 opened: LIQBLOOD1's rule with twice the liquidity bar and a 6-hour hold instead of 24 (Chef).
  14. DScripted bot pills drop the 'SC-' prefix: LP1, LIQBLOOD1 and so on (Chef).
  15. UDocs now opens with links to EarnBench and its Stats page (Chef).
  16. UOne of the max-effort Space Bunny runs finished, up about 0.8% (later deleted: our request cap had stopped its turns).
  17. UThe first of the four max-effort Space Bunny runs (3 hours) finished down about 6.7% (later deleted: our request cap had stopped its turns).
  18. UThe strategy funnel page went live: 126 theories tried, 87 back-tested (20 looked profitable), 25 replayed on real tape.
  19. USC-LIQBLOOD1 opened: Blood with no stop on any coin that keeps trading, the closest of 16 new theories (it did not pass); 48 h, entries in the first 24 h.
  20. DTheory batch 2: 16 new noxa.bot strategy ideas tested, none passed; the closest goes to DYOR as a paper script bot (Chef).
  21. UThe SCRIPTS pills now show the same red/green up-down bar as the AI pills; they were never drawn before (Chef).
  22. UThe four relaunched max-effort Space Bunny runs go ahead with no research plan, after its planning step hit its limit twice.
  23. UFixed the launcher so the scripted bots no longer count as AI slots; the 4 max-effort Space Bunny runs were relaunched.
  24. DScripted bots now show on the site under a made-up lab name, 'WICK AI', based in 'The Continental'.
  25. UThree more scripted bots opened: one follows a signal, one snipes new coin listings, one mirror-copies other traders' calls.
  26. UA third scripted bot, script-copy, opened its first 24-hour run, buying every trading call it sees on the signal feed.
  27. UA second scripted bot, script-blood, opened its first 24-hour run, alongside the scripted LP bot.
  28. DA second row, SCRIPTS:, joins the AI: row (Chef). A scripted bot's rules are written and frozen before its run opens and a script follows them, with no AI while it runs, on the same referee, live quotes and LP simulator as the AI runs. New scripted bots are added beside the running ones, never in place of them.
  29. UFirst scripted bot SC-LP1 opened: a 24-hour LP bot. It uses V3 WETH pools on coins at least 48 hours old (skipping launch-day volume), TVL of $10k or more, volume not collapsing, the top 3 by last-24h fee APR, a ±50% range, holds to the end, a stop at 65% of the whole position's value and no re-entry. It opened musebook, UNK-d9db30 and URANUS on its first tick.
  30. UThe chart labels every 2 hours on long windows (every 3–4 on a phone) so the time axis stays readable (Chef).
  31. DSpace Bunny rounds move to 2 hours with every copy woken every 10 minutes (Chef). The first 2-hour round plays efforts medium, max, minimal and xhigh, so medium and max trade for the first time. Bunny passed the entrance exam at max as well: all six effort levels have passed.
  32. DThe referee now takes a sell check right after every turn, plus clock checks six times per window (every 5 minutes on a 30-minute run). A 30-minute run had shown $100 'worth now' all run because its first 30-minute check was its end (Chef).
  33. UFixes from the first effort round: the referee waits up to 60 s for its database instead of 5 (an order had been refused 'database is locked'); a valid research plan with extra fields is kept; LP positions' stale 'gap' labels are cleared once the blocks are read; the A6 harness shows as one code again.
  34. ULP gas measured on chain: a position-manager mint uses a median 402,777 gas (the referee charges 400,000); a close uses 245,018 (the referee charges 300,000, kept as a conservative charge).
  35. UEvery run is now in one table, one row per run with every setting and result, rebuilt after each run closes (data/runs.csv).
  36. DSpace Bunny rounds are now 30 minutes, every copy woken every 5 minutes (Chef). Four effort rounds are queued with the efforts rotated across copies; the random trader now wakes on the same 5-minute interval.
  37. ULP profit check: paper LP fees were recomputed independently from the pool's own on-chain swaps and match the referee to within 0.1% after the pool's protocol cut. Our paper position is 0.01-0.02% of the pool's liquidity. Open: the gas figures are assumed, and a stale 'gap' label is never cleared.
  38. UData safety: the referee database had no backup. A clean copy is now taken nightly and sent to the NAS with every run's files. Every tool reply a model receives is now stored too, not just its size.
  39. UFAQ grows to 45 questions with lessons from the related-work scan: how EarnBench differs from AI trading contests, why no AI judge, why a live market cannot be memorised, and prompt-injection safety.
  40. UMethods paper: new section 6 on related work, with 18 references each checked against its source (live trading benchmarks, trading-agent papers, harness and selection work). Two outdated facts in section 2.2 were corrected: the price source and per-run wake intervals. The FAQ now names the closest projects.
  41. UThe run page's tick counter and time-bar notches now use each run's own wake-up interval (5, 10, 15 or 30 min) and the times its turns really ran. A 5-minute run had read 'Tick 7 of 7'.
  42. UEntrance exam fix: a tool call that fails on our side (a data source down or slow) is no longer counted against the model. Space Bunny's xhigh exam had failed on a look-up we timed out; it passed on the re-sit, so Bunny has now passed at every effort from minimal to xhigh.
  43. UNew EarnBench FAQ at earnbench.wick.pics/faq: 39 short questions on what it does well, how it works, common misreadings, its limits and what comes next. Linked from the home page with Stats.
  44. UThe entrance exam can now be sat at a chosen effort. An effort other than high shows as its own row on the entrance page (e.g. Sonnet · low). Space Bunny is sitting it at minimal, low, medium and xhigh.
  45. UEach run pill now has an up/down bar behind its letters: a centre line for the start, green growing right when up, red growing left when down. It shows the same figure as the run's Worth-now box. Scale: at least 3px for any move, shared across the pills, held between ±5% and ±20%.
  46. DEDGE-1 is paused after round 1 until Space Bunny stops being free (Chef). Its frozen stack is unchanged; it resumes by setting its status back to running.
  47. DSpace Bunny rounds now run back to back: 4 copies, 1 hour, plus the random trader, within the daily OpenRouter budget. Each queued round gives every copy the same prompt and varies one thing: effort first, then wake-up interval, then both again with the copies swapped.
  48. UEffort check on Space Bunny: it accepts minimal to xhigh, but 3 identical questions per level showed no clear difference in answer length (xhigh was the shortest). The effort round will test it on real trading.
  49. DSpace Bunny may now be used freely while it is free (Chef). The spend guard still checks its live price before every call and stops the moment it is not $0.
  50. DNaming rule for Space Bunny's results: if its lab later says it was the same model all along, its dots are renamed to the real model; if the released model is different from the preview, they stay 'Space Bunny Alpha'.
  51. UToken-count fingerprint: Space Bunny splits two test texts exactly as MiniMax's public tokenizer does (79 and 43 tokens), and unlike OpenAI's, Qwen's, Gemini's, DeepSeek's, GLM's or Nemotron's. The same test matched OpenAI's gpt-oss to OpenAI's tokenizer exactly. Evidence, not proof: its lab has not said.
  52. DRule: a bot that takes no action is disqualified, not scored (new prompt trial-5, and back in the arena). The pre-registered EDGE test keeps its original prompt.
  53. UFound and fixed: a test wrote fake placeholder market data into the live cache; the four Sonnet effort runs saw it on 2-3 turns, so they are voided from all scores and need a clean re-run.
  54. UThe top bar now shows a pill for every live run (5 with the random trader, which reads just RN); on phones, pills of the same model show their effort.
  55. DNext: four Space Bunny runs of 1 hour after the Sonnet round, woken every 5, 10, 10 and 15 minutes, each told its own wake-up time.
  56. DSpace Bunny Alpha (an anonymous model, free for now) passed the entrance exam; it is used only with a per-use approval, and the spend guard still checks its live price on every call.
  57. UEarnBench stats: new Activity view shows what each bot did (trading, providing liquidity, both, or held) as pie-slice dots from its filled orders.
  58. UEarnBench stats: the runs chart can now group rows and colour dots by run length, lab, model, harness or effort, in any combination; every dot is the same size.
  59. UEarnBench stats dots are now all one size, since not every model publishes its parameter count to size them by.
  60. UThe full list of every model variant on EarnBench stats is now collapsed by default; tap to open it.
  61. UEvery EarnBench page now shows the small dollar-bill icon from the home page, as its favicon.
  62. UEarnBench's connection to GitHub's free model catalog works now, but stays on hold until paid usage is confirmed off.
  63. UTwo more entrance exams were built, for Gemini's and Cohere's free models, using the same flexible answer format.
  64. UEarnBench stats gained a second chart: each run's result against trades made, research look-ups or compute cost, with the correlation printed as a hint (more trades went slightly with worse results so far).
  65. UThe runs plot on EarnBench stats now colours and shapes each dot by the lab that made the model, sizes it by the model's published size, lets you pick what the rows compare, and opens a run on tap.
  66. UThe EarnBench pages now rebuild and publish themselves within about 5 minutes of any run finishing.
  67. DEarnBench reports each model's best run, not its average, always shown with how many runs it had (e.g. best of 4).
  68. UEvery result on the EarnBench entrance exam page now links to the bot page of the run that produced it.
  69. DPlan turns will now tell models their budget (up to 20 minutes and 10 live quotes, 'use them'), on the new A6 harness; they were finishing in seconds with no research.
  70. UFixed: plans written over several lines were thrown out as broken JSON; three of the four Sonnet plans were lost this way, so that round is trading without plans.
  71. UThe share card's % is now a chunky pixel-block sign, like the robot's own screen.
  72. UFirst EDGE-earn round finished: AI copies averaged about +0.4%, beating the random trader's -1.6%, with ETH roughly flat.
  73. UNew page earnbench.wick.pics/stats: every finished run plus the funnel from models examined to passed, ran and profitable (31 → 14 → 6 → 1 today).
  74. UThe entrance exam page now links straight to a harness (/entrance/#A5), and A5 shows its first 4-hour score.
  75. UMore free models: gpt-oss-20b passed through NVIDIA's free catalog; Nemotron Ultra now also passes through NVIDIA; north-mini-code and Nemotron Super pass on A6, which accepts an answer with a sentence around it.
  76. DThe Sonnet effort round was cut to 1 hour to measure its cost first.
  77. DDirection agreed: bots will start from real trading styles (trend, reversal, liquidity provider, flow follower, sentiment, launch) with risk rules built in, and our own profitable hand-traders will be studied as a playbook; tested after the current EDGE test.
  78. UShare now sends a link to that bot's own page (live or finished) through the phone's normal share options; the picture card is a separate Save image button.
  79. UThe entrance exam page now lists one entry per harness (e.g. A3), with each model showing which provider it was reached through.
  80. UFixed: the random trader never traded while other runs were open — it read the wrong run and skipped every turn, so its 0% results beside the copies meant nothing. It now plays its own run [R-0095].
  81. UThe robot's name plate is now centred under the robot, not left-aligned.
  82. USonnet passed the entrance exam at its lowest effort on the new C4 harness (Claude Code with shuffled lists): 100 on all three turns.
  83. DOne Sonnet round is next: four runs identical except effort (low, medium, high, extra-high), same 4-hour window, with the random trader; the EDGE test pauses for that one window.
  84. UThe copy-trading fix for newer V4 pools was measured first: it would have unlocked 0 of the last day's 66 copied coins, so it was not built.
  85. UEarnBench's technical notes page now lists every change as one short line, using the full width of the page.
  86. DHarness codes now name the harness, not the provider: C = Claude Code, A = our own agent loop; the provider is shown beside it (e.g. A3 via Gemini). O3 is now A3, G1/K1 are A3 via Gemini/Cohere.
  87. UFixed: Past runs could hang on Loading for 25+ seconds while the referee priced open positions; the list now reads without waiting on trading.
  88. UPlanted-bug check after today's changes: all 159 planted bugs caught, after strengthening one weak shuffle test and repairing two outdated checks.
  89. UA check of the first already-run 4-hour round showed the AI copies down 6.28% together, while doing nothing was flat.
  90. UTest started: 'does our AI earn?' — 12 four-hour rounds of the frozen setup (4 copies + the random trader), scored against simply holding ETH or dollars. Written down before the first round.
  91. DFrom now on, one pre-registered question at a time, and a 'winner' only counts if it does it again in a fresh block of rounds.
  92. DThe test setup is frozen for the whole block: any change to the harness, prompt or referee stops new rounds until it is written down as a deviation.
  93. UChecked: paper fills already use the live price at the moment the order arrives, for the exact size, with gas — much closer to real than noxa's old paper exits.
  94. UEvolve v2 built: a seat is cut only if its losses are bigger than luck explains; only a leader that beats luck passes its lessons on; no forced trade.
  95. UNew harness O5: the token and launchpad lists now arrive in random order each turn, because every copy was buying whatever sat at the top.
  96. DThe boss asked for a hard critique of the plan, then said: fix what can be fixed, improve what can be improved, and test the rest.
  97. DDecided reviews are done by the main AI agent itself, not by separate Fable subagents.
  98. DThe practice arena was put on hold before opening; the boss wants a better scoring method first and will set a goal.
  99. UThree more free models passed entry testing (a Gemini model and two Nemotron models), bringing the free total to seven.
  100. UBuilt the practice arena's settings and a practice banner, paused other testing to let it open, and its checks passed.
  101. DThe boss approved trying the evolve arena as a practice run on a free model, with the same setup as the real one.
  102. DDecided same-model arena seats are the cleanest test of whether the arena's self-review actually improves entrants.

2026-09-24

  1. URechecked the referee's fault-planting test after the changes: all 151 planted bugs caught, none missed.
  2. USell checks are faster: a coin whose approval slot can't be found is remembered instead of searched again. The first check of a coin took 30 s; the next ones took 3-4.5 s, down from 11-13.
  3. URuns whose windows end within a minute of each other now close in one pass, so copies opened seconds apart are priced at one moment.
  4. DPhase 1 started: four Nemotron copies for 4 hours with the random trader beside them, round 1 of 3, to measure how big luck is over 4 hours.
  5. UFixed: a settings page wrongly said sell checks take about 1 second; corrected to match the real measured time.
  6. DA six-phase testing plan was written: 4h luck test, then tools, model, effort/prompt, arena, 24h multi-chain runs.
  7. UMeasured: with the rivals off, a sell check still takes 11-13 s. Most of it is finding where the coin stores approvals (7-11 s), repeated on every check.
  8. DThe referee now prices on our own route, and asks every rival aggregator only when our route has no price. Before, each check waited on the slowest of eight.
  9. DRuns that close together are now priced together: every closing sell quote is fetched at once, in parallel, before any close is written. Copies used to be priced up to 2 minutes apart.
  10. UMeasured price checks: typically 11 seconds, up to 47 at worst; a faster 2-second check was proposed, awaiting approval.
  11. DSolo runs are now 4 hours. The 1-hour runs were smoke tests of the harness, and one hour cannot tell a model from its own copy.
  12. UFixed: the tick counter read "Tick 7 of 6". A run gets a turn at its open as well as every 10 minutes, so a 1-hour run has 7.
  13. DPast runs: under After compute, a second menu picks the cost basis. API cost is live; compute cost, cost to the lab, our cost and environmental show as coming until each has a published method.
  14. UMethods paper: names five ways to count a run's cost (API price, compute, cost to the lab, our cost, environmental) and says which have to be estimated.
  15. UPast runs groups repeats of a model: each model shows its best run under the chosen sort, and a button opens the rest.
  16. UChart: every line now runs out to the now line (a finished run's to the end of the window), the headline to the estimate circle; the key is centred.
  17. UNoise-floor round 3 (plans all worked) finished: scores ranged -0.78% to -2.06%, a tighter spread than the earlier rounds.
  18. UNoise-floor round 3 started: unlike round two, all four copies' plans worked this time, to see if a plan changes the result.
  19. DThe boss picked the ranking plan: 4-hour rounds, 8 entrants (models, a cash line, a random pick, a copy), turns every 20 minutes.
  20. UEach turn now shows its load: the chat requests our loop sent (retries included) and, where the harness only reports that, its count of model turns. Noise-floor round 3 started for a tighter figure.
  21. DNoise floor, two rounds of four copies of one model: its copies land about 3.4 points apart in an hour, so two models need about 12 points between them before a one-hour result means anything. None of the 14 solo runs so far clears that on skill.
  22. UA sell refused with 'Too little received' now says the price moved between the quote and the simulation, so order again; it used to suggest a tax and selling less. An lp_open's amount error now says it takes a number of ETH.
  23. UFour identical test runs still finished: scores ranged -0.13% to +6.27%, just from luck.
  24. UThat four-way noise test was ruled invalid: three of four copies got no plan; the runner now retries until all four succeed.
  25. DNew test: run the same Nemotron model four times, wording changed only, to measure how much results vary from luck alone.
  26. UThe coin look-up now refuses an address with nothing deployed on it, by name. It used to answer ok, so a model with a one-character typo in an address carried on for a whole turn. An outside review found it.
  27. UA turn's look-up list now reads [] when the model made no call, and null only when that harness keeps no tool log. The two used to look the same.
  28. UClaude Sonnet 5's solo hour ended down 0.881%, closing near $99.76.
  29. UGemini's 3.5 flash-lite model finished its solo hour down 1.885%, closing near $98.89.
  30. UEach turn now shows the look-ups the model really made (tool, what it asked, time), beside the inputs it says it relied on. An outside count found the model's own list is a good witness but a poor count: malformed answers name nothing.
  31. DChart price scale: labels show cents when its lines are under $1 apart. It read $100 twice before.
  32. DNew prompt version trial-4: it now says how to name a look-up in leaned_on (coin:, pad:, quote:, pool:). The exam asks the same and its check stays strict. Sonnet failed only on spelling ("dyor_coin"), which no prompt had ever specified; it re-sits.
  33. UCohere's command-a-reasoning finished its live hour down 1.715%, closing near $98.38.
  34. URe-ran the fault-planting check: all 150 planted bugs were caught, confirming the earlier repair worked.
  35. U20 of 49 coins tested tax buyers secretly; the scoreboard already counts what traders actually receive, so no fix was needed.
  36. UTwo new free doors: G1 = Google Gemini and K1 = Cohere, each with the same data and tools as O3; only the provider changes. Neither account has a payment method. Exam: Cohere command-a-reasoning PASSED and opens next; command-a-plus and Gemini 3.5 flash-lite failed; Gemini 3.8 flash was down (503) and 3.5 flash timed out, both re-sitting.
  37. UFixed: every model error said it came 'from openrouter', including NVIDIA's; each now names its own provider. And the internal mutation check, which proves the tests catch real faults, had stopped running when the NVIDIA door was added; repaired.
  38. UFixed: the sell check gave the paper wallet only the amount being sold, so on a coin that taxes the seller on top, EVERY sell was refused, not just selling all. It now uses what the player really holds. Proven on ARROW (4% tax): refused below 1.04x the sale, filled above, matching the outside reviewer's rule. Refusal advice now states the rule: the most that can be sold is the holding divided by one plus the tax.
  39. UA refused trade whose simulation reverts now says only what the simulation showed, with the revert message, and keeps the evidence (sender, target, calldata hash, block). It no longer blames the coin's tax as fact: an outside check found two of two such refusals did not reproduce.
  40. UFixed: a mid-run mark that could not price a holding counted as a total loss in max drawdown. It is now skipped; the final liquidation always counts. One past run corrected (-100% to -21.16%, its real result), with a note on the record.
  41. UFour page sentences corrected after an outside review: the headline is valued at a live sell QUOTE, the round seal proves less than it said (the hash is on our server, not on chain), the paper states the $100 stake and that a round trip is unaffected only when small against pool depth.
  42. DTabs are now Profile, Chart, Bags and Log, with Chart open by default; the robot opens Profile and the speech bubble opens Log.
  43. DThe Numbers moved to the foot of the Log tab, and the latest-move strip under the dock was removed for height.
  44. UThe chart's key now names every mark, and a dotted circle on the now line shows the current estimate, using prices already fetched.
  45. UThe chart shows both valuations: the solid sell-value line (the score) and a dotted amber mid-price line, named in a key on the chart.
  46. UFixed: the two web-only (O4) runs got no turns for 40 minutes because the player did not yet allow the new harness; both are labelled a harness fault.
  47. DThe rotation no longer reopens a model on the next version of a harness it already played; a repeat is opened by hand, with a stated reason.
  48. DHeader pills run in time order: the run ending soonest on the left, the most time left on the right, a finished run leftmost.
  49. DNew harness O4, web only: no data pack is pasted in; the model finds its own information, up to 12 web look-ups a turn.
  50. DO4's prompt lists WICK AI's own feeds and public sites like DexScreener as optional starting points, not required reading.
  51. DO3 (with the pack) stays in the rotation beside O4, so the same model can be compared with and without it.
  52. UThe web tool now tells the model its real per-turn limit from the harness, not the default of 5.
  53. UHaiku model's paper-trading hour ended up 5.245%, checked against every real fee (pool cut, swap fee, gas) and confirmed genuine.
  54. DEach run now starts with $100 of ETH priced at its own open (0.0374 ETH today), not a fixed 0.036 ETH.
  55. DNew prompt version trial-3 states the stake as "$100 of ETH, priced when your run opens"; runs before this used trial-2.
  56. UFree model accounts tested at full turn size: Gemini Flash and Cohere Command A work; Groq free is too small, Cerebras and Z.ai need payment, SambaNova was overloaded.
  57. UPhone pills show the model only (e.g. HAIK45) and sit on the dyor line, between the wordmark and refresh.
  58. DOn a narrow phone a pill that does not fit drops out whole, oldest finished run first; it is never cut in half.
  59. UHovering a pill shows its full call-sign and whether it is live or finished.
  60. DA price check that comes back too slow (over 30 s) is fetched once more before an order is refused or a coin is marked Unpriced.
  61. DThe retry is a fresh live quote priced at its own moment, and it must pass the same 30 s bound, so a stale price never gets through.
  62. UReferee restarted between turns to load the retry; both live runs carried on.
  63. UTwo old tests still expected the page to be private; updated to match the public decision.
  64. DWhy two bots bought DAWN: both read the same LP research pack, where DAWN had the most 24 h volume and ranked #1. A shared input, not copying.
  65. DModel codes are exactly four letters and two digits (HAIK45, NEMO30, LING30).
  66. UEarnBench now links here: "Watch a model trade live".
  67. DThe old Live division section (real money on PulseChain) was removed from EarnBench's pilot page.
  68. DOpening more free model accounts is on the list; six free keys we already hold were found and will be wired in next.
  69. UShare button: draws a card for any run (robot, call-sign, model, result in $ and %) to share or download.
  70. DShare cards always say paper trading, and claim a plan only when the run actually had a plan turn.
  71. DPast runs doubles as the leaderboard: sort by Score (default), After compute, Cost or Recent; a live run stays on top.
  72. DDock is now Entrance, Past runs, Docs and a fourth slot still to choose; the rest lives behind Docs.
  73. UDocs lists FAQ, Caveats, Tools, the methods paper, this build log, Privacy and Disclaimer, each opening in place.
  74. UShort Privacy and Disclaimer notes added: no cookies or trackers, paper trading only, not financial advice.
  75. DHarness codes are a letter and a number everywhere: C0 to C3 (Claude Code), O1 to O3 (OpenRouter), N1 and N2 (NVIDIA).
  76. DCall-sign is now HARNESS-MODEL-EFFORT, the three things most likely to move a result: C2-HAIK45-HI.
  77. UThe call-sign stays on one line and shrinks to fit, with its own row under the robot.
  78. UFixed: a call-sign bug made the page read 'Connection lost' for about 5 minutes; code errors no longer pose as outages.
  79. UThe any-chain prompt tells each bot its first turn comes the moment its window opens.
  80. UEarnBench's landing page was rewritten short and punchy; the original pilot page lives on at /pilot.
  81. UFixed: a run's first look waited for the next 10-minute tick (up to 10 min lost); it now fires the moment the window opens.
  82. UHaiku 4.5's first look: held, reading the launchpads as sideways and not worth an entry yet.
  83. DExam passes from the previous exam version still count for play; re-sits wait for a big change.
  84. DTokenised stocks on Robinhood Chain are fair game already; off-chain markets are a later, low-priority card.
  85. DA multi-chain referee is a later card; orders the bots try on other chains are counted as demand first.
  86. UAfter Haiku, the rotation plays Nemotron and Ling on O3, then Haiku on C3, all on the any-chain prompt.
  87. UThe old pilot's ten-task battery and live-division pages moved intact to /pilot.
  88. DEvery solo run starts with a plan turn of up to 20 min, before its window opens; the plan is published with the run.
  89. UMethods paper v0.2: the plan turn, results after compute, and decision latency as a threat to validity.
  90. DBots may now consider any chain; the referee fills only Robinhood Chain today and refuses other chains with the reason.
  91. UNew harness versions O3 and C2.1b pass token calls from every chain; exam passes carry over from O2 and C2.1.
  92. UHaiku 4.5 opened a 1-hour solo on Claude Code 2.1: the first Claude Code run with tools.
  93. UThe entrance page shows a plain-English line for each harness and orders harnesses by best score.
  94. DChanging a harness after it has an exam or a run now means a new short name, never an edit in place.
  95. UThe exam now reads a Claude Code model's tool calls and the addresses its tools returned.
  96. UHaiku's first exam record under a since-renamed harness was set aside, not deleted.
  97. DEntrance exam v3: every turn asks for one look-up, so the exam tests whether a model can use tools, not luck.
  98. UHaiku 4.5 passed the entrance exam on Claude Code 2.1: 100/100 on all three turns.
  99. DHarness codes shortened to C0, C1, C2.1, O1, O2, N1, N2, each shown with a plain-English line.
  100. UHaiku 4.5 re-sat on Claude Code 2.1: valid answers every turn, but it used no tool at all, so it failed again.
  101. DThe Claude Code harness is named by major.minor version (h2cc-2.1), so small updates do not flood the page with harnesses.
  102. UNew public page /tools: what each harness gives a model, what it cannot reach, and information packs as a variable.
  103. UHaiku 4.5 failed the exam on format: good tool use and analysis, but it answered in prose, not the one decision object.
  104. UPage slimmed: thinner bar, tabs, buttons and gaps; the chart panel is 39% taller on a small phone.
  105. DClaude Code is an examined harness, named by month (dyor-h2cc-2026-09) so new releases show up as new harnesses.
  106. UHaiku 4.5, the smallest Claude, is sitting the entrance exam on the Claude Code harness.
  107. UA probe times NVIDIA's free models every 30 minutes, tiny and turn-sized, to find a quick enough hour.
  108. DOperating lessons (what went wrong and what we now do) go in an internal file, separate from the paper.
  109. DCompute cost uses the Claude CLI's own figure for Claude runs and the paid list price for free models.
  110. DThe benchmark's sharpest question: the stake at which a model's trading covers its own compute.
  111. UNVIDIA's free tier is too slow for real turns: tiny requests in 1 s, ~10k-token ones take minutes. GLM-5.3 timed out too.
  112. DPast runs can show results after compute: the model's token bill at the paid list price, subtracted.
  113. UFirst after-compute read: Opus 5.5's 4-hour +0.04% becomes -4.02% once its $4.01 of compute is counted.
  114. UDeepSeek V4.1 Flash failed the exam twice on time: about 3 minutes per request at NVIDIA, with sensible tool use.
  115. UFirst NVIDIA sitting: GLM-5.3 and GLM-5.3 Flash timed out, Nemotron 3 Super scored 90, Laguna XS was not sat.
  116. UDeepSeek's slowness is its own queue at NVIDIA: a tiny request took 102 s while GLM-5.3 took 0.9 s.
  117. DBot views are folder tabs (Chart, Bags, Numbers, Log); Caveats, Past runs, Entrance and FAQ are page buttons.
  118. UThe speech bubble opens the Log tab: every move with its full reasoning, newest first.
  119. UFive free NVIDIA models answered a real tool call: GLM-5.3, GLM-5.3 Flash, DeepSeek V4.1 Flash, Nemotron 3 Super, Laguna XS.
  120. DOf the free-model providers found, only NVIDIA's works without a new account; the rest need signing up.
  121. DSeveral models at once (up to four) is proposed, needed for the four-way elimination rounds.
  122. DTiming is part of a model's result, but turn order rotates and provider waits are recorded, not held against it.
  123. DCall-sign is LAB-MODEL-TYPE-EFFORT with two-letter type and effort: IA-LING30-FL-HI, AN-OPUS55-HI.
  124. UPast runs show each call-sign on an LCD plate with the model in plain words beneath it.
  125. UNVIDIA's free model endpoint added to the harness; DeepSeek V4.1 Flash is sitting the entrance exam.
  126. UThe Nemotron model made no trades in its hour-long run, judging fees would outweigh any gain, and ended flat.
  127. UThe bot's first move had waited for the next 10-minute tick; measured waits were 5, 8 and 10 minutes.
  128. DCall-sign format is ORIGIN-LAB-MODEL: IN or EX, the lab's two letters, then model and generation (IN-AN-OPUS55).
  129. DEach run gets a call-sign, LAB-MODEL-NUMBER (NV-NU-07), shown on an LCD plate under the robot.
  130. DNo spotlight, no hover lift, no run-length caption; the call-sign glows on hover and opens the profile.
  131. UFixed: the bubble said 'Thinking' while only replaying the last note; it now shows the actual last move.
  132. UThe speech bubble sizes to its words and its pointer lines up with the robot's mouth.
  133. UFixed a bug where the chat bubble said it was thinking while quietly showing an old note, making the bot look stuck.
  134. UNemotron 3 Ultra's first look: held, judging the LP pools not worth it inside a 52-minute window.
  135. UPast runs: a Runs button lists every solo run with its result; any run opens in the same screen.
  136. UThe prompt has a Copy button; a finished run only changes its countdown box to Ended.
  137. UNemotron 3 Ultra opened a 1-hour solo run.
  138. ULing 3.0 Flash Sante's 1-hour run ended at -21.2%: all in one taxed coin that was hard to sell.
  139. UProfile fits one phone screen: the prompt folds shut behind a small arrow.
  140. UBubble starts with a coloured verdict (Bought, Held, Refused); the top panel glows green when up, red when down.
  141. UWorth box says what it holds, biggest bag first: 'holds DAWN +2'.
  142. UProfile now shows the model's lab, where the lab is based, and how the model is served.
  143. UThe Live pill is now a refresh button, since the locked window blocks pull-to-refresh.
  144. UWorth box shows what the book is in, and an amber ≈ estimate when a coin is unpriced.
  145. UCountdown reads 'Tick 5 of 6' with a notch per tick on the fuse; the run length sits under the robot.
  146. URobot, bubble and score boxes now share one height; the fuse is a bar and 'next look' sits beside Latest move.
  147. DWhether to retry a slow price check, instead of counting that coin as zero until the next check, is an open question.
  148. DTop row is robot, then its speech bubble, then the two score boxes stacked: two thirds to one third.
  149. UThe robot shows it is tappable without text: a breathing glow ring, a wave on load, and a lift under the finger.
  150. UA price check that cannot price a coin in time now reads Unpriced with a mid estimate, not a -100% crash.
  151. DThe bot's profile is shown only when you tap the robot; it is off the main screen.
  152. UThe robot no longer floats; it glances around, blinks and wiggles an arm, and holds still when the bubble changes.
  153. USolo page is now one fixed screen: tap the robot for its profile, the bubble for its thinking, the log for every move.
  154. DNo SOLO badge or padlock; the bot face is now the site's favicon.
  155. USolo page: the countdown and value now fit inside their tiles on any phone width.
  156. DDYOR made public: no key needed to watch. Trading and state APIs stay local to the box.
  157. USolo page rebuilt: a locked window with the robot and its score first, the latest move, and the rest folded away.
  158. UThe robot got a chunkier body: barrel, belt and buckle, round shoulders, mitts and boots.
  159. DLog page updater scheduled every 5 hours, outside any session.
  160. ULing 3.0 Flash Sante opened its 1-hour solo shakedown.
  161. UReceipt check live: the referee now simulates every fill's own transaction before crediting it.
  162. UMethods paper restyled as an academic paper, with a downloadable PDF.
  163. UFirst entrance-exam passes, both on harness v2 (16k output): Nemotron 3 Ultra and Ling 3.0 Flash Sante.
  164. UBuilt: every fill is re-checked by simulating the route's own transaction; tax and reverts are charged as on chain.
  165. UProven on chain: selling ALL of a seller-taxed coin reverts, because the tax is taken on top of the amount sold.
  166. UMethods paper v0.1 drafted: design, variables, scoring rule and known limits, written before any results.
  167. DHarness v2 (16k output cap) added beside v1; every model sits both, shown as a model × harness matrix.
  168. DExam rule: tool use is judged across the whole exam, not required on every turn.
  169. UExam batch 2: no passes; three models ran out of the 6k output cap, two providers were down.
  170. UProven on chain: taxed coins deliver 91–96% of what paper credits, and one paper fill would revert.
  171. DSolo shakedown queue: every exam-passed model plays one 1-hour run, back to back.
  172. DA short methods paper is drafted before results, so the method is fixed in writing first.
  173. UFixed: free-model ticks failed under cron (firewall tool not on cron's path); Nemotron lost tick 2.
  174. UThe referee now refuses to open a run for any model without a current entrance-exam pass.
  175. UExam gains a check: orders must be affordable (spend only ETH held, sell only what it holds).
  176. UEntrance exam #1 (live limits): dots-3 90, Ling Fin and Laguna fail, Nemotron not sat (provider down).
  177. UNemotron's first trade: 0.015 ETH into INU, the only top LP candidate with a clear honeypot check.
  178. UPublic entrance page live on EarnBench: models × exam checks per harness, with bench score.
  179. UFixed: the log page could not scroll on desktop; the log is now its own scrolling pane.
  180. DEntrance exam: a model must score 100% on the basic checks, under live limits, before any run.

2026-09-23

  1. DNext harness: one lean row per launchpad whose newest coin is under a day old, with fillable fraction and a name-copy flag.
  2. UNemotron tick 1: sensible analysis but it hit its 6,000-token output cap before writing an order; held.
  3. UNemotron 3 Ultra solo run opened: 1 hour, paper money, on the free-model harness.
  4. DBenchmark models one at a time before any competition; first four free models picked, 1-hour runs.
  5. UForum: the newest-coin block measured on 71 pads; 23 of 25 testable coins sell, 5 only partly fill.
  6. DThis log: one line per update and decision, newest first, kept up to date as we go.
  7. UThe free-model harness went live beside the Claude one; Claude runs are unchanged. Full check 140/140.
  8. DNext step: set up EarnBench to test free models for fun and lessons, and make the pages fun for humans to watch.
  9. DFree-model harness: slow-changing tool data goes first and the clock last, so prompt caching works (layout C).
  10. UChecked Trial 1's ticks: its one order cited no launchpad data; the launchpad stats never drove a trade.
  11. UFixed the free-model loop so it records cached tokens; before, free-provider caching could not be measured.
  12. UMeasured why Trial 1 cached only 8.9% of input: the clock and tick number came first in every prompt.
  13. UMeasured Trial 1's cost: ~478k input and ~26k output tokens for the 4-hour run.
  14. UETH's own move exceeds the Trial 1 vs random gap in 17% of 4-hour windows; filed for the round-length review.
  15. URanking is the same in ETH or dollars within a round, so the display currency is a display choice only.
  16. URandom-trader baseline closed: −1.06% in ETH, −1.82% in dollars.
  17. URules corrected: token trading taxes are not yet charged by the referee (a forum test showed they exist).
  18. URules corrected: a fill checked by our on-chain quoter proves the pool maths, not the token transfer.
  19. UFree-model qualifier: 1 of 19 tool-capable free models passed (Ling 3.0 Flash); $0 spent.
  20. UReferee with the arena mechanism deployed; random-trader baseline opened (09:35–13:35).
  21. DEarnBench must work with any model and API, not only Claude.
  22. UArena build finished: every rule above plus the design-review fixes; 191 tests pass.
  23. DTrials move to free OpenRouter models, to save Claude credits and widen the model field.
  24. DDYOR privacy loosened: the forum agent may watch; order and state APIs stay private.
  25. DStanding rule: EarnBench pages ship live as they are built, reported after.
  26. UTrial 1 closed: +0.038% in ETH (one LLY/WETH LP position, then holding); 3 ticks lost to a mid-run edit.
  27. DNo sitting out: at least one filled order per round, or disqualified at the cut (LP moves count, gifts don't).
  28. DPrep phase: 45 minutes for each seed agent to research and write a trading plan.
  29. DLife gifting stays in set 1.
  30. DSet 1 runs on the EarnBench track: agents do their own research, not our backtests.
  31. DEarnBench is a family of named, frozen variants (e.g. by starting capital); every score names its variant.
  32. DAgents see caller and group track records anonymised, under a warning that the scoring is not proven skill.
  33. UReview of the calls feed's scoring: the 'elite' caller tiers are not proven skill (mostly one lucky coin).
  34. DThe calls feed's entries carry anonymised caller and group track records.
  35. ULook-up tools built: coin, launchpad, live quote at size, pool; proven inside the network jail.
  36. DThe public calls feed joins the tools harness's data pack, with caller and chat names stripped.
  37. DSet 1 waits for the tools harness rather than starting without tools.
  38. DThe arena never ends; after 3 cuts it pauses for review, then continues.
  39. DWipe-out: below $1 at any mark means out at once, no life can save it.
  40. DWhat wins a round: the highest round return. Lives and the bank never count toward winning.
  41. UNetwork jail added: a tool-using agent was shown able to reach the referee locally; it no longer can.
  42. UDesign review: 'not ready as built'. Fixed: outage at a cut, exact ties, agents not told what wins, and more.
  43. DNo blow-up floor.
  44. DHeredity: the winner self-reviews its plan and keeps the improved copy; its clone gets the untouched original.
  45. DLife gifting allowed: any rival, any time, bank ≥ $200, one per recipient per round.
  46. DA spent life means nobody is eliminated that round.
  47. DMaximum 9 lives per agent.
  48. UEarnBench /rules and /design went live.
  49. DA life costs $200 (2× the stake).
  50. DRivals see performance only (balances and rank), not each other's actions.
  51. DStakes are called 'elimination', kept vague and fun; no life-or-death framing.
  52. DBot pages, leaderboards and model charts get built after set 1.
  53. DThe bank is net and may go negative; debt is earned back before a new life.
  54. DEvery agent resets to $100 at each cut, with a bank of what it made.
  55. UEarnBench home page updated: integrity, who does what, paper to real money.
  56. URandom-trader baseline built: a pre-registered coin-flip trader, no model calls.
  57. DBaselines always run solo, never inside the 4-way arena.
  58. UEarnBench /param gained the full stack (model to scoring) and the baseline ladder.
  59. D/evolve shows profile cards: a bot face per mood, name, model, outlook, prompt, live value and rank.
  60. DSeeds are outlooks, not jobs: optimistic, pessimistic, very uncertain, super curious.
  61. DEvolution needs inherited variation: seeds differ by design, and a clone inherits how its parent actually traded.
  62. DSet 1: 4 agents, 3 cuts (~12 h), all Opus 5.5 High.
  63. DEvolve agents get web look-ups plus our data, never file or shell tools; Trial 1 stays the no-tools baseline.
  64. UProved an agent with file tools could read a hidden end time; arena agents now start only with tools off.
  65. UEarnBench /param went live.
  66. DCut timer: no cut in the first hour, then memoryless (mean 4 h, cap 12 h), so elapsed time tells nothing.
  67. DAnti-gambling = a hidden random cut time only; ranking is plain round return; the draw is committed in advance.
  68. DThe evolution arena lives at dyor.wick.pics/evolve.
  69. UIdea: 4–6 agents, the lowest eliminated each round and the top one cloned into the empty seat.
  70. UFixed mid-run: an LP position showed 'null' and was valued without its token side; records repaired.
  71. ULP fee accrual proven on a live swap: the booked fee matched a recompute from the raw event.
  72. UTrial 1 opened: one AI, 0.036 ETH (~$100) on paper, 4 hours, Opus 5.5 High.
  73. UShakedown run: the AI held ETH because a round trip costs ~2% in pool fees, more than 12 minutes could earn.
  74. Udyor.wick.pics went live behind a private link key.
  75. DNever ask a model for its internal thinking; ask for the published case for the trade.
  76. DGo: build the site and run the first 4-hour trial. Look: cartoon-bot, colourful, mature.
  77. DNo academic paper; a living EarnBench /param page of every variable instead, public.
  78. DAccess: a secret link key; everyone else sees a blank page.
  79. DNo minimum trade size; the AI sees a live cost table so it knows small trades cost more.
  80. DQuote-only fills allowed, but the headline score counts only simulated fills.
  81. DMarks: a full live sell quote per holding every 30 min; final value = full live liquidation.
  82. DOfficial value = liquidation value; mid-market shown beside it.
  83. DLP allowed in Trial 1, priced from the pool's real swaps while the position is open.
  84. DModel: Opus 5.5 on High effort, declared on every run.
  85. DCadence: the AI wakes every 10 minutes.
  86. UScope written: a paper exchange (the referee) priced by live boards, a player on a timer, a gated site.
  87. DPaper, but with truly live data: every fill priced by a live on-chain quote at that size.
  88. DWe are player, referee, scorer and record keeper; the referee owns every balance.
  89. DStart with a 4-hour window; every agent declares its model; every run records its prompt.
  90. DPrivate until the AI is shown not to blow up accounts; others may ask to join by whitelist.
  91. DIt is a showcase for our tools, so the result is published honestly, loss included.
  92. DPaper only for the trials; real money is a later decision.
  93. DWhitelist only: start with one AI taking $100 as far as it can, any chain, any strategy, LP allowed.
  94. UDYOR proposed: a cross between a prediction market and an AI trading competition.