{"schema_version":2,"source_url":"https://mimo.xiaomi.com/rl/","updated_at":"2026-09-17T05:41:17+00:00","monitor":{"status":"ok","last_attempt_at":"2026-09-17T05:41:17+00:00","last_success_at":"2026-09-17T05:18:10+00:00","cadence_minutes":60,"message":"本轮复核仍使用 05:18 左右的同一份新鲜官方归档；没有新的采集或 DeepSWE 结果，暂无明显新结论。"},"analysis":{"headline":"暂无明显新结论，等待下一份官方采集","summary":"本轮可用的最新官方归档仍是 05:18 左右那一份，没有新增训练轮次快照或 DeepSWE 结果。上一轮已确认的事实不变：Pro 已完成 13 轮、Flash 重跑后已完成 16 轮；外部测验仍只对应 Pro 第 10 轮、Flash 第 12 轮。","answers":{"progress":"暂无新采集：仍是 Pro 14 · Flash 17","training":"暂无新读数，保留上一份观察","benchmark":"没有新测验；旧 checkpoint 不算新结果"},"facts":["本轮复核的 source-capture.json captured_at 仍为 2026-09-17T05:17:59+00:00，与上一份已核验快照属于同一轮证据，没有生成新的训练快照。","上一份已核验状态仍为：Pro 已完成 13 轮、正在第 14 轮；Flash 已完成重跑后的第 16 轮、正在第 17 轮。","DeepSWE 仍为 Pro 第 10 轮 63.72、Flash 第 12 轮 60.77；本轮没有新增 checkpoint 结果。"],"interpretation":"短时间没有新的归档，只能说明本轮没有新增可验证证据，不能解释为训练停止、能力停滞或指标没有变化。","limits":"本轮没有新的 source-capture，因此不能推断 05:18 之后的训练阶段、练习评分、成本或其他指标。Pro 的阶段冲突仍沿用上一份快照的 unknown 处理。","next_watch":"等待下一次浏览器采集；优先看 Pro 是否完成第 14 轮、Flash 重跑后的第 17 轮是否完成，以及是否出现 Pro 第 10 轮之后或 Flash 重跑后的新 DeepSWE 结果。","evidence_snapshot_ids":["verified-20260917T051810Z"]},"backfills":[{"id":"official-history-20260917T051810Z","observed_at":"2026-09-17T05:18:10+00:00","verification":"verified","source_url":"https://mimo.xiaomi.com/rl/","evidence_url":"https://github.com/jbang2004/mimo/blob/daa8a8295bbd464b89cfdbb5570a3f9873563e90/source-capture.json","note":"首次导入官方历史序列，不是本网站此前逐小时采集的记录。Flash 使用官方回退后保留的 1–16 轮序列；旧版 16、17 轮已被重跑替代，不参与这条趋势。","models":{"flash":{"run_id":"mimo-v2.6-flash@1789485389","training_series_id":"flash-official-retained-after-1789610803","steps":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16],"recorded_at":["2026-09-15T17:07:10+00:00","2026-09-15T18:28:52+00:00","2026-09-15T22:20:35+00:00","2026-09-16T00:12:38+00:00","2026-09-16T01:46:53+00:00","2026-09-16T03:28:46+00:00","2026-09-16T05:14:37+00:00","2026-09-16T06:51:58+00:00","2026-09-16T08:35:19+00:00","2026-09-16T10:18:58+00:00","2026-09-16T12:16:20+00:00","2026-09-16T14:13:17+00:00","2026-09-16T16:12:01+00:00","2026-09-16T18:28:22+00:00","2026-09-16T20:34:17+00:00","2026-09-17T05:02:07+00:00"],"metrics":{"dynsam":[0.513693,0.496314,0.542596,0.50997,0.547289,0.538274,0.540552,0.560746,0.572833,0.569299,0.558439,0.576982,0.575977,0.583968,0.596286,0.594958],"reward":[0.516717,0.531547,0.548014,0.540377,0.535109,0.54286,0.531523,0.563526,0.559313,0.563913,0.567065,0.564307,0.561624,0.562557,0.563659,0.568556],"entropy":[0.413307,0.410635,0.412081,0.418074,0.408574,0.412859,0.417186,0.414308,0.412714,0.416586,0.422298,0.424936,0.433968,0.440154,0.431725,0.44111],"pg_loss":[0.000840813,0.00164193,0.00240887,0.0031375,0.00114159,0.00601586,0.00410864,0.00172965,0.00490317,0.00280688,0.00576059,0.00668219,0.00393784,0.00733537,0.0112836,0.00692291],"grad_norm":[0.00603667,0.00647849,0.00591034,0.00595898,0.00588154,0.00589539,0.00562257,0.00649302,0.00577247,0.00552962,0.00602531,0.00556255,0.00581106,0.00569057,0.00550012,0.00905447],"kl":[0.00286778,0.0041279,0.00460891,0.00529593,0.00740857,0.00880669,0.0091408,0.00934374,0.00926197,0.00953717,0.00982649,0.00910489,0.00883756,0.0091613,0.00832682,0.0032674],"context_tokens":[71633.5,78715.1,77525.6,87621.6,90422.1,92158.8,97512.6,93838.9,93496.9,96030.7,103002.0,103351.0,103978.0,112221.0,107496.0,98849.5],"turns":[47.3148,56.0742,49.9217,63.0969,52.4471,57.6547,59.3484,53.6523,57.3909,57.7261,60.6522,62.1462,60.6645,64.5941,57.0229,54.1008],"tokens_step":[1789670000.0,1963050000.0,1938190000.0,2189550000.0,2259170000.0,2300440000.0,2438180000.0,2345010000.0,2330290000.0,2401220000.0,2574190000.0,2585290000.0,2599210000.0,2805300000.0,2685540000.0,2418430000.0]},"benchmark":{"name":"DeepSWE","version":"1.1","aggregation":"avg@3","harness":"mini-swe-agent","unit":"score","protocol_id":"DeepSWE-v1.1|mini-swe-agent|avg@3|official-num2","observed_at":"2026-09-17T05:18:10+00:00","reported_at":null,"carried_forward":false,"steps":[1,2,3,4,5,6,8,10,11,12],"values":[48.67,53.1,56.78,54.03,57.23,54.57,57.08,60.18,54.13,60.77]}},"pro":{"run_id":"mimo-v2.6-pro@1789468339","training_series_id":"pro-official-retained-v1","steps":[1,2,3,4,5,6,7,8,9,10,11,12,13],"recorded_at":["2026-09-15T17:21:17+00:00","2026-09-15T21:11:21+00:00","2026-09-16T02:36:34+00:00","2026-09-16T04:30:51+00:00","2026-09-16T06:26:50+00:00","2026-09-16T08:22:39+00:00","2026-09-16T10:33:01+00:00","2026-09-16T12:42:01+00:00","2026-09-16T15:06:38+00:00","2026-09-16T17:13:08+00:00","2026-09-16T22:48:21+00:00","2026-09-17T01:19:20+00:00","2026-09-17T03:56:28+00:00"],"metrics":{"dynsam":[0.564662,0.555333,0.566948,0.575625,0.569944,0.585981,0.588339,0.60271,0.603837,0.590406,0.614156,0.608784,0.623849],"reward":[0.552211,0.555953,0.563102,0.561784,0.567325,0.566282,0.573527,0.578199,0.585502,0.582035,0.587047,0.581967,0.587806],"entropy":[0.395265,0.379237,0.38227,0.387387,0.379743,0.385975,0.395814,0.392417,0.387348,0.390081,0.392639,0.392207,0.398121],"pg_loss":[0.00411032,0.00236891,0.00210573,0.00251458,0.00202267,0.0042952,0.00419337,0.00550789,0.00469692,0.00616015,0.0029433,0.00617481,0.0073807],"grad_norm":[0.00533389,0.00585464,0.00557954,0.00567242,0.00580347,0.00606722,0.00545478,0.00622983,0.00872634,0.00609214,0.00582598,0.00543666,0.00485642],"kl":[0.0021894,0.00312041,0.00308185,0.00495296,0.00791114,0.00898885,0.00970035,0.00959044,0.00980071,0.00879815,0.00529786,0.00655812,0.00852191],"context_tokens":[72177.6,69231.9,74191.9,80519.2,80385.5,84799.1,89705.7,91416.5,93699.9,89040.3,86591.4,93076.6,99593.8],"turns":[47.4689,43.4,45.3328,53.0981,53.5462,52.1122,57.7969,55.1437,60.7042,55.0407,46.917,55.0367,64.8149],"tokens_step":[1800520000.0,1730640000.0,1855440000.0,2010940000.0,2009690000.0,2118430000.0,2246470000.0,2266550000.0,2327390000.0,2224380000.0,2164150000.0,2328100000.0,2491010000.0]},"benchmark":{"name":"DeepSWE","version":"1.1","aggregation":"avg@3","harness":"mini-swe-agent","unit":"score","protocol_id":"DeepSWE-v1.1|mini-swe-agent|avg@3|official-num2","observed_at":"2026-09-17T05:18:10+00:00","reported_at":null,"carried_forward":false,"steps":[1,2,3,4,5,6,7,8,10],"values":[58.41,56.25,58.41,60.47,59.59,58.55,57.44,62.24,63.72]}}}}],"snapshots":[{"id":"verified-20260917T051810Z","observed_at":"2026-09-17T05:18:10+00:00","source_updated_at":null,"verification":"verified","source_url":"https://mimo.xiaomi.com/rl/","evidence_url":"https://github.com/jbang2004/mimo/blob/daa8a8295bbd464b89cfdbb5570a3f9873563e90/source-capture.json","models":{"flash":{"run_id":"mimo-v2.6-flash@1789485389","training_series_id":"flash-official-retained-after-1789610803","phase":"rollout","phase_note":"已从第 15 轮重新开始，完成了重跑的第 16 轮；当前第 17 轮正在收集新的尝试。","current_step":17,"completed_steps":16,"metrics_step":16,"metrics_recorded_at":"2026-09-17T05:02:07+00:00","status_recorded_at":"2026-09-17T05:18:01+00:00","metrics":{"dynsam":0.594958,"reward":0.568556,"entropy":0.44111,"pg_loss":0.00692291,"grad_norm":0.00905447,"kl":0.0032674,"context_tokens":98849.5,"turns":54.1008,"tokens_step":2418430000.0,"passrate_zero":null,"passrate_one":null,"infra_error_rate":null,"step_seconds":null,"rollout_seconds":null,"train_seconds":null,"total_tokens":37622730000.0,"cost_usd":390825.9079607856,"samples":401408.0,"prompts_per_step":1568.0,"attempts_per_prompt":16,"accepted":1168,"evaluated":1154},"approximate_metrics":[],"benchmark":{"name":"DeepSWE","version":"1.1","aggregation":"avg@3","harness":"mini-swe-agent","unit":"score","checkpoint_step":12,"value":60.77,"protocol_id":"DeepSWE-v1.1|mini-swe-agent|avg@3|official-num2","observed_at":"2026-09-17T05:18:10+00:00","reported_at":null,"carried_forward":false}},"pro":{"run_id":"mimo-v2.6-pro@1789468339","training_series_id":"pro-official-retained-v1","phase":"unknown","phase_note":"页面显示 training，状态接口显示 rollout；阶段信号不一致，暂不高亮。已完成 13 轮、正在第 14 轮是一致的。","current_step":14,"completed_steps":13,"metrics_step":13,"metrics_recorded_at":"2026-09-17T03:56:28+00:00","status_recorded_at":"2026-09-17T05:17:59+00:00","metrics":{"dynsam":0.623849,"reward":0.587806,"entropy":0.398121,"pg_loss":0.0073807,"grad_norm":0.00485642,"kl":0.00852191,"context_tokens":99593.8,"turns":64.8149,"tokens_step":2491010000.0,"passrate_zero":null,"passrate_one":null,"infra_error_rate":null,"step_seconds":null,"rollout_seconds":null,"train_seconds":null,"total_tokens":27573710000.0,"cost_usd":878998.325364089,"samples":326144.0,"prompts_per_step":1568.0,"attempts_per_prompt":16,"accepted":2304,"evaluated":2172},"approximate_metrics":[],"benchmark":{"name":"DeepSWE","version":"1.1","aggregation":"avg@3","harness":"mini-swe-agent","unit":"score","checkpoint_step":10,"value":63.72,"protocol_id":"DeepSWE-v1.1|mini-swe-agent|avg@3|official-num2","observed_at":"2026-09-17T05:18:10+00:00","reported_at":null,"carried_forward":false}}},"raw_evidence":{"status":"/rl/api/status?run=pro|flash","series":"/rl/api/series?run=pro|flash","benchmarks":"/rl/api/benchmarks","notices":[{"id":"n-b60f90","t":1789612058.9304461,"text":"we restarted the flash run from step 15. reason: a type of infra error on one of datasets was not correctly detected over the past ~3 hours.","run":null},{"id":"n-92030c","t":1789589280.3782058,"text":"the mimo-v2.6-pro run is restarting due to a vram issue on one node.","run":null},{"id":"n-5ef2b1","t":1789581497.4956336,"text":"we have updated the latest deepswe results for flash step 12 & pro step 8. we will keep posting as the offline evaluation results come out.","run":null}],"descriptions":{"dynsam/avg@n":"mean pass rate: for each prompt sampled this step, the fraction of its n attempts that succeed, averaged over prompts","dynsam/avg@n_no_infra":"avg@n with attempts that failed for infrastructure reasons excluded","critic/rewards/mean":"mean reward over trajectories trained on this step","actor/entropy_loss":"mean per-token entropy of the policy","actor/pg_loss":"clipped policy-gradient objective","actor/grad_norm":"global gradient norm before clipping","train_infer_diff/new_infer/kl":"KL between inference-engine and trainer log-probs on the same tokens","ctx_response_length/mean":"tokens generated per trajectory","dynsam/agg_turn/mean":"agent turns per trajectory","perf/total_num_tokens":"tokens trained on this step","timing_s/step":"wall-clock of the whole step","timing_s/outer_gen":"wall-clock of rollout generation","timing_s/trainer_ops":"wall-clock of the trainer","dynsam/passrate/zero":"share of prompts where no attempt succeeded","dynsam/passrate/one":"share of prompts where every attempt succeeded","dynsam/infra_error/seq_rate":"share of sequences lost to infrastructure failures","env/active":"sandbox environments in flight","partial/avg_staleness":"policy versions between sampling and training, on average","dynsam/num_measurable":"prompts with a measurable pass rate this step; per data source under dynsam/<source>/num_measurable","dynsam/passrate/hist9_ratio":"share of prompts by pass rate, in nine bins from none solved to all solved","train/harness/*/training/rollouts":"rollouts in the training batch, per agent harness","ctx_total_length/mean":"total context length per trajectory (prompt + response), in tokens"}}}],"observations":[{"id":"launch-verified-history","time":"2026-09-17T05:22:13+00:00","kind":"data_quality","title":"首次接入官方历史曲线，每个读数都有出处","summary":"已导入两款模型的训练历史与 DeepSWE 历史。曲线来自官方接口回溯，不冒充本网站过去已经执行过的逐小时监控。","facts":["历史序列保留原始轮次和训练记录时间；抓取发生于 2026-09-17T05:18:10+00:00","未取得的耗时、错误率等指标留空，不沿用旧版近似值。"],"interpretation":"从今天起可以在相同记录口径下看真实前后变化。","limits":"官方公开的是过程指标，不是完整训练配方；轨迹内容与模型权重不在本次采集范围内。","next_watch":"后续每小时追加真实观测，保留原始快照和变化依据。","evidence_snapshot_ids":["verified-20260917T051810Z"]},{"id":"benchmark-baseline-verified","time":"2026-09-17T05:18:10+00:00","kind":"benchmark","title":"编程测验已有提升，但“最新成绩”不是“当前模型成绩”","summary":"Pro：58.41 → 63.72（第 1 → 10 轮）；Flash：48.67 → 60.77（第 1 → 12 轮）。当前训练已超过这些测验对应的轮次。","facts":["官方状态记录：Pro 已完成 13 轮、正在第 14 轮；Flash 已完成重跑后的第 16 轮、正在第 17 轮。","Pro 已公布的 DeepSWE 成绩从第 1 轮的 58.41 到第 10 轮的 63.72，增加 5.31；Flash 从第 1 轮的 48.67 到第 12 轮的 60.77，增加 12.10。"],"interpretation":"这支持已测版本在该编程测验上有改善；不需要借助训练奖励来代替能力测试。","limits":"本次是历史结果首次核验，不代表这些结果刚刚发布。未把页面 num2 分数强行换算成每 100 次成功数。","next_watch":"观察新测验是否继续改善，而不是把成绩未更新判为能力停滞。","evidence_snapshot_ids":["verified-20260917T051810Z"]},{"id":"notice-n-b60f90","time":"2026-09-17T02:27:38+00:00","kind":"incident","title":"Flash 为什么“退回去”训练？不是因为分数低","summary":"官方公告：某类环境错误此前未被正确检测，Flash 因此从第 15 轮重新开始。这里将重跑替代的旧第 16、17 轮从有效曲线中分离。","facts":["we restarted the flash run from step 15. reason: a type of infra error on one of datasets was not correctly detected over the past ~3 hours.","状态事件记录包含一次 restart，以及标为 redo 的新第 16 轮。"],"interpretation":"模型失败与环境失败需要区分。错误检测会影响训练数据与评分的可信度，这比单看轮次增减更重要。","limits":"公告没有披露受影响的全部样本或损失细节，不能据此推算模型退步多少。","next_watch":"重跑后的测验、错误检测情况及后续官方说明。","evidence_snapshot_ids":["verified-20260917T051810Z"]},{"id":"routine-20260917T054117Z","time":"2026-09-17T05:41:17+00:00","kind":"routine","title":"暂无明显新结论：本轮仍是同一份 05:18 归档","summary":"本轮没有新的官方采集或 DeepSWE checkpoint，因此不追加训练快照，也不把旧数值刷新成新结果。","facts":["source-capture.json 的 captured_at 仍为 2026-09-17T05:17:59+00:00，capture_ok=true。","当前 data.json 已经包含与这份归档对应的 verified-20260917T051810Z 快照。"],"interpretation":"监控链路正常完成复核，但没有新的数据事件值得更新能力判断。","limits":"无法据此判断 05:18 之后训练是否又完成了新的轮次。","next_watch":"下一份新归档到来后再比较轮次、dynsam/reward 与新增外部测验。","evidence_snapshot_ids":["verified-20260917T051810Z"]}],"legacy":{"updated_at":"2026-09-17 11:58 CST","models":{"pro":{"phase":"Step 11 · training","step":11,"step_delta":"+1","dynsam":"0.590","dynsam_delta":"+0.026 vs Step 1","deepswe":"62.24","deepswe_delta":"checkpoint ~Step 8","tokens":"20.6B+","context":"~89K","cost":"$740K+","reward":"~0.582","zero":"~15.0%","one":"~21.3%","grad_norm":"~0.0061","kl":"~0.0088","turns":"~55.0","infra_error":"~0.61%","step_time":"~2h 06m"},"flash":{"phase":"Step 17 · rollout","step":17,"step_delta":"+1","dynsam":"0.603","dynsam_delta":"+0.089 vs Step 1","deepswe":"60.77","deepswe_delta":"checkpoint ~Step 12","tokens":"37.8B+","context":"~104K","cost":"$321K+","reward":"~0.581","zero":"~14.9%","one":"~21.8%","grad_norm":"~0.0055","kl":"~0.0083","turns":"~52.6","infra_error":"~0.56%","step_time":"~1h 50m"}},"analysis":{"signal":"中","headline":"训练仍在健康推进，Flash 的相对增益更明显","summary":"当前公开曲线没有显示 reward 坍塌、梯度异常或明显的 train/infer 偏离。Flash 的 dynsam 提升幅度显著大于 Pro，更像是起点较低后的快速追赶；但 Pro 的已公开 DeepSWE checkpoint 仍略高。真正值得等待的是下一批外部 benchmark：如果训练内 reward 继续涨而 DeepSWE 不再提升，才可能出现收益递减或过拟合训练分布。"},"history":[{"time":"2026-09-17 11:58 CST","title":"建立基线快照","text":"Pro 已进入 Step 11，Flash 已进入 Step 17。当前最明显的结构性现象是 Flash 训练内指标提升更快，但外部 DeepSWE checkpoint 仍未超过 Pro。"}]},"history":[{"time":"2026-09-17 11:58 CST","title":"建立基线快照","text":"Pro 已进入 Step 11，Flash 已进入 Step 17。当前最明显的结构性现象是 Flash 训练内指标提升更快，但外部 DeepSWE checkpoint 仍未超过 Pro。"}]}
