MONTH · 2026-10

Monthly Digest — 2026-10

155 unique stories across 31 days and 8 sources.

Hacker News(32)

  1. Pi Durable (earendil.com)
  2. Pi 1.0 (earendil.com)
  3. Git 3.0's upcoming SHA-256 default will be a costly mistake (blog.gitbutler.com)
  4. Clef: Open-source decision models, and new RL fine-tuning platform (blog.cloudflare.com)
  5. Zig v0.17.0 (ziglang.org)
  6. Apple Pass Designer (developer.apple.com)
  7. A 12-year sequence of telescope images of a star and four planets orbiting (bsky.app)
  8. Mike Tomlin spent 12 years building a Minecraft city (www.nytimes.com)
  9. Getting the most out of Opus 5.5 in Claude and Claude Code (claude.dev)
  10. Celebrating the 100th birthday of the kidney donated to him as a teenager (www.whec.com)
  11. ADHD, autism or complex trauma? [pdf] (www.cambridge.org)
  12. Hole Punch: Sling your spaceship around gravitational fields (notoriousbfg.com)
  13. Improper redaction reveals Google Data Center water and electricity usage (www.1011now.com)
  14. Turn off Apple Intelligence on macOS 27 and get its disk space back (github.com)
  15. A map of every lighthouse (mapped.earth)
  16. Car is a smartphone on wheels. Here's who's listening (automatictransmission.khoury.northeastern.edu)
  17. Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates (www.vals.ai)
  18. Plain text is still one of the best technologies we have (deadparrotbbs.com)
  19. Beam: Reflection's 501B open-weight model (reflection.ai)
  20. OpenAI "rogue" agent activities found on Wikimedia projects (diff.wikimedia.org)

GitHub Trending(19)

  1. DietrichGebert / ponytail
  2. mattpocock / skills
  3. NVIDIA / OpenShell
  4. firebase / firebase-ios-sdk
  5. Panniantong / Agent-Reach
  6. JuliusBrussee / caveman
  7. obra / superpowers
  8. pbakaus / impeccable
  9. affaan-m / ECC
  10. Effect-TS / effect
  11. tester-army / e2e
  12. coreyhaines31 / marketingskills
  13. thedotmack / claude-mem
  14. earthtojake / text-to-cad
  15. pingdotgg / t3code
  16. boykopovar / AnyPS5
  17. morluto / rea
  18. ayghri / i-have-adhd
  19. cathrynlavery / diagram-design

Product Hunt(32)

  1. Phare C1®

    A smoke alarm that detects fire, not toast.

  2. Buddy Drop

    Drop the files, get a live URL in seconds

  3. Starlie

    A fast, native Jira client for Mac

  4. Bracket

    The memory layer for your business

  5. Teachoo

    AI that helps you learn anything and everything.

  6. Gauth Unlimited Digital Canvas

    An AI tutor on an infinite whiteboard, not a chat thread

  7. Globestudio

    Open-source dotted maps and 3D globes for designers

  8. WMail

    Beautifully native iPhone app for Fastmail

  9. OTPfill

    Autofill OTP codes from your email on Mac

  10. una mano

    A familiar iPhone keyboard that moves to your thumb

  11. Kindle 2026

    A smaller, faster Kindle with a flush-front display

  12. Yubi

    Talk to your Mac and let Yubi do the typing

  13. Clair

    Answer Claude Code from your wrist

  14. CoreSpeed

    One MCP for everything your agents need: apps, memory, tools

  15. Gemini 4 Argon

    Google's frontier model for careful reasoning & complex work

  16. opensend.cc

    The open source email platform that runs on your server

  17. Spira Maxima

    Video model that turns script into viral social media videos

  18. Dots UI

    A React library for morphable particle interfaces

  19. FastRouter.ai

    Route requests to the right LLM for cost, latency & quality

  20. crosswalk

    A third place for people & their agents, starting with inbox

Hugging Face(25)

  1. False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

    Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.

  2. Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

    Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.

  3. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

    Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

  4. EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

    Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.

  5. OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

    Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.

  6. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.

  7. A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

    AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.

  8. E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

    Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.

  9. Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

    Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/

  10. RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

    A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 120 days, released as the conversation and four files derived from it, a profile, a persona, a chat ground truth and a question set, each citing the messages it rests on. Every chat label carries the reasoning trace that produced it, checked stage by stage against the conversation. Three findings follow. First, the past is rarely needed and far away. Pooled measures mislead: a recency window finds the required message for 95.9\% of probes and 2.2\% of those that need memory, and at the natural rate 96\% of the gain from supplying recorded evidence comes from messages that need none. Second, no detector we tried can tell when memory is needed on real messages, authored questions over the same histories leak the cue, and labeling the same messages as memories raises their use by ten to fourteen points. Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.

  11. Does Learning Protein Folding Generalize to Broader Reasoning?

    Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.

  12. MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

    Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

  13. World Action Modeling with Progressive Visual Planning

    World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in https://sii-ferenas.github.io/ProWAM-page.

  14. In-Distribution Forcing for Long Video Generation at Test Time

    Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the training horizon: nothing constrains the construction of KV entries during rollout, giving rise to the KV-provenance problem where cached entries themselves become out-of-distribution (OOD). To address this, we propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations. Its key mechanism, self-caching, prevents OOD KV entries at their source. Each chunk is cached without attending to prior KV entry, keeping the rolling window exactly in-distribution. Consequently, ID-Forcing seamlessly extends short-horizon models to minute-scale video generation. Extensive evaluations show that our method remains competitive on standard video generation benchmark while substantially outperforming prior work in mitigating drifting, as validated by both our drift metrics and a user study.

  15. LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

    LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.

  16. Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

    Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.

  17. Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

    Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.

  18. Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

    On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.

  19. TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

    Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.

  20. CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

    Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.

Techmeme(32)

  1. Source: some of the information that the three OpenAI employees allegedly mishandled pertained to OpenAI's infrastructure architecture (Rachel Metz/Bloomberg)

    Rachel Metz / Bloomberg : Source: some of the information that the three OpenAI employees allegedly mishandled pertained to OpenAI's infrastructure architecture —  OpenAI has parted ways with three employees for violating company policies on how private information should be handled, including for allegedly sharing it with an outside group.

  2. Lyft agrees to pay $272.5M to settle California claims that it mislabeled drivers as independent contractors rather than employees between 2016 and 2020 (Daniel Wiessner/Reuters)

    Daniel Wiessner / Reuters : Lyft agrees to pay $272.5M to settle California claims that it mislabeled drivers as independent contractors rather than employees between 2016 and 2020 —  Lyft (LYFT.O) has agreed to pay $272.5 million to settle claims by the state of California and three of its largest cities …

  3. OpenAI launches two ChatGPT shopping features: a virtual try-on tool for clothing and accessories, and a Favorites feature to save products to a user's Library (Sarah Perez/TechCrunch)

    Sarah Perez / TechCrunch : OpenAI launches two ChatGPT shopping features: a virtual try-on tool for clothing and accessories, and a Favorites feature to save products to a user's Library —  OpenAI is again experimenting with how its conversational AI assistant, ChatGPT, can help users as they shop online.

  4. DoorDash pulls support for a GOP bill that would have limited DC's ability to write its own tax laws after widespread calls for locals to boycott the service (Martin Austermuhle/The Washington Sun)

    Martin Austermuhle / The Washington Sun : DoorDash pulls support for a GOP bill that would have limited DC's ability to write its own tax laws after widespread calls for locals to boycott the service —  The company still wants the D.C. Council to make changes to a new delivery fee taking effect Thursday.  — Copy

  5. Apple releases an update for iPhone 18 Pro Max devices on AT&T to address cellular failures; units that already lost service require hardware replacement (Chance Miller/9to5Mac)

    Chance Miller / 9to5Mac : Apple releases an update for iPhone 18 Pro Max devices on AT&T to address cellular failures; units that already lost service require hardware replacement —  Apple has confirmed an issue causing a “small number” of iPhone 18 Pro Max devices on AT&T to lose cellular service.

  6. Meta announces Muse Gadgets, providing an open source ESP32 microchip firmware and a Linux SDK to let users bring Muse to their hardware, like Raspberry Pi 5 (Karissa Bell/Engadget)

    Karissa Bell / Engadget : Meta announces Muse Gadgets, providing an open source ESP32 microchip firmware and a Linux SDK to let users bring Muse to their hardware, like Raspberry Pi 5 —  The company created its own smart home device, the Muse Home Link  —  We still don't know too much about Meta's tamagotchi-like AI device …

  7. Meta says it is letting go of employees it hired from AI safety startup Virtue AI four months after they joined the company, citing clashing work styles (Ashley Gold/Semafor)

    Ashley Gold / Semafor : Meta says it is letting go of employees it hired from AI safety startup Virtue AI four months after they joined the company, citing clashing work styles —  Meta said it's letting go of employees it hired from AI safety startup Virtue AI four months after they joined the company in June …

  8. Memo: the US Army is creating an autonomous systems command, after Defense Secretary Pete Hegseth announced the Meridian and Agincourt robotic warfare projects (Colin Demarest/Axios)

    Colin Demarest / Axios : Memo: the US Army is creating an autonomous systems command, after Defense Secretary Pete Hegseth announced the Meridian and Agincourt robotic warfare projects —  The U.S. Army is creating a new autonomous systems command and appointing an acquisition executive to prioritize buying …

  9. David Robinson, ex-OpenAI safety and policy: SV lacks a safety-centric culture; labs must study other fields' safety approaches; time for trial and error's over (David Robinson/The Atlantic)

    David Robinson / The Atlantic : David Robinson, ex-OpenAI safety and policy: SV lacks a safety-centric culture; labs must study other fields' safety approaches; time for trial and error's over —  What I'm about to tell you has, I realize, become something of a cliché: I resigned this week from OpenAI.

  10. Sources: ShinyHunters member Saif al-Din Khader, aka "Rey," was detained in Jordan and is cooperating to identify other hackers involved in the FBI breach (Reuters)

    Reuters : Sources: ShinyHunters member Saif al-Din Khader, aka “Rey,” was detained in Jordan and is cooperating to identify other hackers involved in the FBI breach —  A key member of the ShinyHunters hacking group, which claims to have stolen data on every FBI employee …

  11. YouTube says it is adjusting its Shorts recommendations to prioritize "original content" and reduce the reach of channels re-uploading others' content (Andrew Romero/9to5Google)

    Andrew Romero / 9to5Google : YouTube says it is adjusting its Shorts recommendations to prioritize “original content” and reduce the reach of channels re-uploading others' content —  YouTube says it wants to ‘focus on originality’ in a push to surface better Shorts recommendations for users.

  12. OpenAI's DevDay 2026 announcements to turn ChatGPT into a place to discover, launch, and use software could potentially disrupt the traditional app store model (Sarah Perez/TechCrunch)

    Sarah Perez / TechCrunch : OpenAI's DevDay 2026 announcements to turn ChatGPT into a place to discover, launch, and use software could potentially disrupt the traditional app store model —  The focus of OpenAI's Dev Day on Tuesday may have been on its agentic assistants known as Dots, or its new AI models, but combined …

  13. Sam Altman says OpenAI and Anthropic still hold fundamentally different worldviews on AI regulation, arguing that AI's benefits justify accepting some risks (Politico)

    Politico : Sam Altman says OpenAI and Anthropic still hold fundamentally different worldviews on AI regulation, arguing that AI's benefits justify accepting some risks —  OpenAI CEO Sam Altman said there remains a fundamental difference in worldview between his company and rival Anthropic …

  14. Sources: Anthropic's stock match of employee charity gifts hit $660M+ in the six months through March, likely to reach billions post-IPO, diluting shareholders (Cory Weinberg/The Information)

    Cory Weinberg / The Information : Sources: Anthropic's stock match of employee charity gifts hit $660M+ in the six months through March, likely to reach billions post-IPO, diluting shareholders —  When Anthropic shared financial figures recently with prospective investors in its planned initial public offering …

  15. Sources: Schneider Electric is in advanced talks to buy US engineering software firm PTC for ~$20B, its largest acquisition; a deal could come as soon as Monday (Financial Times)

    Financial Times : Sources: Schneider Electric is in advanced talks to buy US engineering software firm PTC for ~$20B, its largest acquisition; a deal could come as soon as Monday —  Acquisition would be French conglomerate's largest and enhance its products focused on manufacturers

  16. Extracted system prompts show Meta's Muse compiles "a page for every person in the user's life", with facts, history, tips to improve relationships, and more (Wired)

    Wired : Extracted system prompts show Meta's Muse compiles “a page for every person in the user's life”, with facts, history, tips to improve relationships, and more —  Millions have downloaded Meta's AI agent Muse.  But getting it to do your bidding comes with privacy costs.

  17. SpaceX stock closes up 7.6% after Morgan Stanley called SPCX "cheap", reaching its highest level since mid-June and returning Elon Musk to trillionaire status (Lora Kolodny/CNBC)

    Lora Kolodny / CNBC : SpaceX stock closes up 7.6% after Morgan Stanley called SPCX “cheap”, reaching its highest level since mid-June and returning Elon Musk to trillionaire status —  SpaceX shares rose almost 8% on Monday, reaching their highest since mid-June, shortly after Elon Musk's company held its record initial public offering.

  18. TikTok debuts Shopping Assistant, a conversational AI agent that helps users find and buy products, and Buy Direct for one-click purchases from the For You feed (Aisha Malik/TechCrunch)

    Aisha Malik / TechCrunch : TikTok debuts Shopping Assistant, a conversational AI agent that helps users find and buy products, and Buy Direct for one-click purchases from the For You feed —  TikTok announced Monday that it's launching an AI shopping assistant and a new in-app checkout feature that lets users buy directly from brands.

  19. NY-based Reflection unveils Beam, an open model it says rivals Qwen3.8-Max on coding and agentic tasks and GLM 5.2 on reasoning, while using 3x-4× less compute (Semafor)

    Semafor : NY-based Reflection unveils Beam, an open model it says rivals Qwen3.8-Max on coding and agentic tasks and GLM 5.2 on reasoning, while using 3x-4× less compute —  Reflection AI, the startup billing itself as America's answer to open-source Chinese AI, is releasing its first model, called Beam.

  20. Sources: Meta and Microsoft are working to cut their employees' use of Claude; Meta employees using Claude Code have dropped to ~30K from ~60K earlier this year (The Information)

    The Information : Sources: Meta and Microsoft are working to cut their employees' use of Claude; Meta employees using Claude Code have dropped to ~30K from ~60K earlier this year —  Meta Platforms and Microsoft, two of Anthropic's biggest corporate customers, are working to cut their employees' use of Claude …

Solidot(15)

  1. PS5 模拟器的开发取得突破

    当前一代游戏机的模拟器通常需要较长时间才能成熟,但 PS5 的模拟器仅仅几个月时间就让许多 PS5 游戏能在 PC 平台上可玩。SharpEmu 从 5 月的极早 Alpha 阶段到现在具备加载真实游戏 eboot.bin 文件、执行原生 CPU 指令以及部分处理 GPU 相关功能的能力。已有 10 款游戏被标记为可玩,其中包括简单 2D 游戏如 Tetris Forever,也有复杂 3D 大作如 Astro Bot 和 Demon’s Souls 重制版。另一款 PS5 模拟器 KytyPS5 也于上周发布了首个公开版本,它是 PS4 模拟器 Kyty 的扩展版,能启动 2D 游戏及部分 3D 游戏,包括使用虚幻引擎 4/5、Unity 以及自研引擎开发的作品。有 134 款游戏标记为可玩,但很多存在严重 bug。

  2. 二手 CPU 导致玩家被 Riot 封禁

    一名玩家购买了一个二手 CPU Ryzen 7 5800X3D,结果发现无法启动 Riot 工作室旗下的多款游戏,每次启动游戏就被踢出,在联络了 Riot 的客服之后才知道该 CPU 被列入了封禁黑名单,因为其前任主人有作弊行为。Riot 工作室旗下的所有游戏都受到影响,其中包括了 Valorant、League of Legends、Teamfight Tactics、Legends of Runeterra、2XKO 等。Riot 使用了内核级反作弊系统 Vanguard,它深度嵌入在 Windows 内核模式中。暂时不清楚 Riot 是如何唯一标识 CPU 的,CPU 会向操作系统报告硬件 ID,但该硬件 ID 不是唯一标识符,只是告诉操作系统其 CPU 型号。它可能是通过 CPU 运行产生的独特指纹标记 CPU,每个 CPU 在运行时候都会有微小的差异。

  3. 新加坡推出面向公务员的约会软件 FirstDate

    为了提高生育率,新加坡试点推出了为公务员牵线搭桥的约会应用 FirstDate。新加坡的总和生育率已降至每名女性生育 0.87 个孩子,而十年前这一数字为 1.24。政府最近成立了一个专门研究生育率下降问题的工作组,预计该作组将在 2027 年初发表研究结果。FirstDate 面向 21-35 岁的单身人士,目前仅向公务员开放,申请截止日期为 10 月 5 日。用户无需浏览海量的个人资料,而是填写一份关于兴趣、习惯、价值观和偏好的问卷。FirstDate 会在每个周期内(即双方互相接受并预计见面所需的时间)为用户推送一个匹配对象。每项匹配结果都包含匹配度评分、对方的简介和一段个人留言。该服务使用了 Gale-Shapley 稳定婚姻算法。有公务员认为该应用的一大优势可能是被诈骗的可能性较低。

  4. 猫与幸福感正相关

    根据发表在 PLOS One 期刊上的一项研究,猫的数量与幸福感正相关,一个国家的猫越多,其幸福感通常越高。但主要通过猫传播的弓形虫感染率与幸福感负相关,弓形虫感染率较高的国家的幸福指数通常较低。两个关系似乎矛盾,但也可能与卫生、经济条件相关联。研究分析了 93 个国家和地区的数据,幸福指数平均值为 5.72,每万人拥有猫数量的平均值为 989 只,其中中国的幸福指数为 5.97,略高于平均水平;每万人拥有猫的数量为 376 只,显著低于平均值,中国的弓形虫感染率与平均水平相当。研究作者提出,猫对幸福感的影响可能与经济因素有关。也就是随着收入水平的提高,公共卫生条件会得到改善,虽然猫的数量增加,但感染风险不会有相应的增加。

  5. 科学家识别出三种叫声最响亮的鸟

    生物学家在巴西识别出三种叫声最响亮的鸟。至于这些鸟如何演化出最响亮的叫声则仍然是个迷。白钟雀(white bellbird)叫声能达到 125 分贝,裸喉钟雀(bare-throated bellbird)和红腿叫鹤(red-legged seriema)叫声都超过 120 分贝。三种鸟类有着共同的生理特征:即宽大的喙口和肌肉厚实的大块头身体。研究人员表示,120 分贝可能是鸟类发声的生理极限;巴西之外的地方无疑也存在叫声洪亮的鸟类,但其音量不太可能比这些鸟高出太多。根据精确的测量标准,这三种鸟都可以被视为叫声最响亮者。但对科学家而言,谁是赢家并不是重点。

  6. 因涌入大量 AI 报告 Google 冻结其 Bug 悬赏计划

    因涌入大量无效的 AI bug 报告,Google 宣布冻结其 Bug 悬赏计划 Open Source Software Vulnerability Reward Program (OSS VRP),该决定于 10 月 1 日生效。Google 承诺将于 2027 年第一季度提供相关更新,期间将对该项目进行调整。Google 鼓励参与者探索其它 bug 奖励计划。大模型以及 Bug 搜寻自动化脚本的流行,导致了 OSS VRP 项目涌入了大量低质量的报告,Google 和开源项目维护者因此不堪重负,很多报告声称发现了 bug,但实际上无效,而维护者们浪费了大量时间去验证这些报告,无法专注于修复真正重要的 bug。整个行业都存在相同的问题。

  7. 2026 年诺贝尔生理学或医学奖授予了三位研究光遗传学的科学家

    2026 年诺贝尔生理学或医学奖授予了美国斯坦福-霍华德·休斯医学研究所的 Karl Deisseroth、德国柏林洪堡大学的 Peter Hegemann 和维尔茨堡大学的 Georg Nagel,以表彰他们在光门控离子通道和光遗传学方面的发现。大脑如何掌管情感、行为和身体功能,长期以来一直是个谜。20世纪,研究人员开始探究大脑的哪些区域影响哪些功能,但所用的方法意味着他们无法证明因果关系。光遗传学——一种能揭示神经细胞如何在活体大脑中塑造记忆、情感和行为的方法——改变了这一状况。Peter Hegemann 和 Georg Nagel 发现了通道视紫红质——一种具有独特性质的藻类蛋白质,存在于细胞表面。当它被蓝光照射时,蛋白质上会打开一个通道。带电离子随即流入细胞,产生电脉冲。他们发现,无论将这种蛋白质放入哪种细胞,那些细胞都会变得对光敏感。Karl Deisseroth 将通道视紫红质的基因导入大鼠的神经细胞中。通过用蓝光照射这些细胞,他能够触发神经信号。

  8. Riot Games 否认根据 CPU 封禁玩家

    一位玩家声称在购买了一块二手 CPU Ryzen 7 5800X3D 之后,因前拥有者有作弊行为这块 CPU 被列入了封禁黑名单,导致 Riot Games 旗下的所有游戏都无法启动。Riot Games 工作室负责反作弊的高管 Phillip Koskinas 通过社交媒体否认了这一说法, 他称没有找到相关记录,该公司的硬件封禁最长持续四个月,且只针对特定游戏,不会波及该公司的其它游戏,不会因为作弊者使用了一个硬件组件就将其它硬件组件加入到封禁名单。

  9. Anthropic 举报了与 Claude 聊天中发出威胁的佛罗里达女子

    Anthropic 举报了一名佛罗里达女子,原因是这名女子在与 Claude 聊天中威胁要枪击 Lee 县警长办公室,并在后续聊天中提及自己正在搞一把新枪。该女子在家中被捕,法庭记录显示她于 9 月 30 日受到了一项重罪指控。逮捕报告显示,Anthropic 会监控聊天内容,查找可能被视为有威胁性的关键词句,根据威胁程度相关内容可能会被提交给人工进行审核。在本案中,审核团队决定将发现的情况报告给执法部门。她的罪名是通过书面或电子形式发出大规模枪击或恐怖主义行动的威胁。

  10. 2026 年诺贝尔物理奖授予了冰立方中微子天文台提出者 Francis Halzen

    2026 年诺贝尔物理奖授予了美国科学家 Francis Halzen,以表彰其“对冰立方中微子天文台的决定性贡献以及发现具有天体物理起源的高能中微子”。中微子无处不在,它们径直穿过地球,穿过人体,而我们毫无察觉。极少数情况下,一个中微子会与一个原子核发生相互作用,这使得拥有合适设备的人有可能发现它们。Francis Halzen 于 1988 年首次提出了在南极捕获中微子的构想。当中微子与原子核碰撞时,会产生一闪光,这种光可以被清澈冰川冰中的传感器追踪到。南极的冰具有许多优势,因为它不受各种干扰的影响,而且该地区地质稳定,没有地震。他的想法很快得到了其他研究人员的支持,仅仅几年后,就在冰中的传感器上进行了初步测试。冰立方中微子天文台覆盖了整整一立方公里,于 2011 年完工。研究人员很快发现了第一批高能中微子,寻找宇宙中微子来源的工作由此可以正式开始了。

  11. 俄罗斯女研究员因感染肺鼠疫死亡

    在引发广泛关注之后,俄罗斯证实一名女研究员因感染肺鼠疫死亡,否认有其他人感染。俄罗斯称,在西伯利亚和远东地区抗鼠疫科学研究所工作的 28 岁女死者希皮洛娃(Darya Shipilova),是因为意外打破试管而被流出的肺鼠疫(pneumonic plague)感染。曾跟她接触的人接受检验,发现两人感染冠病,两人感染鼻病毒(rhinovirus),但并未发现有人染上肺鼠疫或其他疫病。研究所所在的伊尔库茨克(Irkutsk)州州长科布泽夫(Igor Kobzev) 则发表声明称,希皮洛娃是死于“未能确认病因的肺炎”,而研究所员工过后几天也未出现新的感染病例。鼠疫为细菌感染疾病,可通过几种方式传播,其中以肺鼠疫的传染力最大,也是唯一可经由呼吸道迅速人传人的病种。若迅速检测证实感染,可使用抗生素治疗,但如果不获治疗,则可在一至六天内致命。

  12. Manus 成功融资逾 5 亿美元

    经历收购风波的中国 AI 企业 Manus 完成超过 5 亿美元融资,创始人肖弘也已解除边控,让这家一度卷入中美科技博弈、前途未卜的公司迎来“重启”,也彰显了中国资本市场对 AI 的热情。 Manus 的母公司蝴蝶效应,星期四(10月8日)在公众号宣布融资消息。这是中国 AI 应用领域迄今规模最大的单轮融资之一,由中国私募股权基金博裕资本和老牌美元基金 IDG 领投,老股东腾讯、红杉中国、真格基金跟投。 公司投后估值达到 40 亿美元,也让 Manus 成为中国估值最高的 AI 智能体初创企业。 去年 12 月 Meta 斥资 20 亿美元收购 Manus,但交易被中国监管部门要求撤回。

  13. 北欧饮食与长寿相关

    众所周知,地中海饮食有利于健康长寿。现在研究人员报告另一种欧洲饮食——北欧饮食也与长寿相关。丹麦、芬兰、冰岛、挪威和瑞典等国的传统饮食与地中海饮食有很多相似之处,差不多是其寒冷版本,因此又名北方的地中海饮食。北欧饮食以植物为主,主要食用富含脂肪的鱼和根茎蔬菜。研究人员分析了于 64,000 名瑞典中老年人的健康数据,发现饮食习惯更符合北欧饮食的人的全因死亡率、心血管死亡率和癌症死亡率更低。

  14. 已知最早的游泳哺乳动物

    对一件早白垩世哺乳动物化石的分析证实约 1.25 亿年前的小型哺乳动物已经具备明确的半水生适应特征,这也是目前可确认的、最早具备游泳能力的哺乳动物。新发现的哺乳动物被命名为“板尾董尖齿兽”(Dongoconodon platycauda)。板尾董尖齿兽展现出了独特的半水生适应特征,是目前已知哺乳动物冠群中最早具有游泳能力的代表。其前后足具有发达的侧向扩展结构,与现生鸭嘴兽的蹼足高度相似,表明其生前可能具有发达的蹼膜;与此同时,董尖齿兽的掌骨、跖骨和指趾骨排列能够使手指和脚趾向外展开,从而增加划水时与水接触的面积,有利于游泳推进。 虽然板尾董尖齿兽的手足与鸭嘴兽十分相似,但它的尾巴却采用了完全不同的结构。化石保存了至少19节尾椎,其中靠近尾巴基部和中段的尾椎具有明显扁平的椎体和较宽的横突,表明这只动物具有背腹方向扁平的尾部。不过,它的尾巴并不像现代河狸和鸭嘴兽那样形成宽大的“桨状尾”,而是从基部向末端逐渐变细,更接近现代半水生啮齿类的海狸鼠。这只生活在恐龙时代的小型哺乳动物,可能拥有“鸭嘴兽式的手足”和“海狸鼠式的尾巴”,形成了一种此前未知的游泳方式组合。

  15. 美国男子因利用 AI 生成音乐和机器人账号欺诈播放被判 18 个月

    54 岁的美国北卡罗来纳州男子 Michael Smith 因 AI 辅助欺诈罪被判入狱 18 个月。他是首位涉嫌 AI 辅助流媒体欺诈而被刑事起诉的美国人。Smith 从 2017 年起利用虚假电邮账户和通过欺诈获取的借记卡在 Amazon Music 、Apple Music、Spotify 和 YouTube Music 等平台创建了数千个账号,使用软件操纵机器人程序,循环播放他声称拥有版权的歌曲——这些歌曲实际上都是 AI 生成的。到 2024 年被捕时他创作了数十万首 AI 生成歌曲,播放量多达数十亿次,从流媒体获取了数百万美元的收入。他的歌曲的总播放量甚至超过了当今最炙手可热的歌星。比如 2023 年 4 月其歌曲的播放量达到了 8090 万次,相比下 Taylor Swift 所有歌曲的总播放量同期仅为 930 万次。Smith 除了服刑外还必须上缴其非法获取的 8,091,843.64 美元收入。