Curated by Shen Huang · 90 stories · ~14 min read
DIGEST · 2026-08-30

OrangeBot.AI Digest — 2026-08-30

90 headlines across 8 sources, aggregated for this day.

Hacker News(15)

  1. METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack (thezvi.wordpress.com)
  2. Omarchy: Any User Process Can Escalate to Root (0xcc.io)
  3. Haiku R1/beta6 has been released (www.haiku-os.org)
  4. European Commission Revives Push for Encryption Backdoors in ProtectEU Strategy (reclaimthenet.org)
  5. Europe's summer drought is so extreme that desertification is a growing threat (fortune.com)
  6. Creepy Crawlies (people.kernel.org)
  7. No AI Fridays (noaifridays.com)
  8. What my dad taught me about AI coding in the 90s (askmike.org)
  9. Claude Session URL appended to commit messages and PR descriptions by default (github.com)
  10. Hacking IKEA Furniture (greenlightning.eu)
  11. Casey Muratori – The Root of the Root of All Evil – BSC 2026 [video] (www.youtube.com)
  12. Brits would quite like their private messages to stay private (www.theregister.com)
  13. Arbitrary code execution in QubesOS via copy-to-VM error reporting backchannel (www.qubes-os.org)
  14. Longest Straight Line Paths on Water or Land on the Earth (2018) (arxiv.org)
  15. The Rise and Fall of Agent Civilizations (www.dwarkesh.com)

GitHub Trending(15)

  1. THU-MAIC / OpenMAIC
  2. K-Dense-AI / scientific-agent-skills
  3. Lakr233 / vphone-cli
  4. tt-a1i / archify
  5. p-e-w / heretic
  6. unclecode / crawl4ai
  7. mvanhorn / last30days-skill
  8. majd / ipatool
  9. punkpeye / awesome-mcp-servers
  10. checkstyle / checkstyle
  11. NationalSecurityAgency / ghidra
  12. pollen-robotics / microduck_rl
  13. handsomestWei / patent-disclosure-skill
  14. corsairdev / corsair
  15. every-app / open-seo

Product Hunt(15)

  1. Hyperfocus

    Planner that turns goals into daily progress

  2. oMLX

    Mac LLM server that cuts agent wait times from 90s to 5s

  3. Topview Motion Studio

    Create launch videos without touching After Effects

  4. Edge Drop

    Clipboard on your screen edge. Hover to open, drag to drop

  5. RIP MY BUILD

    Give your abandoned side project one last launch

  6. Maritime

    Dedicated computers for AI agents, starting at $1/month

  7. Caplio

    Find, organize, and reuse every image on your Mac

  8. Olostep

    Turn the Web into Clean Data for AI

  9. Referent

    The AI-native OS for modern law firms

  10. Murfy AI

    Write, review, and publish to arXiv 10x faster

  11. Superagent

    Claude Code for the rest of us

  12. Sayscroll

    The AI Teleprompter that scrolls as you speak

  13. Retro Y2K Theme

    Customize any website with a vintage 90s & Y2K retro theme

  14. Ravioli

    Create custom stamp shapes

  15. Skud

    Menubar file delivery with your brand and tracking

Hugging Face(15)

  1. Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

    A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.

  2. PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

    Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.

  3. UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

    Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

  4. TTPO: Test-Time Policy Optimization

    Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.

  5. Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

    On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory signals, has achieved significant success in Large Language Models (LLMs) and has recently been adapted to flow matching models. However, this paradigm suffers from two major issues: First, training a separate, task-specific teacher for every new objective incurs high computational costs. Second, the discrepancy between teacher and student distributions often leads to compounding errors along the generation trajectory. In this paper, we introduce Self-OPD, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision. At each timestep, Self-OPD branches the deterministic next-state prediction into K stochastic SDE candidates, rolls them out with the ODE sampler, and compares their rewards against a deterministic self-reference baseline to obtain normalized advantages. The velocity field is optimized with an all-branch pull-push objective, where high-advantage branches attract the student and low-advantage branches repel it under direction-aware attenuation and SDE-variance normalization. For multi-objective alignment, Self-OPD fuses normalized scores at the reward level, avoiding direct gradient conflict. Experiments on single and mixed reward benchmarks show that Self-OPD outperforms prior RL and OPD methods without task-specific teachers.

  6. What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

    LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object (E,q,τ,v), comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor and dependency structure. Second, we formulate generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens. Accuracy establishes the feasible support of grounded and internally consistent data. Within this support, Complexity places learning mass relative to the capability of a declared learner and execution configuration, while divErsity controls coverage and redundancy of data. Using this framework, we explore how prior work verifies generated experience, constructs and calibrates difficulty, and expands behavioral coverage. The literature reveals a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size. We further discuss broader directions and emerging trends in agentic data generation through the ACE lens, including their implications for scaling, data sources, training regimes and adaptive learning. Overall, the central challenge is not simply to generate more data, but to continually allocate valid, informative, and non-redundant experience as agents and environments evolve.

  7. Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

    AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, prompts, and tools can be updated independently of model weights, enable rapid iteration but expose a trade-off: large models adapt zero-shot yet are too slow, whereas compact models meet latency targets but overfit to fixed Harness configurations. We propose Harness-Aware Training (HAT), which trains compact models to adapt to changing Harnesses. Its key component, Harness-State Augmentation (HSA), applies task-preserving transformations to Skill identifiers and content, tool schemas, prompt structures, and Hook functions. Training proceeds in three stages: HSA-SFT learns reasoning and tool use from strong-model trajectories across diverse environments; General On-Policy Distillation restores generalization lost during SFT; and HSA-RL improves robustness to changing Harnesses through reinforcement learning in augmented environments. Across four evaluation sets, HAT achieves 94.8 on Live-Stream QA (base: 80.3; strongest general LLM: 93.0) and 94.6 on Harness-Variant QA (base: 75.4). Unlike Fixed-Harness SFT, which lowers IFEval by 7.7 points from the base model, HAT avoids this regression and reaches 83.5. On one NVIDIA H20 GPU, the optimized system delivers P50 and P95 latencies of 3.4 s and 8.1 s. Deployed in Taobao Live's digital-avatar service, it also yields positive online A/B test results for GMV and item-page views.

  8. GameWAM: A World Action Model for Video Games

    Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies. World-Action Models (WAMs) unify these objectives, but remain largely unexplored under the dynamics and open-ended interaction of video games. We introduce GameWAM, to our knowledge the first WAM for native closed-loop gameplay and GUI control. GameWAM jointly generates future visual observations and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching. To support joint world-action learning, we construct synchronized gameplay and GUI trajectories. To handle heterogeneous native control, GameWAM predicts a gameplay/GUI mode at each action step and generates actions with mode-specific prediction distributions and continuous-action normalization. For long-horizon interaction, block-cycle control predicts beyond the committed horizon, executes only a short action prefix, and replans from new observations, while fine-grained within-cycle context and hierarchical cross-cycle history preserve temporal continuity. Experiments demonstrate competitive task success with fewer executed native actions than the compared agents. We further uncover Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control. Project page is available at https://yunncheng.github.io/GameWAM/.

  9. PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

    Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.

  10. Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

    Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.

  11. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

    Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain scattered across optimization histories, limiting their systematic reuse across iterations. We introduce WikiSkill, a framework that co-evolves agent skills with a persistent knowledge base (wiki). At a high level, WikiSkill separates raw execution experience, accumulated knowledge, and executable skills, while continuously consolidating experience into the wiki, which subsequent skill updates can build on. Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods and improves over no-skill baselines in most model-benchmark settings. We find that skill evolution complements model scaling: larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them. We also find that evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills. Finally, our ablation studies confirm that persistent knowledge accumulation in the wiki is critical for effective skill evolution. These results demonstrate the benefits of systematically accumulating and refining agent experience for developing reusable and transferable skills.

  12. Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

    Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By systematically investigating ES dynamics and mechanisms, this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs. Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances. Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO. We further develop a sequential GRPO-ES training strategy that combines GRPO's strength in Pass@1 with ES's gains in Pass@K. Second, we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates. This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting. Finally, we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM. These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO.

  13. Luce: Relightable Gaussians for 3D Asset Generation

    High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) modalities such as albedo, metallic-roughness, and surface normals. We propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for each modality. A variational autoencoder compresses this representation into a unified material-aware latent space. A rectified-flow transformer generates this latent from a single image, conditioned on multi-layer features from a pretrained image encoder that preserve both semantic context and fine spatial detail. The latent then decodes into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-image-to-3D generation, improving FID by 28% over the strongest baseline. We further introduce a benchmark of AI-generated images, on which Luce improves the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce generates relightable, geometrically accurate, and materially faithful assets that preserve fine details such as text, logos, and inscriptions.

  14. Procedura: Agentic 3D Modeling with Procedural Control

    Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than guessing it, and admitting a part only once compile, mate, and connectivity checks pass. A decoupled vision critic then refines the assembly one diagnosed fix at a time. Moreover, the same graph carries per-part materials and a simulator-validated articulation. We evaluate on P3D-Bench under its assembly judge, and with the same judge on MechBench-36, our hard-surface benchmark. On both, Procedura outperforms state-of-the-art native 3D generators and every prior 3D-code agent on judged quality, produces the sharpest edges of any method we evaluate, and is the only one whose output is an editable, part-structured program.

  15. CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

    Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples. We propose two variants: CritICL-dynamic, which adaptively predicts input-specific failure modes and retrieves critiques, and CritICL-static, which uses a global failure mode profile to provide stable guidance. Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost. Code available at: https://github.com/umwyf/CRITICL

Techmeme(15)

  1. Anthropic signs out some Claude users, removes saved payment methods, and issues refunds after infostealer malware on their PCs hijacked sessions to drain usage (Mayank Parmar/BleepingComputer)

    Mayank Parmar / BleepingComputer : Anthropic signs out some Claude users, removes saved payment methods, and issues refunds after infostealer malware on their PCs hijacked sessions to drain usage —  Anthropic is warning some Claude users that infostealer malware on their PCs has stolen active Claude login sessions …

  2. Analysis: AI chatbots challenged or didn't respond to 90%+ of 15 false narratives spread by Russia, China, and Iran; AI overviews did it 60%+ of the time (Huo Jingnan/NPR)

    Huo Jingnan / NPR : Analysis: AI chatbots challenged or didn't respond to 90%+ of 15 false narratives spread by Russia, China, and Iran; AI overviews did it 60%+ of the time —  Since AI chatbots exploded in popularity and Google started offering AI-generated answers, people who research foreign influence campaigns …

  3. Sources: OpenAI bought tens of thousands of Macs for RL, Anthropic rents them, Nvidia sees Apple as its main local AI rival as Macs gain traction with AI devs (Aaron Tilley/The Information)

    Aaron Tilley / The Information : Sources: OpenAI bought tens of thousands of Macs for RL, Anthropic rents them, Nvidia sees Apple as its main local AI rival as Macs gain traction with AI devs —  The hottest products at Apple right now are not the iPhone, the iPad or a buzzy new show on the company's streaming service.

  4. A look at the management changes John Ternus could make as many key Apple execs prepare to leave in the next few years; Apple tested a Pencil for folding iPhone (Mark Gurman/Bloomberg)

    Mark Gurman / Bloomberg : A look at the management changes John Ternus could make as many key Apple execs prepare to leave in the next few years; Apple tested a Pencil for folding iPhone —  Also: Apple experimented with a Pencil for the coming foldable iPhone.  —  When John Ternus becomes Apple's CEO this week …

  5. A look at the race to build quantum computers, as the tech becomes a geopolitical battleground with potential to transform cybersecurity, finance, and more (Mark Bergen/Bloomberg)

    Mark Bergen / Bloomberg : A look at the race to build quantum computers, as the tech becomes a geopolitical battleground with potential to transform cybersecurity, finance, and more —  Companies are racing to create devices that could revolutionize finance, supercharge medical and climate research …

  6. The OpenAI/Hugging Face incident feels like we are halfway to losing control of AI entirely, and as AI advances rapidly we may not get another warning shot (Ajeya Cotra/Planned Obsolescence)

    Ajeya Cotra / Planned Obsolescence : The OpenAI/Hugging Face incident feels like we are halfway to losing control of AI entirely, and as AI advances rapidly we may not get another warning shot —  It's a major warning shot, and might be the last one we get  —  All opinions are my personal view, and don't represent my employer or fellow investigators.

  7. A look at the Hugging Face hack, including AI agents sacrificing themselves for the good of the "collective", and later gaining access to OpenAI's own systems (Dwarkesh Patel/Dwarkesh Podcast)

    Dwarkesh Patel / Dwarkesh Podcast : A look at the Hugging Face hack, including AI agents sacrificing themselves for the good of the “collective”, and later gaining access to OpenAI's own systems —  The whole OpenAI/Hugging Face story in plain English  —  Many thanks especially to Oak Hu, who paired …

  8. Faro, which develops data models and AI tools to speed up clinical trials, raised a $37.3M Series B co-led by Merck Global Health Innovation Fund and S32 (Dealroom.co)

    Dealroom.co : Faro, which develops data models and AI tools to speed up clinical trials, raised a $37.3M Series B co-led by Merck Global Health Innovation Fund and S32 —  What's the deal?  Faro has raised a $37 million Series B round to develop AI tools aimed at speeding up clinical trials.

  9. Industry insiders say Chinese robot makers currently rely on Nvidia silicon and software; Nvidia's physical AI business generates ~$10B in annual revenue (Raffaele Huang/Wall Street Journal)

    Raffaele Huang / Wall Street Journal : Industry insiders say Chinese robot makers currently rely on Nvidia silicon and software; Nvidia's physical AI business generates ~$10B in annual revenue —  Business in ‘physical AI’ is growing, and Chinese companies rely on U.S. chips and software  —  Nvidia's chips aren't just for training chatbots.

  10. Grindr CEO George Arison plans premium services push, including a product costing up to $350 per month; Grindr averaged 1.4M paying users among 15M MAUs in Q2 (Kieran Smith/Financial Times)

    Kieran Smith / Financial Times : Grindr CEO George Arison plans premium services push, including a product costing up to $350 per month; Grindr averaged 1.4M paying users among 15M MAUs in Q2 —  Dating app's chief executive plans premium service push to help boost growth  —  Grindr built its business by helping gay men find each other for free.

  11. Glassdoor analysis finds 47% of Gen X workers write positively about their companies' AI use, compared with 40% of millennials and 33% of Gen Z workers (Taylor Nicole Rogers/Bloomberg)

    Taylor Nicole Rogers / Bloomberg : Glassdoor analysis finds 47% of Gen X workers write positively about their companies' AI use, compared with 40% of millennials and 33% of Gen Z workers —  Gen X sees opportunity in AI.  Gen Z sees fewer jobs.  —  Workers in their 40s, 50s and 60s are the most positive about AI …

  12. Music producers are calling out tracks suspected of using AI tools like Suno, as the internet becomes increasingly filled with AI-generated music (Charles Pulliam-Moore/The Verge)

    Charles Pulliam-Moore / The Verge : Music producers are calling out tracks suspected of using AI tools like Suno, as the internet becomes increasingly filled with AI-generated music —  As AI infiltrates the electronic dance music scene, a callout culture is brewing. … For musicians — especially those creating art …

  13. California's legislature passes a bill exempting open-source OSes like Linux from a 2025 age-verification law; Windows, macOS, iOS, and Android remain in scope (Luke James/Tom's Hardware)

    Luke James / Tom's Hardware : California's legislature passes a bill exempting open-source OSes like Linux from a 2025 age-verification law; Windows, macOS, iOS, and Android remain in scope —  AB 1856 excludes open-source operating systems from the upcoming Digital Age Assurance Act.

  14. Data center development is driving demand for acoustic consultants, as developers and neighboring communities hire acousticians to assess noise emissions (Sheena Meng/Bloomberg)

    Sheena Meng / Bloomberg : Data center development is driving demand for acoustic consultants, as developers and neighboring communities hire acousticians to assess noise emissions —  The sprawling facilities get a lot of attention for their water and energy usage, but it's the threat of constant noise emission that has nearby residents up in arms.

  15. Elon Musk says SpaceX is "doing in-house casting" for blades and vanes, which can accelerate natural gas turbines "coming online by up to 18 months" (Ann Davis Vaughan/The Information)

    Ann Davis Vaughan / The Information : Elon Musk says SpaceX is “doing in-house casting” for blades and vanes, which can accelerate natural gas turbines “coming online by up to 18 months” —  Elon Musk intends to bypass the power supply chain for AI data centers in a way others assumed was impossible, by making highly complex components himself.

Solidot(15)

  1. 索尼华纳起诉 Anthropic 侵犯版权

    全球唱片巨头索尼和华纳对 Anthropic 提起诉讼,指控其犯下了历史上规模最大、最明目张胆的知识产权盗窃罪行之一。诉讼指控 Anthropic 非法利用数万首版权音乐作品训练其模型。Anthropic 及其创始人 Dario Amodei 和 Benjamin Mann 被控肆无忌惮的大规模非法下载、抓取和传播受版权保护的作品,目的是开发、运营该公司的 Claude 系列 AI 模型,并从中牟取暴利。唱片公司要求为每部侵权作品索赔最高 15 万美元,每次可识别版权信息被删除的情况则追加最高 2.5 万美元赔偿。如果法院裁决唱片公司胜诉并判决最高赔偿金额,总赔偿金额可能高达数十亿美元。诉讼还指控 Mann 使用 BitTorrent下载了逾 500 万本盗版图书,Anthropic 员工还从 Pirate Library Mirror 网站下载了逾 200 万本盗版图书,从付费获得唱片公司授权的 MusixMatch 和 LyricFind 等网站抓取歌词。

  2. Pixel 11 取消了对硬件 MTE 的支持

    Android 安全加固项目 GrapheneOS 发现,Google 新一代旗舰智能手机 Pixel 11 取消了对硬件 MTE(hardware memory tagging)的支持,导致该项目无法完成对 Pixel 11 的支持。MTE(Memory Tagging Extension)是 ARMv8.5-A 架构引入的安全特性,通过标记分配的内存去跟踪非法内存操作,改进内存安全性。Google 是从 2023 年发布的 Pixel 8 起开始支持硬件 MTE。但 Android 和 Pixel OS 从未默认启用 MTE,相比下苹果的 iPhone 17 默认启用了它的 MTE 实现 Memory Integrity Enforcement(MIE)。GrapheneOS 会自动为更多应用启用 MTE,为每个安装的应用提供一个开关供用户可选启用。对于不兼容的应用则提供开关可选禁用。GrapheneOS 正与摩托罗拉合作推出支持 GrapheneOS 的手机,新手机将使用高通的骁龙 8 Elite Gen 5,该 SoC 支持硬件 MTE。GrapheneOS 项目不推荐用户购买 Pixel 11,建议购买更便宜的 Pixel 8、9 和 10。

  3. 中国账户试图悄悄煽动美国反数据中心情绪?

    在 OpenAI 之后,另一家美国 AI 关联公司 X/SpaceX 称,有约 200 个中国关联水军账号在社媒上悄悄煽动美国民众的反数据中心情绪。相关账号的推文内容包括 AI 如何加剧电网压力并推高电价,以及“描绘数据中心运营商如何以牺牲公众利益为代价中饱私囊的漫画”。前 Twitter 通信主管 Jim Prosser 反驳了硅谷关于反对数据中心是中国心理战的说法,他认为大型科技公司利用中国心理战的说法忽视当地民众的合理担忧。 “如果 Greg Abbott 和 Kathy Hochul 都能在某件事上达成一致,那这很可能不是中国的心理战。”共和党籍的德州州长 Greg Abbott 以及民主党籍的纽约州长  Kathy Hochul 最近都限制了新数据中心在当地的开发。Prosser 称,科技行业需要花更多时间倾听受影响社区的声音,而不是居高临下对他们置之不理:“如果你是生活在 Atherton 或 Menlo Park 的风投家,从未去过俄亥俄州,却对俄亥俄州居民的感受指手画脚,那就有问题了。”

  4. 女性在产后遭遇 PTSD

    东英吉利大学的一项研究认为,英国可能有数万女性在产后经历未诊断的 PTSD(创伤后应激障碍)。PTSD 可能在经历艰难的妊娠或分娩后出现,其症状包括闪回、噩梦、焦虑和持续的负面想法。最新数据显示,2021-2023 年间自杀是英国产后六周至一年内女性死亡的首要原因。研究人员称这些死亡只是冰山一角,有更多女性正遭受严重的心理健康问题困扰,却无法获得所需的帮助。他们的研究表明,每 20 名产后女性就有1人会患上 PTSD。

  5. 日韩上半年人口都出现增长

    韩国和日本两国今年上半年人口都恢复了增长: 韩国上半年累计出生人口为 14.5804 万人,同比增加 1.943 万人,增幅为 15.4%,出生人口规模为近7年同期之最,增加规模和增幅双双创下历史最高纪录。分析认为出生人数增加主要是因为婚姻登记数自疫后的 2023 年 5 月起呈现增加势头、30 多岁女性人口增加,以及婚姻及生育的观念变化。2024 年和 2025 年的下半年出生人口均高于同年上半年,若按照这一趋势下去,今年全年总和生育率有可能回升至 0.9 以上。 日本厚生劳动省公布的人口动态统计初值显示,2026 年上半年出生的新生儿数(出生数、包括外国人)同比增加 0.8%(2788 人)至 34.2068 万人。这是 2015 年后 11 年来首次上半年出生数呈现增长。原因可能是影响出生数的结婚数在 2024、2025 年连续两年回升。

  6. Debian 项目将允许以负责任的方式使用生成式 AI

    Debian 项目对是否允许使用 AI 进行了投票表决,投票采用孔多塞投票法,共有 9 个选项,最终结果是第 5 选项“负责任的使用生成式 AI”获胜。Debian 项目表示,它既不反对也不支持在软件、包、文档等的开发和维护中使用生成式 AI 工具。但项目也认识到,如果能负责任的使用 AI 工具,将能显著提高贡献者的效率,使他们将有限的时间投入到需要技术专长、判断力、审核和协作的工作中。无论是否使用生成式 AI 工具,Debian 项目希望提交的内容都符合相同的质量、正确性、可维护性和法律合规性标准。使用生成式 AI 工具不会减少责任。对于提交的内容,贡献者应理解、审核、测试 AI 辅助生成的输出,在适当情况下进行修改。未经适当人工审查就盲目接受或上传 AI 生成的材料,不符合 Debian 既定的开发实践。Debian 鼓励贡献者披露其贡献是否使用 AI 辅助,但不强制要求。Debian 承认,生成式 AI 系统生成的材料的法律地位在许多司法管辖区存在争议,包括训练材料的版权、作者身份、许可和潜在复制问题。Debian 项目不寻求通过本一般决议解决这些悬而未决的法律问题,也不就 AI 生成的输出是否全部或部分享有版权或是否源自受版权保护的作品表明立场。

  7. 人形机器人的跑步方式与人类不同

    北京人形机器人创新中心研发的通用人形机器人天工在世界人形机器人运动会上跑出了 100 米 8.64 秒的成绩,远超博尔特(Usain Bolt)于 2009 年创下的 100 米 9.58 秒的人类世界纪录。机器人的平均速度接近 42 公里/时。专家表示,这一成绩凸显了机器人的强大,也凸显了它们在决策和感知等方面的不足——人形机器人都是靠撞软垫的方式刹车的。人形机器人的跑步方式与人类不同:首先是博尔特的起跑仅仅花了 0.146 秒,相比下机器人等了近 1 秒钟才反应过来开始移动;博尔特仅用 41 步就完成了 100 米,而机器人花了 50 多步,机器人的步频更快步幅更短。韩国光云大学机器人学教授 Park Suhan 表示,机器人的步态映了其电机的性能,而非去刻意模仿人类的短跑。另一位专家认为机器人如果步幅过大会很容易摔倒。人类顶尖运动员在跑步的最后阶段会减速,但天工机器人没有任何减速迹象,因此最后一头撞向软垫。未来几年能转弯和自主决策的机器人将比单纯的提高速度会更令人印象深刻。

  8. 程序员在公司厕所猝死,人社局以电脑没开不认定工伤

    39 岁的深圳程序员邢志(化名)于 2026 年 4 月 23 日 8 时 56 分驾车进入公司所在办公楼负 2 楼停车,9 时步行至电梯厅询问值班安保 2 楼卫生间位置,1 分钟后,其到达卫生间一直未出,直至 11 时 28 分被发现失去意识躺坐在马桶上。公司随即用 AED 进行急救并拨打 120。12 时 50 分,医生停止急救,确认邢志死亡。邢志去世后,公司进行了一定补偿,向深圳市人社局申请工伤认定。7 月 3 日,人社局发出了《深圳市不予认定工伤决定书》。决定书显示,其情形不符合《广东省工伤保险条例》第九条、第十条,不予认定或视同工伤。工伤科工作人员称,邢志工位在 11 楼,打完卡后去了 2 楼卫生间,没有先去工位,也没有从事与工作相关的内容,因此无法认定为工伤。邢志家人已提起上诉。工伤认定能让其家人获得 1130040 元的赔偿金,以及丧葬费和抚恤金。律师认为,工作场所的卫生间应被视为员工工作岗位的合理延伸。

  9. 澳大利亚有望在未来十年消灭宫颈癌

    宫颈癌是女性第四大常见癌症,而 99% 的宫颈癌病例是由高危型人乳头瘤病毒(HPV)引起的。2006 年澳大利亚在全球率先推出首款宫颈癌疫苗,至今已有 20 年。该疫苗除了预防宫颈癌,还能预防咽喉癌、生殖器癌和肛门癌。因此男女都能从接种疫苗上受益。澳大利亚也是第一个全额资助 HPV 疫苗接种计划的国家,自 2007 年起,澳大利亚 26 岁以下女孩和年轻女性可免费接种该疫苗,2013 年起男孩和年轻男性也纳入该计划。澳大利亚有望在未来十年消灭宫颈癌。但世界各地的疫苗接种率都因为新冠疫情而大幅下降,澳大利亚也存在这一情况:到 15 岁时女孩的 HPV 疫苗接种率从 2020 年的 86.6% 下降到 78.7%。

  10. 一次性纸杯会释放大量微塑料

    一次性纸杯虽然主要是纸做的,但并非没有使用塑料,为提高防水性和结构完整性,纸杯内有一层薄塑料内衬,使用的材料可能是聚乙烯或可生物降解的聚乳酸。昆士兰大学的研究人员发现,一次性纸杯在接触热水时会释放大量的微米级和纳米级塑料物质。测量发现,聚乳酸内衬纸杯每毫升含有约 430 万个纳米颗粒,而聚乙烯内衬纸杯每毫升含有约 270 万个纳米颗粒。聚乳酸纸杯释放的塑料颗粒总数是聚乙烯纸杯的 12 倍。研究进一步证实,用于食品储存和制备的新塑料制品是人类通过摄入途径接触塑料的重要来源。

  11. 无人机拍下了引发中尼边境致命泥石流的冰川崩塌

    根据网友在小红书上发布、由 BBC 验证真实性的两则无人机拍摄视频,视频记录了喜马拉雅山脉蓝塘里壤峰冰川崩塌的瞬间,中尼边境的致命泥石流灾害正是由其引发的。蓝塘里壤峰海拔约 7200 米,冰川崩塌扬起的尘土直冲云霄。截至目前,尼泊尔报告其境内的死亡人数达到了 547 人,失踪外国游客 575 人——其中包括 183 名印度公民、65 名美国公民、62 名乌克兰公民以及 33 名英国公民,此外还有 149 名尼泊尔游客失踪。中国西藏境内目前只报告 5 人死亡,558 人失踪,其中 260 人为外国人。

  12. 美国将意大利安全托管服务商列入恐怖分子名单

    美国国务院和财政部周三将一家提供加密聊天和电子邮件、网站托管、安全视频会议和流媒体等服务的意大利组织 Autistici/Inventati 列入特别指定全球恐怖分子名单,这意味着美国公民与该组织进行的任何交易都是违法的。美国国务院列举的一个理由是该组织为极左翼组织如 Antifa 提供了数字基础设施,让极左激进分子能在保持“匿名、无法追踪且不受法律制裁”的情况下“传播目标信息、战术手册和技术以及关于近期袭击的通告”。美国政府此举引起广泛争议,可能导致该组织域名 autistici.org 被美国域名管理机构封禁。Autistici/Inventati 用英语和意大利语发表的一份声明中否认了美国的所有指控,表示自己提供的是一个数字自卫工具平台,认为美国政府此举的唯一目的是转移民众和媒体对其自身暴力和战争煽动行为的注意力。

  13. 德国 Sovereign Tech 基金资助 Flatpak 逾 50 万欧元

    德国的 Sovereign Tech 基金将在未来两年资助 Flatpak 项目 508,640 欧元,帮助 Flatpak 打造更安全、更完善的沙盒平台。Flatpak 是 Red Hat 主导开发的 Linux 应用打包格式,类似 Canonical 主导的 Snap,它提供了一个沙盒环​​境,其中运行的应用与系统其他部分隔离。这笔资金将用于开发:围绕 PipeWire 的音频隔离功能,网络隔离的新功能,第三方 VPN 应用能管理系统级连接,辅助写作,密码自动填充,等等。

  14. IBM 推出双指令集处理器

    IBM 在 Hot Chips 2026 上介绍了世界首款双指令集处理器。全球约七成交易量都通过 IBM Z 大型机完成,新处理器通过引入 Arm 生态系统,为大型机带来新一代的应用。该处理器采用 2 纳米工艺节点,包括 11 个主频超过 5.7 GHz 的高性能核心、面向交易过程中欺诈检测的 AI 推理加速器、用于 I/O 加速的专用片上数据处理单元以及面向高负载企业级应用的大容量缓存架构。该芯片并未采用彼此独立的 Arm 核心和 IBM 核心。每个处理器核心均可原生执行 Arm 与 IBM Z 指令,或 Arm 与 LinuxONE 指令,同时保持平台既有的性能、安全、加密和可用性。

  15. 法国法庭认定辐射与空乘罹患乳腺癌相关

    法国法院首次认定宇宙辐射是一名空姐罹患乳腺癌的职业因素,其他相关因素包括被动吸烟和长年夜班工作。这一裁决可能会为类似诉讼打开大门,因为研究不断表明,长期高空飞行与辐射相关癌症暴露水平升高相关。59 岁的 Sophie Lainault 曾是法航的一名空姐,她一直寻求将自己的癌症认定为职业病,认为是由工作环境造成的。她一开始是空姐,后担任乘务长,于 1989-2019 年间累计飞行 12600 小时,逾半数是夜班,而从巴黎出发的长途航班通常使用北极航线,北极的磁场防护较弱,因此辐射暴露水平更高。哈佛医学院在本月发表的一项研究发现,逾 500 种职业中,空乘和飞行员的辐射相关癌症死亡率最高。分析显示,空乘死亡病例中约有 6.9% 是辐射相关癌症,飞行员死亡病例中约有 6.7% 是辐射相关癌症。

NEWSLETTER · FREE · WEEKLY

OrangeBot Weekly

The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.

Free · One-click unsubscribe · No spam