OrangeBot.AI Digest — 2026-08-23
90 headlines across 8 sources, aggregated for this day.
Hacker News(15)
- How I find problems to solve as a staff engineer (lalitm.com)
- A website for debloated open source alternatives (debloat.dev)
- Why Sal Khan't: On Learning by Making but Teaching by Telling (punyamishra.com)
- GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost (reinvently.co.uk)
- Coconut oil jet fuel matches kerosene's efficiency in engine tests (studyfinds.com)
- I spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a day (ericpardee.github.io)
- How Complex Systems Fail (1998) (how.complexsystems.fail)
- What Is a Harness? (earendil.com)
- Slovakia finds Russian backdoor in traffic speed cameras (risky.biz)
- My favorite nonfiction books about cults, scams, and schemes (bookdna.com)
- Malware infects Android-based automotive head unit firmware (securelist.com)
- I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes (www.xda-developers.com)
- JIT Compiling Code in 5μs (malisper.me)
- Wi-Fi 8 is the first wireless upgrade in years that isn't chasing speed (www.xda-developers.com)
- The End of an Athlon (www.os2museum.com)
GitHub Trending(15)
- openai / codex
- freestylefly / awesome-gpt-image-2
- mattpocock / skills
- basecamp / omarchy
- AprilNEA / OpenLogi
- block / buzz
- apache / maka
- Alishahryar1 / free-claude-code
- tinyhumansai / openhuman
- affaan-m / ECC
- ruvnet / ruflo
- VoltAgent / awesome-agent-skills
- virgiliojr94 / book-to-skill
- dani-garcia / vaultwarden
- anthropics / claude-plugins-community
Product Hunt(15)
- Aximote
Your car data, finally in your pocket
- Yattayo
A physical slider to-do board, faithfully rebuilt in 3D
- Claude Academy
The Official Learning Hub by Anthropic
- Local Music Organizer for Mac
Ultimate toolkit for local Apple Music library maintenance
- KanaSensei
Read Japanese kana in two weeks
- FetchSandbox MCP
The MCP that proves your AI's integration fixes work
- Tab Notes
Turn your browser new tab into a distraction-free notepad
- Yatko
The download button Github forgot to add
- Flown
Every flight you've ever taken, on one private map
- Construct Computer
Your AI coworker gets a computer. You get your day back.
- Plask
Have little ducks show how deep you dive on your Apple Watch
- OpenLogi
A local-first alternative to Logitech Options+
- ANCBuddy for Bose QC Ultra
Control Bose QC Ultra from your macOS menu bar
- Toplify
Track your App Store ranking worldwide
- AutoClaw
An AI work agent across desktop, browser, and chat
Hugging Face(15)
- EnvHarness: Awakening Static Worlds for Agent Learning
LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross-artifact consistency. FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, while analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment. These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.
- 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bottlenecks. On the reference-context side, conditioning on all previously generated views grows as O(N), weakening cross-view appearance guidance. On the target-context side, disjoint groups cannot directly exchange information, causing global structural drift. 4DAnyone addresses both bottlenecks with two complementary designs: Reference Context Packing (RCP) compresses growing reference views into a fixed-length mixed-resolution context with O(1) reference-context complexity, while Target Context Routing (TCR) rotates target-view groupings during denoising to share context across groups at high-noise steps and stabilize details at low-noise steps. We further build the MVGameHuman dataset using our in-house game engine and combine it with light-stage and in-the-wild video datasets for training. Experiments on DNA-Rendering and DyMVHumans show that 4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization. See our project page for video results and source code: https://4danyone.github.io.
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce SWE-bench Science, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, Claude Code with Opus-5 (max), achieves a pass@1 below 50\%, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.
- WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.
- ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.
- Repo0: Design-Driven Zero-to-All Code Generation
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.
- τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
- Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.
- EXIMO: VLM Guided Exploration of VLA Policies
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular, collecting teleoperation datasets requires hundreds of hours of expensive human labour and the alternative, reinforcement learning (RL), can be notoriously sample-inefficient especially for long-horizon tasks. In addition, RL with VLAs imposes several challenges due to the model's size and architectural design. In this work, we propose EXIMO, an efficient algorithm for finetuning of VLA policies. EXIMO operates in three stages: explore, imitate, and optimize. During the explore phase, EXIMO equips the VLA with a vision language model (VLM) that acts as a planner. The VLM thinks and breaks down challenging long-horizon problems into shorter ones for the VLA. The VLM, together with the VLA, is used to collect an orchestrated dataset on new tasks. During the imitate phase, the VLA is finetuned with the orchestrated data. Finally, during the optimize stage, we use residual off-policy RL to further finetune the policy. In our experiments, we ablate all three stages of EXIMO and show that it outperforms existing approaches significantly in terms of sample-efficiency and final performance.
- The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM costs up to 1,431x more than an embedding model of comparable quality (USD 154 vs. USD 0.11 per benchmark pass), and the open LLMs tested process tokens 2.5 to 736x more slowly on the same GPU. Reasoning tokens account for 28 to 81% of LLM inference cost; lower reasoning budgets preserve or improve retrieval quality for most models in our ablation. The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro. These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval. Our code, datasets, and results are publicly available at https://github.com/embeddings-benchmark/embedders-dilemma.
- Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harness---is typically treated as a fixed artifact after deployment. This work studies an alternative where the harness is task-specific and continuously evolvable: each task family maintains its own harness, which is hot-swapped across iterations through a fixed task-injection seam and rewritten using environment feedback. We introduce Hierarchical Self-Improvement (HSI), a framework in which a single frozen LLM M operates across three hierarchical scopes: a task harness H that executes tasks, an evolver that rewrites H, and a meta-evolver that rewrites the evolver's strategy code under a frozen outer anchor. A thinking-on/off design isolates the contribution of harness evolution by disabling reasoning during task execution while enabling it during self-modification. HSI is bounded by two factors: a feedback-fidelity bound, since evolution requires informative reward signals to guide selection, and a backbone capability bound, since harness redesign cannot overcome limitations of the frozen model. On BALROG with DeepSeek-V4-Flash-Preview as the frozen backbone, HSI achieves consistent gains over the initial harness on moderate-difficulty tasks (+39.3 on BabyAI, +33.0 on Crafter, +25.0 on TextWorld, and +15.0 on MiniHack, all in raw \% Progress), while obtaining strong held-out generalization on BabaIsAI sub-suites (0.98 best-test on BreakStop and 1.00 on GoTo from a 20% unseen split). On tasks beyond the backbone's capability (NLE), harness evolution provides no improvement. These results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits. Code is available at https://github.com/TailinZhou/hsi.
Techmeme(15)
- Greg Abbott says data center companies "dug their own grave" by moving into communities without first gaining support, signaling growing Republican backlash (Axios)
Axios : Greg Abbott says data center companies “dug their own grave” by moving into communities without first gaining support, signaling growing Republican backlash — Texas Gov. Greg Abbott delivered one of the starkest warnings yet from a Republican to the AI industry …
- Sources: Hugging Face is exploring a sale that could value it at $13B+, up from $4.5B in 2023, and has been working with a bank to evaluate bidders' interest (Katie Roof/Business Insider)
Katie Roof / Business Insider : Sources: Hugging Face is exploring a sale that could value it at $13B+, up from $4.5B in 2023, and has been working with a bank to evaluate bidders' interest — The AI industry's next blockbuster acquisition may not be another model maker. — Hugging Face, whose platform helps developers discover …
- Ramp data: Fable 5, launched in June, has plateaued at ~11% of spending on Anthropic tools, as companies shift to cheaper models; Opus 5 surpassed Fable 5 (George Hammond/Financial Times)
George Hammond / Financial Times : Ramp data: Fable 5, launched in June, has plateaued at ~11% of spending on Anthropic tools, as companies shift to cheaper models; Opus 5 surpassed Fable 5 — AI lab's Fable 5 has met with sluggish demand from corporate clients — Anthropic's US customers are using cheaper alternatives …
- Alibaba plans to raise ~$10B in a follow-on share offering to fund AI investments; sources: it plans to offer 710M shares at a 3.6% discount to Friday's close (Reuters)
Reuters : Alibaba plans to raise ~$10B in a follow-on share offering to fund AI investments; sources: it plans to offer 710M shares at a 3.6% discount to Friday's close — China's Alibaba (9988.HK) on Sunday launched a HK$80-billion ($10.2 billion) share placement to fund artificial intelligence-related development.
- Sources: foldable iPhone feels durable, fits well in a pocket, has useful iPad-like app layouts, excels as a camera viewfinder but lacks telephoto and Face ID (Mark Gurman/Bloomberg)
Mark Gurman / Bloomberg : Sources: foldable iPhone feels durable, fits well in a pocket, has useful iPad-like app layouts, excels as a camera viewfinder but lacks telephoto and Face ID — Also: Get ready for iPhone price hikes. — Apple is about to bring some of its product magic back with its first foldable iPhone.
- The careers of Z.ai's Tang Jie and Moonshot AI's Yang Zhilin, once teacher and pupil at Tsinghua University, show that China's AI leap is no sudden development (Raffaele Huang/Wall Street Journal)
Raffaele Huang / Wall Street Journal : The careers of Z.ai's Tang Jie and Moonshot AI's Yang Zhilin, once teacher and pupil at Tsinghua University, show that China's AI leap is no sudden development — University lab nurtured the computer scientists who are using ingenuity and imitation to chase down Anthropic and OpenAI; ‘they know perfectly how to monetize their work’
- A profile of Judge Yvonne Gonzalez Rogers, who is presiding over US state AGs' social media addiction lawsuit against Meta and oversaw the Musk v. Altman trial (Jeffrey Kopp/CNBC)
Jeffrey Kopp / CNBC : A profile of Judge Yvonne Gonzalez Rogers, who is presiding over US state AGs' social media addiction lawsuit against Meta and oversaw the Musk v. Altman trial — It's been a crazy four months for Yvonne Gonzalez Rogers. — The judge in the Northern District of California spent late April …
- The popularity of risky leveraged chip ETFs in South Korea prompted regulators to cap individual exposure and mandate a weeklong investor education course (Financial Times)
Financial Times : The popularity of risky leveraged chip ETFs in South Korea prompted regulators to cap individual exposure and mandate a weeklong investor education course — Leveraged single-stock ETFs attracted billions of dollars of net inflows even as they plunged during market sell-off
- Sources: Flipkart Minutes, the quick commerce service of Walmart's Flipkart, is now delivering 1.1M-1.2M orders per day, up from ~390K-400K in November 2025 (Jagmeet Singh/TechCrunch)
Jagmeet Singh / TechCrunch : Sources: Flipkart Minutes, the quick commerce service of Walmart's Flipkart, is now delivering 1.1M-1.2M orders per day, up from ~390K-400K in November 2025 — Indian startups spent years getting consumers accustomed to having groceries and everyday goods delivered within minutes.
- Semiconductor cram schools are flourishing in Seoul as applications to Samsung and SK Hynix surge, fuelled by record earnings and eye-catching worker bonuses (Financial Times)
Financial Times : Semiconductor cram schools are flourishing in Seoul as applications to Samsung and SK Hynix surge, fuelled by record earnings and eye-catching worker bonuses — Private tutors teach semiconductor basics and help polish CVs for hopefuls looking to join lucrative sector
- Sources: Nvidia plans to use its $6B deal with Poolside to build an open-weight AI model to compete with Chinese models like DeepSeek and Kimi (Robbie Whelan/Wall Street Journal)
Robbie Whelan / Wall Street Journal : Sources: Nvidia plans to use its $6B deal with Poolside to build an open-weight AI model to compete with Chinese models like DeepSeek and Kimi — A sweeping agreement with startup Poolside aims to build an open AI ecosystem in the U.S. to compete with Chinese heavyweights and American AI giants
- Sources: Iran-linked hackers shut down a small UK power plant for four days, coinciding with a wave of Iran-affiliated attacks on US water utilities (Telegraph)
Telegraph : Sources: Iran-linked hackers shut down a small UK power plant for four days, coinciding with a wave of Iran-affiliated attacks on US water utilities — Unprecedented cyber attack believed to be most successful of its kind — Tony Diver , Political Editor. Rozina Sabur , National Security Editor.
- Sources and documents detail how Tether's plan to build two bitcoin mining sites in Uruguay fell apart amid a dispute with state utility UTE over power supply (Reuters)
Reuters : Sources and documents detail how Tether's plan to build two bitcoin mining sites in Uruguay fell apart amid a dispute with state utility UTE over power supply — Uruguay seemed like the perfect place for cryptocurrency giant Tether to launch a bitcoin mining operation.
- AI agents' growing capabilities are driving productivity FOMO among some startup founders, who feel compelled to work long hours managing and guiding the agents (Katherine Bindley/Wall Street Journal)
Katherine Bindley / Wall Street Journal : AI agents' growing capabilities are driving productivity FOMO among some startup founders, who feel compelled to work long hours managing and guiding the agents — The growing capabilities of AI give new meaning to working yourself to the bone — Seductive. Intoxicating. All-consuming.
- London-based Inherent, founded by DeepMind alumni and with $50M in seed funding, says its new Faraday agent beats GPT-5.5 at reproducing research paper findings (Anna Heim/TechCrunch)
Anna Heim / TechCrunch : London-based Inherent, founded by DeepMind alumni and with $50M in seed funding, says its new Faraday agent beats GPT-5.5 at reproducing research paper findings — Inherent, a London AI lab founded by Google DeepMind alumni, says its AI agent just outperformed much larger models from Anthropic and OpenAI using a fraction of the size.
Solidot(15)
- 机器人短跑超越人类,但刹住是问题
为期五天的世界人形机器人运动会于周六在北京开幕。运动会共设 51 个项目,包括 30 项体育竞技和 21 项场景化竞赛。超过 40% 的项目要求机器人完全自主运行。在 8 月 22 日的首日赛事中,两台机器人跑出了比人类百米世界纪录(9.58秒)保持者博尔特(Usain Bolt)更快的成绩,相比去年百米短跑仍然耗时 20 秒以上的机器人,可谓进步巨大。另一台机器人则在 400 米短跑中实现了 39.7 秒的成绩,超越了南非运动员 Wayde van Niekerk 43.03 秒的世界纪录。不过,在冲过终点线后,这些机器人的制动能力依然存在很大缺陷。现场画面显示,机器人纷纷撞向十几米开外的巨大软垫,然后跌倒在地,多台机器人翻倒后甚至出现了明显损坏。这次运动会还设置了在模拟工厂、餐厅、办公室及紧急情况场景下测试机器人性能的比赛项目。比如人形机器人能否在处理包装和仓储作业的同时,可靠地完成诸如线缆连接等精密任务?它们能否应对角度不当的线缆、刚好够不到的物体、发生位移的包裹等工厂中常见的困难?这些任务旨在评估它们在那些不那么引人注目的岗位上像人类一样工作的能力。
- 卡巴斯基发现第一种针对汽车的 Android 恶意程序
俄罗斯安全公司卡巴斯基的研究人员报告他们发现第一种针对汽车的 Android 恶意程序。恶意程序通过基于 Android 的兜风出行汽车主机(head unit)固件的内置更新程序传播,被认为与 MoYu Group 黑客组织有关,该组织与 BADBOX 僵尸网络有关联。卡巴斯基称它已经通知了兜风出行,对方表示已修复相关安全问题。这一汽车恶意程序传播案例类似廉价电视盒,攻击者旨在创建住宅代理僵尸网络,因此使用了相同的网络基础设施。
- 柳树和杨树释放出的化合物会恶化城市空气质量
数百万棵垂柳和白杨树将北京装缀成一个绿色的大都市。然而根据《Science Advances》上发表的一项研究,柳树和杨树释放出的化合物是城市空气污染的重要来源。广州暨南大学的研究人员最初想要了解人类活动对臭氧污染的影响,结果意外发现城市植被是臭氧的重要来源。植物会释放出挥发性有机化合物,作为植物光合作用的副产品,被广泛种植的柳树和杨树会释放出大量的异戊二烯。研究团队发现,北京 35% 的树木会释放异戊二烯。研究人员在北京各地采集空气样本,测量挥发性有机化合物浓度,在城市各监测站收集臭氧数据。研究发现,在 2021 年 5 月至 7 月期间,植物排放的化合物约占北京总排放量的 10%,其余来自人类活动如汽车尾气和工业化学品。植物释放出的挥发性有机化合物与大气中的羟基反应生成过氧自由基,过氧自由基再与空气中的氮氧化物(NOx)反应生成化合物,这些化合物在阳光照射下会转化为臭氧。研究发现,植物排放的有机化合物占最终生成臭氧的化学物质的 52%,其中异戊二烯是主要贡献者。人类活动产生的挥发性有机化合物总排放量高于植物排放,但对生成臭氧的贡献远小于植物。研究人员还调查了 24 个特大城市种植的树种,发现澳大利亚悉尼和墨尔本所种植树的异戊二烯排放量预计会高于北京,悉尼有 65% 的树木会排放异戊二烯。
- 波兰加密货币交易所 CEO 在 2022 年失踪,4 年后他的继任者也失踪了
Nicole Suszek 最后一次收到哥哥 Sylwester 的电话语音留言是在 2022 年,在留言中 Sylwester 急迫的请求她给他寄去比特币,否则以后就永远见不到面了。Sylwester 从此杳无音信,家人认为他已经遇害。Sylwester 是东欧和中欧最大加密货币交易所 Zondacrypto 的创始人,他在 2014 年创办了 Zondacrypto 的前身 BitBay。接替 Sylwester 担任 Zondacrypto CEO 的波兰律师 Przemyslaw Kral 在今年四月也失踪了,这一事件让 Sylwester 案再次浮出水面。波兰总理 Donald Tusk 则指责 Zondacrypto 与俄罗斯情报机构、有组织犯罪和右 翼政客有关联。Zondacrypto 网站在 4 月关闭,导致数十万客户无法提现或交易。Zondacrypto 发行的代币 ZND 已贬值逾 99.9%。Przemyslaw Kral 最后一次露面是在 4 月 16 日,他通过社媒发表了一则视频,呼吁客户不要恐慌,不要对交易所失去信心,称公司还有 4000 比特币,但这些比特币所在的钱包密钥只有前 CEO 才知道。加密货币专家对此表示怀疑,因为该钱包已有近十年没有活动了。对于 Kral 身在何处,有人据称曾在以色列、博茨瓦纳和迪拜等地目击到他,但这些说法都未经证实。代表 Zondacrypto 账户被冻结客户的华沙律师 Robert Nogacki 认为 Kral 在东南亚,他表示关于 Zondacrypto 他唯一确定的就是:“它从一开始就是个骗局。”
- 3 分钟冲刺跑产生的分子反应与 90 分钟中等强度运动截然不同
3 分钟冲刺跑产生的分子反应与 90 分钟中等强度运动截然不同。洛克菲勒大学的研究人员比较了人体对不同强度运动的反应。他们发现,六次 30 秒全力冲刺跑后,血液中近四分之一的蛋白质发生了变化。相比之下,90 分钟持续中等强度骑行仅改变了不到 0.25% 的蛋白质。中等强度的跑步机运动对蛋白质的影响比骑行更大,但仍然远小于短暂的冲刺跑。冲刺跑还改变了逾 200 种代谢物,迅速提升了参与血管生长、组织重塑和激素信号传导的蛋白质水平。部分蛋白质是通过一种名为胞外域脱落(ectodomain shedding)的快速细胞信号传导过程进入血液的——蛋白质并非新产生并释放,而是细胞表面已有的蛋白质片段被切除并迅速进入血液循环。33 种与降低心血管和代谢疾病风险相关的蛋白质有 32 种会因短暂的冲刺跑发生改变,只有 3 种会受到中等强度运动的影响。逾四分之一蛋白质还与延缓生物衰老相关。研究结果表明,运动强度可能会强烈影响释放到血液中的蛋白质和代谢物,进而影响全身组织的反应方式。
- Rockstar 向微软和 Discord 发去法庭传票以识别 GTA6 泄密者身份
2022 年 9 月一名黑客泄漏了当时尚未宣布的 GTA6 的图片和视频,此事促使开发商 Rockstar Games 加强了安全措施。然而到了 2026 年 8 月游戏还有 3 个月即将发售时,自称 CyberLeek 的个人或组织发布了 GTA6 的一系列新视频,视频显示泄密者手中可能有一个可运行的版本,也就是游戏本体被盗了。彭博社援引知情人士的消息称,Rockstar 尚未确定泄密者身份,也不知道游戏本体是如何泄漏出去的。该公司目前正全力查明泄露源头并追踪泄密者。为了识别泄密者,Rockstar 母公司 Take-Two 的律师正向法院申请传票,要求微软和 Discord 提供信息帮助识别泄密者身份。Take-Two 要求微软在 9 月 4 日前提供信息,要求 Discord 在相同的截止日期前提供 CYBERLEEK、CINEMATICROCKSTAR 和 Surfer24k™ 等账号的信息。
- 因门把手安全隐患特斯拉在华召回近 300 万辆车
特斯拉和另外 8 家汽车制造商 21 日宣布,将在中国召回总计约 430 万辆汽车,创下中国汽车召回规模纪录。此次召回的整改措施包括软件更新、加贴警示标签,以及改进门把手周围的标识等。大多数车企还将通过 OTA 远程升级软件。根据国家市场监督管理总局发布的公告,特斯拉将从 9 月 25 日起召回 298 万辆进口及中国制造的 Model 3、Model Y、Model S 和 Model X 汽车。特斯拉的召回规模最大,这也反映出该公司采用此类门把手设计的车型销量巨大。除特斯拉外,此次召回行动涉及车企包括中国一汽、北汽蓝谷、东风汽车、奇瑞、吉利、小鹏、零跑和小米。小米将召回约 39 万辆汽车,零跑约 37.1 万辆,小鹏约 26.4 万辆。零跑、小鹏和吉利此次召回的规模也均创下各自公司的历史纪录。监管机构表示,在发生严重碰撞并导致车辆电气系统失效后,机械式紧急车门解锁装置可能难以识别。车内人员可能难以打开车门逃生,救援人员也可能难以进入车内。
- 使用胁迫密码删除手机数据的美国公民被控妨碍联邦执法的重罪
2025 年 1 月,Samuel Tunick 从多米尼加共和国度假返回美国时,在亚特兰大 Hartsfield-Jackson 国际机场被拦下,美国海关和边境保护局官员要求搜查他的手机。他最终交出了手机以及一个密码,该密码删除了手机上的数据。他的 Pixel 智能手机运行的是安全加固的 Android 操作系统 GrapheneOS,它内置了被称为胁迫密码的安全功能,输入该密码后会删除手机上的数据。美国检方以妨碍联邦执法的重罪起诉了他,他因此面临最高五年的监禁。这是已知首个因输入特定密码删除设备数据而遭到起诉的案例。佐治亚州北区联邦检察官 Theodore Hertzberg 在一份声明中表示:“妨碍联邦执法是性质严重、有严重后果的罪行。任何销毁或试图销毁财产(包括数据)以阻止合法搜查和扣押的人,都应预料到会因其行为受到起诉和惩罚。”Tunick 在接受《纽约时报》采访时表示:“政府不拥有我们的数据。政府不拥有我们的通信、我们的人际关系,无论他们多么努力尝试。我们必须捍卫对隐私的基本权利;否则我们无法真正说自己生活在一个民主社会中。”
- 中国要求政府部门提前停用 Windows 10 政府版改用 Linux
彭博社报道,中国政府下令部分机构提前停止使用 Windows 10 中国政府版,改用国产 Linux 发行版。为维护数字主权,中国已不再信任美国公司的软件。微软回应彭博社的询问时表示它没有发现影响该 Windows 系统的安全事件。Windows 10 中国政府版由微软和中国电子科技集团的合资企业神州网信开发。神州网信原计划到 2027 年 2 月停止支持该版本,但其生命结束时间被提前到今年下半年。中国政府机构采用的国产 Linux 发行版可能包括了麒麟操作系统(Kylin OS)和统信 UOS。统信 UOS 桌面版源自 Deepin 和 Debian Linux。
- 中国准备发射嫦娥七号,前往月球南极寻找水冰
中国准备发射嫦娥七号,它将尝试首次直接在月球南极着陆,搜寻阴影区的陨石坑去寻找水冰。嫦娥七号使用的运载火箭为长征五号,计划从海南文昌航天发射场发射,发射窗口为 2026 年 8 月 24 日上午。嫦娥七号由一个轨道器和一个着陆器组成,而着陆器搭载了漫游的巡视器和飞跃器,其中飞跃器具备重复起飞着陆、月面飞行、月面行走功能。在阳照区完成探测并充电后,它将飞入有永久阴影区的撞击坑进行探测。嫦娥七号探测器将耗时时六天抵达月球轨道,随后将在轨道上展开为期两个月的准备工作,计划于 11 月着陆月球南极,预定着落地点为沙克尔顿撞击坑,它是一个直径 21 公里的环形山,边缘接近月球南极。月球两极被认为蕴藏了巨大的冰库,但其规模有多大、以及实际分布情况,都需要等待实地观察。
- 微软隐藏 OneDrive Photos,但该应用并未删除
微软最近被发现悄悄向 Windows 11 用户推送了一款新的照片应用 OneDrive Photos,与 OneDrive 位于同一文件夹内,无法单独卸载。事情曝光之后,微软表示这是一次意外,他们原本无意如此大范围的推送 OneDrive Photos。为了减少对用户的“曝光”,微软在开始菜单应用列表或 Windows 搜索中移除了“OneDrive Photos”,但它本身并没有删除,只是不让用户发现。
- 微软调查部分用户在安装 Windows 11 八月安全更新后遭遇游戏崩溃的报告
微软正在调查部分用户在安装 Windows 11 八月例行安全更新后遭遇游戏崩溃的报告。根据发布在 Release Health 上的声明,受影响的游戏可能会失去响应、意外关闭、引发“EXCEPTION_ACCESS_VIOLATION”错误或触发设备意外重启。不是所有游戏都受到影响,微软列出的受影响游戏包括了《ARC Raiders》、《MARVEL Tōkon: Fighting Souls》和《The Finals》。微软表示正在调查问题是否由它引起的,它请求受影响用户提供反馈。
- 混合型 T 细胞在超级百岁老人血液中显著增加
当代人类的平均寿命约为 71 岁,有少数人能迎来百岁生日,而能活过 110 岁的人则更稀有,他们被称为“超级百岁老人”。根据发表在《Cell Reports》期刊上的一项研究,日本大阪大学研究团队发现,一种罕见的免疫细胞会随着极端高龄而显著增加。这类细胞兼具识别威胁和杀伤危险细胞的能力,或有助于超级百岁老人应对随着年龄增长而增加的持续性健康威胁。随着年龄增长,一些疾病的患病风险会增加,人体抵御感染的免疫能力也会逐渐减弱。T细胞是人体免疫系统的一类重要细胞,主要分为两类:辅助性T细胞负责协调免疫反应,杀伤性T细胞则负责清除受感染或癌变细胞。研究人员发现,超级百岁老人会积累一种不同寻常的“混合型”T细胞,即CD4细胞毒性T淋巴细胞(CD4 CTL)。这类罕见细胞同时具备识别威胁和摧毁危险细胞的能力。研究人员分析了不同年龄组人群的免疫细胞,包括70—90岁人群、百岁老人以及超级百岁老人。结果发现,在生命的大部分阶段,这类细胞始终十分少见,但在接近100岁时开始显著增加。在超级百岁老人中,CD4 CTL占血液中全部T细胞的比例接近1/5,而在较年轻的研究参与者中,这一比例仅约4%。进一步分析发现,部分CD4 CTL发生了明显的克隆扩增,即少数细胞不断复制,形成了数量庞大的同源细胞群。这一现象提示,这些细胞可能长期受到某些特定抗原的反复刺激,并在持续的免疫应答过程中不断增殖。研究人员还发现,超级百岁老人这类细胞所携带的部分T细胞受体,与癌症患者肿瘤组织中的T细胞受体高度相似,但这些超级百岁老人均无癌症病史。研究人员表示,这一发现提示,这些免疫细胞可能具有识别肿瘤细胞的能力,甚至可能在肿瘤尚未发展到临床可检测阶段时,就已对其产生免疫反应。这些发现意味着,极端高龄时期的免疫变化可能并非免疫系统单纯“衰老”和“耗竭”,而更像是免疫系统为适应长期健康生存而进行的一种重新组织。
- 达斯·维达赞美 Flock 车牌跟踪系统
Flock 的车牌跟踪系统最近在美国引发了激烈争论,媒体同一时间报道了大量警官利用 Flock 摄像头跟踪女友/前女友、妻子/前妻的新闻。但在一片争论之中,皇帝最忠实的助手、西斯尊主达斯·维达则大肆赞美了 Flock。周三晚上加州圣地亚哥公共安全与宜居社区委员会会议(Public Safety and Livable Neighborhoods Committee Meeting)的公众评论期间,达斯·维达在台上说,“皇帝是 Flock 的粉丝,我们必须继续利用 Flock 技术,如此才能跟踪和监视那些叛军渣滓,看着他们从一个游乐场到另一个游乐场,从游乐场到游泳池,从游泳池到体育馆。因为我们都知道,Flock 摄像头不仅跟踪车牌;它们还跟踪孩子。它们在公园和体育馆里跟踪孩子,我们需要这个,我需要它来跟踪前女友。”
- Bilibili 进军国际市场
Bilibili 本周重新发布了国际版应用,准备推出英文版本,进军全球市场。新的国际版应用将不需要身份验证,用户无需提供护照或身份证件即可注册。Bilibili 此前已积极邀请西方知名主播如 MrBeast 在其有 3.76 亿月活用户的中文主站发布视频。更大规模的全球扩张可能会挑战 YouTube 的霸主地位,但也面临类似 TikTok 的审查、内容审核和数据安全等棘手问题。根据招聘信息,B 站正在洛杉矶、伦敦、墨西哥城、圣保罗、伊斯坦布尔和东京招聘社区经理。
OrangeBot Weekly
The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.
Free · One-click unsubscribe · No spam