OrangeBot.AI Digest — 2026-08-22
90 headlines across 8 sources, aggregated for this day.
Hacker News(15)
- hdiutil is deprecated in macOS 27 Golden Gate (lapcatsoftware.com)
- Scrap (twitter.com)
- Hister – A private, full content search index that you control (hister.org)
- typ.ing (typ.ing)
- Anthropic appears to be A/B testing reduced effort levels in Claude Code (twitter.com)
- ElevenLabs, TwelveLabs, ThirteenLabs (quantumi.sh)
- A Friendly Introduction to Racket (geometridae.bearblog.dev)
- New MCP Roadmap (blog.modelcontextprotocol.io)
- A Kantian Critique of "Sorry" by Justin Bieber (decodingvibes.com)
- Hook, hold, harvest and hide: Meta's alleged strategy laid out in first week (www.theguardian.com)
- Munder Difflin – Agent harness to run an office of your clones (munderdiffl.in)
- Statement by Prime Minister Carney on Canada-U.S. trade negotiations (www.pm.gc.ca)
- Canada will match US tariffs 'dollar for dollar' as trade talks break down (www.bbc.com)
- Early-life stress leaves a 'scar' inside brain cells in mice (medicine.washu.edu)
- There's no reason for software to be slow anymore (danluu.com)
GitHub Trending(15)
- openai / codex
- mattpocock / skills
- affaan-m / ECC
- obra / superpowers
- Wei-Shaw / sub2api
- makeplane / plane
- n8n-io / n8n
- anthropics / claude-code
- AprilNEA / OpenLogi
- modular / modular
- multica-ai / andrej-karpathy-skills
- mahlernim / google-timeline-visualizer
- ripienaar / free-for-dev
- microsoft / TypeScript
- cursor / plugins
Product Hunt(15)
- Pocket by Meta
Vibe-code games, then share them like TikToks
- SubtitleGenerator
From video to publish-ready AI subtitles—all in one browser
- Open Analytics
AI-Native Google Analytics alternative for the modern web
- Maccess
Your Mac, in your pocket — trackpad, screen, and AI
- Port Radar for macOS
An AI port manager for your Mac.
- AutoClaw
An AI work agent across desktop, browser, and chat
- VeloFiler
Keyboard-first dual-pane file manager for macOS
- Agents Never Sleep
Agents keep running with the lid closed
- Toplify
Track your App Store ranking worldwide
- KerasFormers
Keras 3 collection of pretrained models
- Pawvis
Control your Mac via camera & train gestures, local & FOSS
- Zero
Vercel's programming language built for AI agents
- ShogunAI
Your personal AGI on your PC. Built to finish real work.
- Actx0
Memory infrastructure for AI agents.
- Epho
Run Claude Code, Codex or Opencode in cloud with your repo
Hugging Face(15)
- EnvHarness: Awakening Static Worlds for Agent Learning
LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross-artifact consistency. FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, while analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment. These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.
- 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bottlenecks. On the reference-context side, conditioning on all previously generated views grows as O(N), weakening cross-view appearance guidance. On the target-context side, disjoint groups cannot directly exchange information, causing global structural drift. 4DAnyone addresses both bottlenecks with two complementary designs: Reference Context Packing (RCP) compresses growing reference views into a fixed-length mixed-resolution context with O(1) reference-context complexity, while Target Context Routing (TCR) rotates target-view groupings during denoising to share context across groups at high-noise steps and stabilize details at low-noise steps. We further build the MVGameHuman dataset using our in-house game engine and combine it with light-stage and in-the-wild video datasets for training. Experiments on DNA-Rendering and DyMVHumans show that 4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization. See our project page for video results and source code: https://4danyone.github.io.
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce SWE-bench Science, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, Claude Code with Opus-5 (max), achieves a pass@1 below 50\%, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.
- WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.
- ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.
- Repo0: Design-Driven Zero-to-All Code Generation
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.
- Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.
- Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harness---is typically treated as a fixed artifact after deployment. This work studies an alternative where the harness is task-specific and continuously evolvable: each task family maintains its own harness, which is hot-swapped across iterations through a fixed task-injection seam and rewritten using environment feedback. We introduce Hierarchical Self-Improvement (HSI), a framework in which a single frozen LLM M operates across three hierarchical scopes: a task harness H that executes tasks, an evolver that rewrites H, and a meta-evolver that rewrites the evolver's strategy code under a frozen outer anchor. A thinking-on/off design isolates the contribution of harness evolution by disabling reasoning during task execution while enabling it during self-modification. HSI is bounded by two factors: a feedback-fidelity bound, since evolution requires informative reward signals to guide selection, and a backbone capability bound, since harness redesign cannot overcome limitations of the frozen model. On BALROG with DeepSeek-V4-Flash-Preview as the frozen backbone, HSI achieves consistent gains over the initial harness on moderate-difficulty tasks (+39.3 on BabyAI, +33.0 on Crafter, +25.0 on TextWorld, and +15.0 on MiniHack, all in raw \% Progress), while obtaining strong held-out generalization on BabaIsAI sub-suites (0.98 best-test on BreakStop and 1.00 on GoTo from a 20% unseen split). On tasks beyond the backbone's capability (NLE), harness evolution provides no improvement. These results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits. Code is available at https://github.com/TailinZhou/hsi.
- Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.
- EXIMO: VLM Guided Exploration of VLA Policies
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular, collecting teleoperation datasets requires hundreds of hours of expensive human labour and the alternative, reinforcement learning (RL), can be notoriously sample-inefficient especially for long-horizon tasks. In addition, RL with VLAs imposes several challenges due to the model's size and architectural design. In this work, we propose EXIMO, an efficient algorithm for finetuning of VLA policies. EXIMO operates in three stages: explore, imitate, and optimize. During the explore phase, EXIMO equips the VLA with a vision language model (VLM) that acts as a planner. The VLM thinks and breaks down challenging long-horizon problems into shorter ones for the VLA. The VLM, together with the VLA, is used to collect an orchestrated dataset on new tasks. During the imitate phase, the VLA is finetuned with the orchestrated data. Finally, during the optimize stage, we use residual off-policy RL to further finetune the policy. In our experiments, we ablate all three stages of EXIMO and show that it outperforms existing approaches significantly in terms of sample-efficiency and final performance.
- τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
Techmeme(15)
- AI agents' growing capabilities are driving productivity FOMO among some startup founders, who feel compelled to work long hours managing and guiding the agents (Katherine Bindley/Wall Street Journal)
Katherine Bindley / Wall Street Journal : AI agents' growing capabilities are driving productivity FOMO among some startup founders, who feel compelled to work long hours managing and guiding the agents — The growing capabilities of AI give new meaning to working yourself to the bone — Seductive. Intoxicating. All-consuming.
- London-based Inherent, founded by DeepMind alumni and with $50M in seed funding, says its new Faraday agent beats GPT-5.5 at reproducing research paper findings (Anna Heim/TechCrunch)
Anna Heim / TechCrunch : London-based Inherent, founded by DeepMind alumni and with $50M in seed funding, says its new Faraday agent beats GPT-5.5 at reproducing research paper findings — Inherent, a London AI lab founded by Google DeepMind alumni, says its AI agent just outperformed much larger models from Anthropic and OpenAI using a fraction of the size.
- Sources: some of Nvidia's top customers have been told that prices will jump 15%+ on systems, including Vera Rubin and Grace Blackwell, starting in early 2027 (Bloomberg)
Bloomberg : Sources: some of Nvidia's top customers have been told that prices will jump 15%+ on systems, including Vera Rubin and Grace Blackwell, starting in early 2027 — Some of Nvidia Corp.'s biggest customers have been told that the prices of servers containing its artificial intelligence chips …
- Apparel retailers like Zalando, Zara, and ASOS are betting on AI virtual fitting rooms to create a better online shopping experience and cut costly returns (Sonja Wind/Bloomberg)
Sonja Wind / Bloomberg : Apparel retailers like Zalando, Zara, and ASOS are betting on AI virtual fitting rooms to create a better online shopping experience and cut costly returns — Buying clothes online can leave shoppers frustrated and retailers drowning in returned goods. Now brands hope that AI-powered tools will deliver the perfect fit.
- Carrier Pidge hit 75K users and Roost 650K downloads, as slow messaging apps delivering texts at pigeon speeds attract users tired of constant notifications (Emmett Lindner/New York Times)
Emmett Lindner / New York Times : Carrier Pidge hit 75K users and Roost 650K downloads, as slow messaging apps delivering texts at pigeon speeds attract users tired of constant notifications — Apps that let users send digital messages at ultraslow speeds have taken off, with users reveling in a slower pace of life.
- A look at the narrowing US-China AI gap, as a spate of compelling, low-cost releases makes Chinese AI models increasingly attractive to businesses (Bloomberg)
Bloomberg : A look at the narrowing US-China AI gap, as a spate of compelling, low-cost releases makes Chinese AI models increasingly attractive to businesses — A spate of compelling releases at budget prices have made China the frontrunner in the race for global adoption.
- Ox Alpha, a "stealth model" from an unknown AI lab with a 1M-token multimodal context and capacity for 100T tokens/day, goes viral after launching on OpenRouter (Rohail Saleem/Wccftech)
Rohail Saleem / Wccftech : Ox Alpha, a “stealth model” from an unknown AI lab with a 1M-token multimodal context and capacity for 100T tokens/day, goes viral after launching on OpenRouter — When an unknown AI lab drops an anonymous model - Ox Alpha - for free, while declaring that they have the capacity …
- Cheap energy, abundant land, and proximity to Beijing have made Ulanqab, Inner Mongolia, a data center hub, with ~100 data centers built or under construction (Zeyi Yang/Wired)
Zeyi Yang / Wired : Cheap energy, abundant land, and proximity to Beijing have made Ulanqab, Inner Mongolia, a data center hub, with ~100 data centers built or under construction — Cheap energy, abundant land, and proximity to Beijing have turned a city in Inner Mongolia into a crucial hub for data centers.
- Inside the World Robot Conference in Beijing, drawing over 300 exhibitors; Unitree founder Wang Xingxing said the industry's "ChatGPT moment" has yet to come (Financial Times)
Financial Times : Inside the World Robot Conference in Beijing, drawing over 300 exhibitors; Unitree founder Wang Xingxing said the industry's “ChatGPT moment” has yet to come — Beijing policymakers have made robotics a ‘strategic priority’ — At the World Robot Conference in Beijing this week …
- With each successive era of LLMs, from early scaling, to reasoning, to agentic, open models have taken half as long to catch up to the first closed model (SemiAnalysis)
SemiAnalysis : With each successive era of LLMs, from early scaling, to reasoning, to agentic, open models have taken half as long to catch up to the first closed model — Comparing open vs. closed models across the eras of frontier models, Is the gap narrowing? — ∙ Paid
- Rundoo, a provider of AI-powered business operations software for independent supply stores, raised a $30M Series B led by Battery Ventures (Mike Wheatley/SiliconANGLE)
Mike Wheatley / SiliconANGLE : Rundoo, a provider of AI-powered business operations software for independent supply stores, raised a $30M Series B led by Battery Ventures — Rundoo Inc., the creator of an artificial intelligence-native system-of-record for independent supply stores, said today it has closed …
- OpenAI President Greg Brockman's role has expanded significantly, giving him control over its product and scaling teams following a wave of executive departures (Hayden Field/The Verge)
Hayden Field / The Verge : OpenAI President Greg Brockman's role has expanded significantly, giving him control over its product and scaling teams following a wave of executive departures — OpenAI has had a hell of a year. The company spent months battling former cofounder Elon Musk in a sensational jury trial …
- Following abandoned IPO attempts in NYC and London, Shein is struggling to grow as it nears a Hong Kong listing at a fraction of its peak valuation of $100B (Sui-Lee Wee/New York Times)
Sui-Lee Wee / New York Times : Following abandoned IPO attempts in NYC and London, Shein is struggling to grow as it nears a Hong Kong listing at a fraction of its peak valuation of $100B — Facing pressure in the United States and Europe, Shein is struggling to find new ways to grow ahead of its much-delayed initial public offering.
- How politicians who once championed data centers, including Greg Abbott and Josh Shapiro, are now slowing their development as the issue becomes a liability (Wall Street Journal)
Wall Street Journal : How politicians who once championed data centers, including Greg Abbott and Josh Shapiro, are now slowing their development as the issue becomes a liability — State governors, including Greg Abbott and Josh Shapiro, slowing development of the facilities as anger over AI spreads
- Sources: Anthropic's bankers said the company could raise $100B+ in its IPO, which could value it at $2T, in recent discussions with potential investors (New York Times)
New York Times : Sources: Anthropic's bankers said the company could raise $100B+ in its IPO, which could value it at $2T, in recent discussions with potential investors — The offering could value the five-year-old A.I. start-up at $2 trillion, its bankers have told potential investors, which would exceed Elon Musk's SpaceX.
Solidot(15)
- 3 分钟冲刺跑产生的分子反应与 90 分钟中等强度运动截然不同
3 分钟冲刺跑产生的分子反应与 90 分钟中等强度运动截然不同。洛克菲勒大学的研究人员比较了人体对不同强度运动的反应。他们发现,六次 30 秒全力冲刺跑后,血液中近四分之一的蛋白质发生了变化。相比之下,90 分钟持续中等强度骑行仅改变了不到 0.25% 的蛋白质。中等强度的跑步机运动对蛋白质的影响比骑行更大,但仍然远小于短暂的冲刺跑。冲刺跑还改变了逾 200 种代谢物,迅速提升了参与血管生长、组织重塑和激素信号传导的蛋白质水平。部分蛋白质是通过一种名为胞外域脱落(ectodomain shedding)的快速细胞信号传导过程进入血液的——蛋白质并非新产生并释放,而是细胞表面已有的蛋白质片段被切除并迅速进入血液循环。33 种与降低心血管和代谢疾病风险相关的蛋白质有 32 种会因短暂的冲刺跑发生改变,只有 3 种会受到中等强度运动的影响。逾四分之一蛋白质还与延缓生物衰老相关。研究结果表明,运动强度可能会强烈影响释放到血液中的蛋白质和代谢物,进而影响全身组织的反应方式。
- Rockstar 向微软和 Discord 发去法庭传票以识别 GTA6 泄密者身份
2022 年 9 月一名黑客泄漏了当时尚未宣布的 GTA6 的图片和视频,此事促使开发商 Rockstar Games 加强了安全措施。然而到了 2026 年 8 月游戏还有 3 个月即将发售时,自称 CyberLeek 的个人或组织发布了 GTA6 的一系列新视频,视频显示泄密者手中可能有一个可运行的版本,也就是游戏本体被盗了。彭博社援引知情人士的消息称,Rockstar 尚未确定泄密者身份,也不知道游戏本体是如何泄漏出去的。该公司目前正全力查明泄露源头并追踪泄密者。为了识别泄密者,Rockstar 母公司 Take-Two 的律师正向法院申请传票,要求微软和 Discord 提供信息帮助识别泄密者身份。Take-Two 要求微软在 9 月 4 日前提供信息,要求 Discord 在相同的截止日期前提供 CYBERLEEK、CINEMATICROCKSTAR 和 Surfer24k™ 等账号的信息。
- 因门把手安全隐患特斯拉在华召回近 300 万辆车
特斯拉和另外 8 家汽车制造商 21 日宣布,将在中国召回总计约 430 万辆汽车,创下中国汽车召回规模纪录。此次召回的整改措施包括软件更新、加贴警示标签,以及改进门把手周围的标识等。大多数车企还将通过 OTA 远程升级软件。根据国家市场监督管理总局发布的公告,特斯拉将从 9 月 25 日起召回 298 万辆进口及中国制造的 Model 3、Model Y、Model S 和 Model X 汽车。特斯拉的召回规模最大,这也反映出该公司采用此类门把手设计的车型销量巨大。除特斯拉外,此次召回行动涉及车企包括中国一汽、北汽蓝谷、东风汽车、奇瑞、吉利、小鹏、零跑和小米。小米将召回约 39 万辆汽车,零跑约 37.1 万辆,小鹏约 26.4 万辆。零跑、小鹏和吉利此次召回的规模也均创下各自公司的历史纪录。监管机构表示,在发生严重碰撞并导致车辆电气系统失效后,机械式紧急车门解锁装置可能难以识别。车内人员可能难以打开车门逃生,救援人员也可能难以进入车内。
- 使用胁迫密码删除手机数据的美国公民被控妨碍联邦执法的重罪
2025 年 1 月,Samuel Tunick 从多米尼加共和国度假返回美国时,在亚特兰大 Hartsfield-Jackson 国际机场被拦下,美国海关和边境保护局官员要求搜查他的手机。他最终交出了手机以及一个密码,该密码删除了手机上的数据。他的 Pixel 智能手机运行的是安全加固的 Android 操作系统 GrapheneOS,它内置了被称为胁迫密码的安全功能,输入该密码后会删除手机上的数据。美国检方以妨碍联邦执法的重罪起诉了他,他因此面临最高五年的监禁。这是已知首个因输入特定密码删除设备数据而遭到起诉的案例。佐治亚州北区联邦检察官 Theodore Hertzberg 在一份声明中表示:“妨碍联邦执法是性质严重、有严重后果的罪行。任何销毁或试图销毁财产(包括数据)以阻止合法搜查和扣押的人,都应预料到会因其行为受到起诉和惩罚。”Tunick 在接受《纽约时报》采访时表示:“政府不拥有我们的数据。政府不拥有我们的通信、我们的人际关系,无论他们多么努力尝试。我们必须捍卫对隐私的基本权利;否则我们无法真正说自己生活在一个民主社会中。”
- 中国要求政府部门提前停用 Windows 10 政府版改用 Linux
彭博社报道,中国政府下令部分机构提前停止使用 Windows 10 中国政府版,改用国产 Linux 发行版。为维护数字主权,中国已不再信任美国公司的软件。微软回应彭博社的询问时表示它没有发现影响该 Windows 系统的安全事件。Windows 10 中国政府版由微软和中国电子科技集团的合资企业神州网信开发。神州网信原计划到 2027 年 2 月停止支持该版本,但其生命结束时间被提前到今年下半年。中国政府机构采用的国产 Linux 发行版可能包括了麒麟操作系统(Kylin OS)和统信 UOS。统信 UOS 桌面版源自 Deepin 和 Debian Linux。
- 中国准备发射嫦娥七号,前往月球南极寻找水冰
中国准备发射嫦娥七号,它将尝试首次直接在月球南极着陆,搜寻阴影区的陨石坑去寻找水冰。嫦娥七号使用的运载火箭为长征五号,计划从海南文昌航天发射场发射,发射窗口为 2026 年 8 月 24 日上午。嫦娥七号由一个轨道器和一个着陆器组成,而着陆器搭载了漫游的巡视器和飞跃器,其中飞跃器具备重复起飞着陆、月面飞行、月面行走功能。在阳照区完成探测并充电后,它将飞入有永久阴影区的撞击坑进行探测。嫦娥七号探测器将耗时时六天抵达月球轨道,随后将在轨道上展开为期两个月的准备工作,计划于 11 月着陆月球南极,预定着落地点为沙克尔顿撞击坑,它是一个直径 21 公里的环形山,边缘接近月球南极。月球两极被认为蕴藏了巨大的冰库,但其规模有多大、以及实际分布情况,都需要等待实地观察。
- 微软隐藏 OneDrive Photos,但该应用并未删除
微软最近被发现悄悄向 Windows 11 用户推送了一款新的照片应用 OneDrive Photos,与 OneDrive 位于同一文件夹内,无法单独卸载。事情曝光之后,微软表示这是一次意外,他们原本无意如此大范围的推送 OneDrive Photos。为了减少对用户的“曝光”,微软在开始菜单应用列表或 Windows 搜索中移除了“OneDrive Photos”,但它本身并没有删除,只是不让用户发现。
- 微软调查部分用户在安装 Windows 11 八月安全更新后遭遇游戏崩溃的报告
微软正在调查部分用户在安装 Windows 11 八月例行安全更新后遭遇游戏崩溃的报告。根据发布在 Release Health 上的声明,受影响的游戏可能会失去响应、意外关闭、引发“EXCEPTION_ACCESS_VIOLATION”错误或触发设备意外重启。不是所有游戏都受到影响,微软列出的受影响游戏包括了《ARC Raiders》、《MARVEL Tōkon: Fighting Souls》和《The Finals》。微软表示正在调查问题是否由它引起的,它请求受影响用户提供反馈。
- 混合型 T 细胞在超级百岁老人血液中显著增加
当代人类的平均寿命约为 71 岁,有少数人能迎来百岁生日,而能活过 110 岁的人则更稀有,他们被称为“超级百岁老人”。根据发表在《Cell Reports》期刊上的一项研究,日本大阪大学研究团队发现,一种罕见的免疫细胞会随着极端高龄而显著增加。这类细胞兼具识别威胁和杀伤危险细胞的能力,或有助于超级百岁老人应对随着年龄增长而增加的持续性健康威胁。随着年龄增长,一些疾病的患病风险会增加,人体抵御感染的免疫能力也会逐渐减弱。T细胞是人体免疫系统的一类重要细胞,主要分为两类:辅助性T细胞负责协调免疫反应,杀伤性T细胞则负责清除受感染或癌变细胞。研究人员发现,超级百岁老人会积累一种不同寻常的“混合型”T细胞,即CD4细胞毒性T淋巴细胞(CD4 CTL)。这类罕见细胞同时具备识别威胁和摧毁危险细胞的能力。研究人员分析了不同年龄组人群的免疫细胞,包括70—90岁人群、百岁老人以及超级百岁老人。结果发现,在生命的大部分阶段,这类细胞始终十分少见,但在接近100岁时开始显著增加。在超级百岁老人中,CD4 CTL占血液中全部T细胞的比例接近1/5,而在较年轻的研究参与者中,这一比例仅约4%。进一步分析发现,部分CD4 CTL发生了明显的克隆扩增,即少数细胞不断复制,形成了数量庞大的同源细胞群。这一现象提示,这些细胞可能长期受到某些特定抗原的反复刺激,并在持续的免疫应答过程中不断增殖。研究人员还发现,超级百岁老人这类细胞所携带的部分T细胞受体,与癌症患者肿瘤组织中的T细胞受体高度相似,但这些超级百岁老人均无癌症病史。研究人员表示,这一发现提示,这些免疫细胞可能具有识别肿瘤细胞的能力,甚至可能在肿瘤尚未发展到临床可检测阶段时,就已对其产生免疫反应。这些发现意味着,极端高龄时期的免疫变化可能并非免疫系统单纯“衰老”和“耗竭”,而更像是免疫系统为适应长期健康生存而进行的一种重新组织。
- 达斯·维达赞美 Flock 车牌跟踪系统
Flock 的车牌跟踪系统最近在美国引发了激烈争论,媒体同一时间报道了大量警官利用 Flock 摄像头跟踪女友/前女友、妻子/前妻的新闻。但在一片争论之中,皇帝最忠实的助手、西斯尊主达斯·维达则大肆赞美了 Flock。周三晚上加州圣地亚哥公共安全与宜居社区委员会会议(Public Safety and Livable Neighborhoods Committee Meeting)的公众评论期间,达斯·维达在台上说,“皇帝是 Flock 的粉丝,我们必须继续利用 Flock 技术,如此才能跟踪和监视那些叛军渣滓,看着他们从一个游乐场到另一个游乐场,从游乐场到游泳池,从游泳池到体育馆。因为我们都知道,Flock 摄像头不仅跟踪车牌;它们还跟踪孩子。它们在公园和体育馆里跟踪孩子,我们需要这个,我需要它来跟踪前女友。”
- Bilibili 进军国际市场
Bilibili 本周重新发布了国际版应用,准备推出英文版本,进军全球市场。新的国际版应用将不需要身份验证,用户无需提供护照或身份证件即可注册。Bilibili 此前已积极邀请西方知名主播如 MrBeast 在其有 3.76 亿月活用户的中文主站发布视频。更大规模的全球扩张可能会挑战 YouTube 的霸主地位,但也面临类似 TikTok 的审查、内容审核和数据安全等棘手问题。根据招聘信息,B 站正在洛杉矶、伦敦、墨西哥城、圣保罗、伊斯坦布尔和东京招聘社区经理。
- 天文学家发现银河系已知最快的恒星
天文学家发现了银河系已知运行速度最快的恒星。这颗名为 S301 的恒星围绕银河系中心的超大质量黑洞——人马座A*运行,最快速度达到每秒 2.5 万公里,超过光速的 8%。它的运行轨道非常接近人马座A*,其运动有望帮助科学家首次直接测量大质量黑洞的自转,并为检验爱因斯坦广义相对论提供新的机会。S301 绕人马座A* 公转周期为 8.7年。在它距离人马座 A*最近时——类似于太阳到土星的距离——恒星的运行速度超过光速的 8%。研究人员认为,S301的轨道特征以及恒星无法在如此靠近超大质量黑洞的位置形成,表明它很可能原本属于一个双星系统。当这个双星系统靠近人马座A*时,黑洞强大的潮汐力将两颗恒星撕裂,其中一颗被黑洞引力捕获,成为如今的 S301;另一颗则被高速抛出,其速度可能高到足以逃离银河系。
- Thunderbird 跟随 Firefox 采用双周发布模式
Mozilla 工程总监 Sylvestre Ledru 上月宣布,从 2026 年 9 月起 Firefox 桌面版和 Android 版本的发布周期从 4 周减少到 2 周。本周释出的 Firefox 154 是最后一个按四周发布模式释出的版本,九月初释出的 v155 则是第一个双周发布版本。由 Mozilla 子公司 MZLA 开发的开源邮件客户端 Thunderbird 宣布也将采用双周发布模式。MZLA 的 Corey Bryant 称,从 9 月起 Thunderbird 采用相同的更新频率。
- 太阳能扩张政策与鸟类多样性下降相关
南京信息工程大学的研究人员在《科学》上发表研究报告,称全球对太阳能发展的推动可能带来隐性的损害生物多样性的代价。可再生能源的扩张有助于应对气候变化,但大规模太阳能开发也可能因栖息地改变或破碎化而导致生物多样性丧失,从而引发新的环境得失权衡。研究人员汇编了一个大型数据集,它涵盖了 2014 年至 2023 年中国的 2344 个县。他们的数据集整合了鸟类观测数据、太阳能政策、环境条件和社会经济信息。他们还考察了土地利用、植被状况和农业生产力变化所带来的影响。研究结果表明,太阳能扩张政策的力度加大与鸟类多样性的显著下降有关:政策强度每增加一个标准差,鸟类生物多样性指数便会下降 2.10%。这些影响在较富裕地区和非沙漠地区最为显著,且对地理分布广泛的物种影响尤为严重。这主要应归因于土地的迅速转化,特别是将农田和草地转化为开发区,后者降低了植被的多样性。
- 海冰消失巨型鲸鱼进入格陵兰
由于海冰融化,巨型鲸鱼如座头鲸进入到了以前难以抵达的东格陵兰沿海地区。这是东格陵兰海洋生态系统发生重大转变的一部分。直到 2006 年该地区才首次记录到座头鲸的踪迹。2007 年记录到了 7 头座头鲸,2024 年船载设备就记录到了 150 头。研究人员结合卫星标记鲸鱼的追踪轨迹和因纽特猎人的证词,估计 2024 年夏天大约有 4000 头座头鲸、6000 头长须鲸和 6000 头小须鲸造访了格陵兰海。这三种鲸鱼在夏季觅食季节至少会消耗 80 万吨鱼类和磷虾。北极原有的鲸鱼要么被迫适应要么被迫离开。
OrangeBot Weekly
The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.
Free · One-click unsubscribe · No spam