Curated by Shen Huang · 90 stories · ~14 min read
DIGEST · 2026-08-21

OrangeBot.AI Digest — 2026-08-21

90 headlines across 8 sources, aggregated for this day.

Hacker News(15)

  1. AI boosted homework scores, then exam scores dropped: study (www.economist.com)
  2. Kobo can run apps now (bandarlabs.github.io)
  3. Felony Bench (www.felonybench.com)
  4. AI boosted homework scores, then exam scores dropped: Study (canews24.online)
  5. Claudette: Make Claude stop talking like a BuzzFeed article (github.com)
  6. New Worlds: We are living in the future of J.G. Ballard or William Gibson (precastreinforced.co.uk)
  7. I accidentally logged hundreds of thousands of phone calls to military bases (lina.sh)
  8. Kagi added a setting for removing paywalled links from search results (kagi.com)
  9. I'm becoming AI-blind (cymerys.com)
  10. Grand jury declines to indict Ohio man charged with destroying Flock camera (san.com)
  11. Felony charges for citizen deleting phone data at US Border (www.nytimes.com)
  12. AI companies destroy physical books – let's scan rare books before it's too late (annas-archive.pk)
  13. DeepSeek-v4-flash-vision-exp (api-docs.deepseek.com)
  14. Small, native web tricks worth remembering (htmlcat.net)
  15. The Lost Treasure of Sid Meier's Pirates (remapradio.com)

GitHub Trending(15)

  1. mattpocock / skills
  2. mahlernim / google-timeline-visualizer
  3. harry0703 / MoneyPrinterTurbo
  4. AprilNEA / OpenLogi
  5. PostHog / posthog
  6. microsoft / TypeScript
  7. obra / superpowers
  8. santifer / career-ops
  9. cursor / plugins
  10. modular / modular
  11. affaan-m / ECC
  12. TryGhost / Ghost
  13. ruvnet / ruflo
  14. apache / maka
  15. protocolbuffers / protobuf

Product Hunt(15)

  1. Wizstar

    Digital avatars that move and act like professional actors

  2. Mindcase

    Extract data from anywhere on the web within minutes

  3. fx (by Vercel)

    Vercel's tiny, open-source coding agent

  4. Supernova

    All your data in Claude and Codex

  5. Actx0

    Memory infrastructure for AI agents.

  6. Epho

    Run Claude Code, Codex or Opencode in cloud with your repo

  7. Dockhand

    Docker management for everyone

  8. Project SKY

    Your ambient AI companion for Windows.

  9. ShogunAI

    Your personal AGI on your PC. Built to finish real work.

  10. Flunkey

    Voice-first AI layer for Windows (beta)

  11. Router by Ramp

    Tokens are money. Save both.

  12. OneCLI

    Give every employee a secured, sandboxed pro assistant agent

  13. Outlook Google Calendar Sync for Mac

    Sync Outlook calendars to Google on your Mac

  14. Plow Latch

    Run AI agents on your Mac with scoped access

  15. PixelRead AI OCR

    Capture, translate, and understand any text on your Mac

Hugging Face(15)

  1. EnvHarness: Awakening Static Worlds for Agent Learning

    LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.

  2. FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

    Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross-artifact consistency. FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, while analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment. These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.

  3. 4DAnyone: Create Anyone in 4D from a Casual Monocular Video

    We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bottlenecks. On the reference-context side, conditioning on all previously generated views grows as O(N), weakening cross-view appearance guidance. On the target-context side, disjoint groups cannot directly exchange information, causing global structural drift. 4DAnyone addresses both bottlenecks with two complementary designs: Reference Context Packing (RCP) compresses growing reference views into a fixed-length mixed-resolution context with O(1) reference-context complexity, while Target Context Routing (TCR) rotates target-view groupings during denoising to share context across groups at high-noise steps and stabilize details at low-noise steps. We further build the MVGameHuman dataset using our in-house game engine and combine it with light-stage and in-the-wild video datasets for training. Experiments on DNA-Rendering and DyMVHumans show that 4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization. See our project page for video results and source code: https://4danyone.github.io.

  4. SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

    Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce SWE-bench Science, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, Claude Code with Opus-5 (max), achieves a pass@1 below 50\%, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.

  5. WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

    Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.

  6. MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

    Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.

  7. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

    Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.

  8. ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

    Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.

  9. Repo0: Design-Driven Zero-to-All Code Generation

    Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.

  10. FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

    Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.

  11. Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

    Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.

  12. Towards Quantifying Benchmark Optimization in ASR Models

    Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

  13. Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

    Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.

  14. EXIMO: VLM Guided Exploration of VLA Policies

    How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular, collecting teleoperation datasets requires hundreds of hours of expensive human labour and the alternative, reinforcement learning (RL), can be notoriously sample-inefficient especially for long-horizon tasks. In addition, RL with VLAs imposes several challenges due to the model's size and architectural design. In this work, we propose EXIMO, an efficient algorithm for finetuning of VLA policies. EXIMO operates in three stages: explore, imitate, and optimize. During the explore phase, EXIMO equips the VLA with a vision language model (VLM) that acts as a planner. The VLM thinks and breaks down challenging long-horizon problems into shorter ones for the VLA. The VLM, together with the VLA, is used to collect an orchestrated dataset on new tasks. During the imitate phase, the VLA is finetuned with the orchestrated data. Finally, during the optimize stage, we use residual off-policy RL to further finetune the policy. In our experiments, we ablate all three stages of EXIMO and show that it outperforms existing approaches significantly in terms of sample-efficiency and final performance.

  15. Chain-of-Experience for Continual LLM Improvement

    Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.

Techmeme(15)

  1. Sources: Devoted Health, which uses AI to help coordinate medical care for those enrolled in Medicare Advantage, is raising new funding at a $25B valuation (Katie Roof/Business Insider)

    Katie Roof / Business Insider : Sources: Devoted Health, which uses AI to help coordinate medical care for those enrolled in Medicare Advantage, is raising new funding at a $25B valuation —  Devoted Health is the latest startup to show that combining health and AI commands a premium valuation.

  2. Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels (Terry Chen/NVIDIA Technical Blog)

    Terry Chen / NVIDIA Technical Blog : Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels —  The research project elevates Claude Opus 5 from a 30% model baseline to 100% as part of the complete AVO agent system, showing that system design …

  3. Anthropic hires Amir Salek, who ran Google's TPU business until 2022, to join its compute team as part of a push to develop its own chips (Dina Bass/Bloomberg)

    Dina Bass / Bloomberg : Anthropic hires Amir Salek, who ran Google's TPU business until 2022, to join its compute team as part of a push to develop its own chips —  Anthropic PBC has hired Amir Salek, a founder of the custom chip program at Alphabet Inc.'s Google, as the AI lab lays the groundwork for a push into making its own semiconductors.

  4. DHH launches Omacom Foundation, a non-profit supporting his Omarchy Linux distro, with $8M in funding from Michael Dell, Jack Dorsey, Patrick Collison, others (Omarchy News)

    Omarchy News : DHH launches Omacom Foundation, a non-profit supporting his Omarchy Linux distro, with $8M in funding from Michael Dell, Jack Dorsey, Patrick Collison, others —  It's time to dream big.  Omarchy Quattro has given people a chance to experience what the malleable computer of the future looks like, and they like it (a lot!).

  5. Source: OpenAI's VP of sales in the Americas, Kaylin Voss, resigned a week after Denise Dresser left, prompting others on the sales team to consider resigning (Laura Bratton/The Information)

    Laura Bratton / The Information : Source: OpenAI's VP of sales in the Americas, Kaylin Voss, resigned a week after Denise Dresser left, prompting others on the sales team to consider resigning —  Kaylin Voss, OpenAI's vice president of sales in the Americas, has resigned a week after her former boss, Denise Dresser …

  6. The US DOJ and TikTok reach a $400M settlement to resolve allegations TikTok violated COPPA; the DOJ filed the lawsuit in 2024 (Ashley Gold/Axios)

    Ashley Gold / Axios : The US DOJ and TikTok reach a $400M settlement to resolve allegations TikTok violated COPPA; the DOJ filed the lawsuit in 2024 —  The Justice Department and TikTok, along with parent company ByteDance, have reached a $400 million settlement to resolve allegations of violating children's online privacy laws …

  7. Sources: Apple is cutting 200+ jobs, including ~100 positions from the Vision Pro unit and another 100 from the Siri team, as it focuses on new devices and AI (Mark Gurman/Bloomberg)

    Mark Gurman / Bloomberg : Sources: Apple is cutting 200+ jobs, including ~100 positions from the Vision Pro unit and another 100 from the Siri team, as it focuses on new devices and AI —  Apple Inc. is cutting jobs across teams responsible for the Siri digital assistant and the Vision Pro headset …

  8. Anthropic says Mythos 5 is now in public beta in Claude Security for Enterprise users, and it is working with providers to embed Mythos 5 in defensive tools (Claude)

    Claude : Anthropic says Mythos 5 is now in public beta in Claude Security for Enterprise users, and it is working with providers to embed Mythos 5 in defensive tools —  No items found.  —  Reading time  —  min  —  https://claude.com/blog/bringing- claude-mythos-5-to-more-defenders

  9. Sources: London-based AI infrastructure startup Nscale is seeking to raise as much as $3B in its US IPO, which could take place as soon as September (Bailey Lipschultz/Bloomberg)

    Bailey Lipschultz / Bloomberg : Sources: London-based AI infrastructure startup Nscale is seeking to raise as much as $3B in its US IPO, which could take place as soon as September —  Nscale is seeking to raise as much as $3 billion in its US IPO, according to people familiar with the matter, joining a rush …

  10. Sources: the US Department of Energy is investigating whether Chinese lidar sensors might pose a security risk if they become widely used on vehicles in the US (Sean O'Kane/TechCrunch)

    Sean O'Kane / TechCrunch : Sources: the US Department of Energy is investigating whether Chinese lidar sensors might pose a security risk if they become widely used on vehicles in the US —  A Department of Energy lab is investigating whether Chinese lidar sensors might pose a security risk if they become widely used …

  11. Walmart, which has long resisted contactless Tap to Pay payments, says its stores will get the tech, which supports Apple Pay and Google Pay, by the end of 2026 (Sarah Perez/TechCrunch)

    Sarah Perez / TechCrunch : Walmart, which has long resisted contactless Tap to Pay payments, says its stores will get the tech, which supports Apple Pay and Google Pay, by the end of 2026 —  Apparently, hell has frozen over.  Walmart on Friday said it will finally accept payments via both Apple Pay and Google Pay at its stores …

  12. Filings: Apple paid Ireland $17B in taxes in 2025, 40% of its $43B in global corporate income tax total, after an EU court ordered it to pay €13B in back taxes (Jamie John/Financial Times)

    Jamie John / Financial Times : Filings: Apple paid Ireland $17B in taxes in 2025, 40% of its $43B in global corporate income tax total, after an EU court ordered it to pay €13B in back taxes —  New filings highlight iPhone maker's global tax liabilities  —  Apple paid Ireland $17bn in taxes last year …

  13. A Dutch regulator fines Uber €825M for deactivating driver accounts through automated systems without informing them, in the second largest fine under GDPR (Financial Times)

    Financial Times : A Dutch regulator fines Uber €825M for deactivating driver accounts through automated systems without informing them, in the second largest fine under GDPR —  Regulator says ride-hailing group deactivated driver accounts through automated systems without adequately informing them

  14. Space data center startup Starcloud raised a $250M extension, at a $2.3B valuation, to its March $170M Series A; source: Nvidia invested $25M (Tim Fernholz/TechCrunch)

    Tim Fernholz / TechCrunch : Space data center startup Starcloud raised a $250M extension, at a $2.3B valuation, to its March $170M Series A; source: Nvidia invested $25M —  Starcloud, a startup developing satellites that can perform AI inference in orbit, told TechCrunch that it has added a $250 million extension to its March $170 million Series A funding round.

  15. Kakao plans to spin off its chat app-based platform business into a company tentatively named KakaoAI and relist it on the Korea Exchange on January 27, 2027 (Reuters)

    Reuters : Kakao plans to spin off its chat app-based platform business into a company tentatively named KakaoAI and relist it on the Korea Exchange on January 27, 2027 —  South Korea's dominant chat app operator Kakao Corp (035720.KS) said on Friday it plans to spin off its chat app-based platform business …

Solidot(15)

  1. 中国准备发射嫦娥七号,前往月球南极寻找水冰

    中国准备发射嫦娥七号,它将尝试首次直接在月球南极着陆,搜寻阴影区的陨石坑去寻找水冰。嫦娥七号使用的运载火箭为长征五号,计划从海南文昌航天发射场发射,发射窗口为 2026 年 8 月 24 日上午。嫦娥七号由一个轨道器和一个着陆器组成,而着陆器搭载了漫游的巡视器和飞跃器,其中飞跃器具备重复起飞着陆、月面飞行、月面行走功能。在阳照区完成探测并充电后,它将飞入有永久阴影区的撞击坑进行探测。嫦娥七号探测器将耗时时六天抵达月球轨道,随后将在轨道上展开为期两个月的准备工作,计划于 11 月着陆月球南极,预定着落地点为沙克尔顿撞击坑,它是一个直径 21 公里的环形山,边缘接近月球南极。月球两极被认为蕴藏了巨大的冰库,但其规模有多大、以及实际分布情况,都需要等待实地观察。

  2. 微软隐藏 OneDrive Photos,但该应用并未删除

    微软最近被发现悄悄向 Windows 11 用户推送了一款新的照片应用 OneDrive Photos,与 OneDrive 位于同一文件夹内,无法单独卸载。事情曝光之后,微软表示这是一次意外,他们原本无意如此大范围的推送 OneDrive Photos。为了减少对用户的“曝光”,微软在开始菜单应用列表或 Windows 搜索中移除了“OneDrive Photos”,但它本身并没有删除,只是不让用户发现。

  3. 微软调查部分用户在安装 Windows 11 八月安全更新后遭遇游戏崩溃的报告

    微软正在调查部分用户在安装 Windows 11 八月例行安全更新后遭遇游戏崩溃的报告。根据发布在 Release Health 上的声明,受影响的游戏可能会失去响应、意外关闭、引发“EXCEPTION_ACCESS_VIOLATION”错误或触发设备意外重启。不是所有游戏都受到影响,微软列出的受影响游戏包括了《ARC Raiders》、《MARVEL Tōkon: Fighting Souls》和《The Finals》。微软表示正在调查问题是否由它引起的,它请求受影响用户提供反馈。

  4. 混合型 T 细胞在超级百岁老人血液中显著增加

    当代人类的平均寿命约为 71 岁,有少数人能迎来百岁生日,而能活过 110 岁的人则更稀有,他们被称为“超级百岁老人”。根据发表在《Cell Reports》期刊上的一项研究,日本大阪大学研究团队发现,一种罕见的免疫细胞会随着极端高龄而显著增加。这类细胞兼具识别威胁和杀伤危险细胞的能力,或有助于超级百岁老人应对随着年龄增长而增加的持续性健康威胁。随着年龄增长,一些疾病的患病风险会增加,人体抵御感染的免疫能力也会逐渐减弱。T细胞是人体免疫系统的一类重要细胞,主要分为两类:辅助性T细胞负责协调免疫反应,杀伤性T细胞则负责清除受感染或癌变细胞。研究人员发现,超级百岁老人会积累一种不同寻常的“混合型”T细胞,即CD4细胞毒性T淋巴细胞(CD4 CTL)。这类罕见细胞同时具备识别威胁和摧毁危险细胞的能力。研究人员分析了不同年龄组人群的免疫细胞,包括70—90岁人群、百岁老人以及超级百岁老人。结果发现,在生命的大部分阶段,这类细胞始终十分少见,但在接近100岁时开始显著增加。在超级百岁老人中,CD4 CTL占血液中全部T细胞的比例接近1/5,而在较年轻的研究参与者中,这一比例仅约4%。进一步分析发现,部分CD4 CTL发生了明显的克隆扩增,即少数细胞不断复制,形成了数量庞大的同源细胞群。这一现象提示,这些细胞可能长期受到某些特定抗原的反复刺激,并在持续的免疫应答过程中不断增殖。研究人员还发现,超级百岁老人这类细胞所携带的部分T细胞受体,与癌症患者肿瘤组织中的T细胞受体高度相似,但这些超级百岁老人均无癌症病史。研究人员表示,这一发现提示,这些免疫细胞可能具有识别肿瘤细胞的能力,甚至可能在肿瘤尚未发展到临床可检测阶段时,就已对其产生免疫反应。这些发现意味着,极端高龄时期的免疫变化可能并非免疫系统单纯“衰老”和“耗竭”,而更像是免疫系统为适应长期健康生存而进行的一种重新组织。

  5. 达斯·维达赞美 Flock 车牌跟踪系统

    Flock 的车牌跟踪系统最近在美国引发了激烈争论,媒体同一时间报道了大量警官利用 Flock 摄像头跟踪女友/前女友、妻子/前妻的新闻。但在一片争论之中,皇帝最忠实的助手、西斯尊主达斯·维达则大肆赞美了 Flock。周三晚上加州圣地亚哥公共安全与宜居社区委员会会议(Public Safety and Livable Neighborhoods Committee Meeting)的公众评论期间,达斯·维达在台上说,“皇帝是 Flock 的粉丝,我们必须继续利用 Flock 技术,如此才能跟踪和监视那些叛军渣滓,看着他们从一个游乐场到另一个游乐场,从游乐场到游泳池,从游泳池到体育馆。因为我们都知道,Flock 摄像头不仅跟踪车牌;它们还跟踪孩子。它们在公园和体育馆里跟踪孩子,我们需要这个,我需要它来跟踪前女友。”

  6. Bilibili 进军国际市场

    Bilibili 本周重新发布了国际版应用,准备推出英文版本,进军全球市场。新的国际版应用将不需要身份验证,用户无需提供护照或身份证件即可注册。Bilibili 此前已积极邀请西方知名主播如 MrBeast 在其有 3.76 亿月活用户的中文主站发布视频。更大规模的全球扩张可能会挑战 YouTube 的霸主地位,但也面临类似 TikTok 的审查、内容审核和数据安全等棘手问题。根据招聘信息,B 站正在洛杉矶、伦敦、墨西哥城、圣保罗、伊斯坦布尔和东京招聘社区经理。

  7. 天文学家发现银河系已知最快的恒星

    天文学家发现了银河系已知运行速度最快的恒星。这颗名为 S301 的恒星围绕银河系中心的超大质量黑洞——人马座A*运行,最快速度达到每秒 2.5 万公里,超过光速的 8%。它的运行轨道非常接近人马座A*,其运动有望帮助科学家首次直接测量大质量黑洞的自转,并为检验爱因斯坦广义相对论提供新的机会。S301 绕人马座A* 公转周期为 8.7年。在它距离人马座 A*最近时——类似于太阳到土星的距离——恒星的运行速度超过光速的 8%。研究人员认为,S301的轨道特征以及恒星无法在如此靠近超大质量黑洞的位置形成,表明它很可能原本属于一个双星系统。当这个双星系统靠近人马座A*时,黑洞强大的潮汐力将两颗恒星撕裂,其中一颗被黑洞引力捕获,成为如今的 S301;另一颗则被高速抛出,其速度可能高到足以逃离银河系。

  8. Thunderbird 跟随 Firefox 采用双周发布模式

    Mozilla 工程总监 Sylvestre Ledru 上月宣布,从 2026 年 9 月起 Firefox 桌面版和 Android 版本的发布周期从 4 周减少到 2 周。本周释出的 Firefox 154 是最后一个按四周发布模式释出的版本,九月初释出的 v155 则是第一个双周发布版本。由 Mozilla 子公司 MZLA 开发的开源邮件客户端 Thunderbird 宣布也将采用双周发布模式。MZLA 的 Corey Bryant 称,从 9 月起 Thunderbird 采用相同的更新频率。

  9. 太阳能扩张政策与鸟类多样性下降相关

    南京信息工程大学的研究人员在《科学》上发表研究报告,称全球对太阳能发展的推动可能带来隐性的损害生物多样性的代价。可再生能源的扩张有助于应对气候变化,但大规模太阳能开发也可能因栖息地改变或破碎化而导致生物多样性丧失,从而引发新的环境得失权衡。研究人员汇编了一个大型数据集,它涵盖了 2014 年至 2023 年中国的 2344 个县。他们的数据集整合了鸟类观测数据、太阳能政策、环境条件和社会经济信息。他们还考察了土地利用、植被状况和农业生产力变化所带来的影响。研究结果表明,太阳能扩张政策的力度加大与鸟类多样性的显著下降有关:政策强度每增加一个标准差,鸟类生物多样性指数便会下降 2.10%。这些影响在较富裕地区和非沙漠地区最为显著,且对地理分布广泛的物种影响尤为严重。这主要应归因于土地的迅速转化,特别是将农田和草地转化为开发区,后者降低了植被的多样性。

  10. 海冰消失巨型鲸鱼进入格陵兰

    由于海冰融化,巨型鲸鱼如座头鲸进入到了以前难以抵达的东格陵兰沿海地区。这是东格陵兰海洋生态系统发生重大转变的一部分。直到 2006 年该地区才首次记录到座头鲸的踪迹。2007 年记录到了 7 头座头鲸,2024 年船载设备就记录到了 150 头。研究人员结合卫星标记鲸鱼的追踪轨迹和因纽特猎人的证词,估计 2024 年夏天大约有 4000 头座头鲸、6000 头长须鲸和 6000 头小须鲸造访了格陵兰海。这三种鲸鱼在夏季觅食季节至少会消耗 80 万吨鱼类和磷虾。北极原有的鲸鱼要么被迫适应要么被迫离开。

  11. AliExpress 被发现静默运行 WebAudio 指纹

    有开发者注意到一个奇怪的现象:蓝牙耳机支持多点蓝牙音频,能同时连接 PC 和手机,PC 通常优先播放音频,只有在 PC 没有播放内容时手机才会播放音频。这位开发者注意到,在 Firefox 或 Chrome 浏览器中打开 AliExpress 网页后,手机会停止播放音频,关闭网页则会恢复。这位开发者随后展开了调查,发现高度混淆的阿里巴巴安全脚本会创建两个 WebAudio 图形,成为浏览器指纹的一部分,该静默运行的 WebAudio 指纹会干扰多点蓝牙音频。用户可利用 uBlock Origin 扩展屏蔽阿里巴巴的脚本 collina.js 和 fireyejs.js 关闭这一指纹。

  12. Google 通过 Google Drive 提供 Android 特定源代码

    Android 安全加固项目 GrapheneOS 抨击 Google 违反了 GPLv2 许可证,原因是 Android 的部分源代码需要通过表单(Google Forms)递交申请然后通过云盘 Google Drive 获取,而且 Google 处理申请的速度越来越慢。GrapheneOS 指出,AOSP(Android 开源项目)现在只提供年度版本和季度更新版本 QPR2,以及针对这两个版本的安全回溯移植。Google 也停止向 AOSP 项目推送 Pixel 智能手机相关的特定代码,而 GrapheneOS 目前只支持 Pixel 智能手机,Google 此举严重影响了 GrapheneOS 对 Pixel 支持,这一状况促使 GrapheneOS 项目转而与摩托罗拉合作,预计支持 GrapheneOS 的摩托罗拉设备将在 2027 年推出。GrapheneOS 称,以前 Google 通常会在数小时内响应特定内核源代码的请求,如今需要数周甚至更长时间。Android 的内核使用的是 GPLv2 许可证,根据该许可证,如果用户索要修改后的源代码,Google 需要提供。但 GPLv2 没有规定多长时间提供。Google 作为全球科技巨头之一,它至少应在合理时间内提供源代码,不应该故意拖延。

  13. AI 记录员捏造了患者服用迷幻蘑菇的经历

    当 Rebecca Green 去看泌尿科医生时,医生询问是否可以用 AI 记录就诊经历,她同意了。但这一决定给她带来了巨大压力。因为 AI 抄写员捏造了她服用迷幻蘑菇的经历,她说自己从未碰过迷幻蘑菇。Green 女士直到三月肾结石手术后才发现这个错误,她阅读了专科医生发给她全科医生的术后信,信中称她曾服用过微剂量迷幻蘑菇,可能是之前肾脏周围出血的原因。在 Green 投诉之后,她的泌尿科医生回了封致歉信,猜测迷幻蘑菇的记录是 AI 听写错误的结果。Royal Australian College of GPs (RACGP)去年估计,四成的全科医生使用 AI 医疗记录员。这个比例数字还是一个保守估计。AI 记录员可以减轻医生的负担,但也会犯下导致临床决策改变的错误信息。

  14. GitHub 公布本周八小时宕机事故原因

    最大的代码托管平台 GitHub 本周发生了一次持续了近八小时的宕机事故,再次在开发者中间引发了寻找替代平台的讨论。GitHub 今年频繁发生宕机事故,已促使多个知名开源项目宣布迁移出去。本周的宕机事故始于 8 月 17 日 13:28 UTC,直至 21:15 UTC 才完全解决——持续 7 小时 47 分钟的事故导致 Issues、Pull Requests、API、Actions 和 Copilot 等服务大量出错。GitHub 解释说,事故直接原因是位于公司美国中部数据中心的负载均衡器网络饱和,而自动扩容策略的配置错误,以及 Visual Studio Code 中一个导致流量放大 10 倍的重试 bug 等一系列连锁反应导致了此次事故持续了如此长时间。

  15. 科学家首次实验观测到真空涨落对超导的增强效应

    在量子电动力学世界中,真空并非空无一物,而是伴随虚粒子的不断产生与湮灭。然而自由空间中的真空涨落通常较弱,难以对宏观凝聚态体系产生可观测影响。如何将其转化为调控量子物态的资源,是凝聚态物理与腔量子电动力学交叉领域的重要课题。研究团队探索利用真空涨落,实现对宏观量子物态的可控调节。研究团队为此引入太赫兹分裂环谐振器构成的“暗腔”,通过重塑电磁环境显著增强真空涨落。研究团队将超导体二硒化铌嵌入暗腔,发现暗腔中二硒化铌的超导临界温度获得实质性提高。在六层二硒化铌器件中临界温度最高提升 5.4%。同时,超导体的临界电流和临界磁场在超导转变附近也显著增强。这是国际上首次实验观测到真空涨落增强超导。

NEWSLETTER · FREE · WEEKLY

OrangeBot Weekly

The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.

Free · One-click unsubscribe · No spam