OrangeBot.AI Digest — 2026-08-27
90 headlines across 8 sources, aggregated for this day.
Hacker News(15)
- Gemini Omni 1.1 Flash (blog.google)
- We found a division by zero bug in FFmpeg with a vibecoded fuzzer (code.ffmpeg.org)
- Two German airport workers die of malaria after 'mosquito arrives on plane' (www.bbc.com)
- Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache (blog.cloudflare.com)
- The turbulent AI era is here (www.gatesnotes.com)
- Small Models Have Arrived (calv.info)
- Decompiling a Nintendo 64 game in 84 days (blog.chrislewis.au)
- Suica, Japan's First IC Transit Card (www.tokyodev.com)
- Show HN: The load-bearing vocabulary of Claude (louisabraham.github.io)
- 507 Mechanical Movements (507movements.com)
- US Government designates host of noblogs.org a "global terrorist" (crimethinc.com)
- Trade (and Tariffs) (xkcd.com)
- Microduck (pollen-robotics.com)
- Xcancel and Nitter have been taken down
- Tell HN: PayPal blocks GrapheneOS
GitHub Trending(15)
- bilawalsidhu / gods-eye-view
- zedeus / nitter
- freestylefly / awesome-gpt-image-2
- tt-a1i / archify
- JetBrains / go-modern-guidelines
- anthropics / claude-plugins-official
- K-Dense-AI / scientific-agent-skills
- DietrichGebert / ponytail
- calesthio / OpenMontage
- rohitg00 / ai-engineering-from-scratch
- ConardLi / garden-skills
- thedotmack / claude-mem
- google / googletest
- AgriciDaniel / claude-obsidian
- marin-community / marin
Product Hunt(15)
- Savvy
An assistant that whispers what to say during a meeting
- Kira Community
Curate moments with #hashtag
- Sendra
Design emails in Figma, export HTML that works everywhere
- Ojin
Talk to an AI Agent with a real face and voice, in real time
- SpacebarX
A keyboard-first outliner for notes, tasks and projects
- Enter Pro
The AI-native platform to build and scale your apps
- Skydive
Build cloud agents that work across your tools
- Yomi
A little cat who loves being read to
- Wondering Canvas
Visual ChatGPT in Parallel
- The Million Sad Ducks
$1 makes one of them permanently happy
- GitNexus (Akon Labs)
The open-source kernel for coding agents
- Lenz
Independent, multi-model fact-checking API for AI workflows
- Kraa 2.0
Text editor and a publishing platform
- Speko
OpenRouter for Voice
- GLM-5.3-Flash
The first natively multimodal model in GLM-5 series
Hugging Face(15)
- VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
- VGI-Bench: Probing Visual Intelligence in Video Generation Models
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/
- FrontierChallenge: Evaluating Scientific Workflow Completion
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
- WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.
- VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.
- D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D^3-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D^3-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D^3-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3times reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.
- Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G^2, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G^2 on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G^2 outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.
- Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method trains a model to generate implicit reasoning traces and rewards them by their ability to predict the next chunk of text. While promising, existing evaluations primarily compare against conventional SFT baselines, leaving open whether the gains come from the RL formulation itself or from more effectively exposing the model to no-CoT data. We address this question with a controlled study of next-chunk reasoning RL and a simple but previously overlooked alternative: Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data. Despite its simplicity, Mixed SFT achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute. The advantage is consistent across in-domain mathematical reasoning and out-of-domain reasoning tasks. Moreover, we show that higher pre-RLVR accuracy does not necessarily translate into higher post-RLVR accuracy, highlighting the need to evaluate no-CoT training strategies in the context of the full post-training pipeline.
- Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
- Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we establish a controlled M-OPD benchmark on SmolLM3-3B-Base with oracle routing, isolating capability integration from routing ambiguity. Our investigation reveals a pronounced capability integration gap: standard M-OPD captures only 35.6% of the available headroom relative to a domain-routed oracle ensemble, with concise tasks such as instruction following suffering severe degradation and premature stagnation. Crucially, we show that this failure stems not from gradient conflict, but from a severe misallocation of the token-level optimization budget. This pathology is driven by three orthogonal factors: structural sequence-length disparities across domains, dynamic convergence drift due to non-uniform learning rates, and multi-step reward staleness from asynchronous policy updates. To resolve these imbalances, we introduce Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student. We fully open-source our end-to-end post-training recipe, training trajectories, and evaluation suites on an academically accessible hardware budget.
- StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.
- Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.
- Code World Model: Coding Agent as World Brain
World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying knowledge, rules, and mechanisms governing world evolution. This makes it difficult to maintain persistent consequences and support coherent, open-ended evolution. We introduce Code World Model, a framework that separates world evolution from visual realization by combining the reasoning and coding capabilities of language models with the generative priors of video models. A coding agent serves as the world brain, reasoning about events and their consequences and generating executable code to maintain persistent world state and perform rule-consistent evolution. To connect executable state with visual generation, we introduce a proxy representation that encodes frame-wise spatiotemporal constraints and is compiled into a proxy video, which conditions a video model to render high-fidelity visual observations. We further develop data pipelines for constructing aligned proxy-observation pairs from gameplay and real-world videos. After fine-tuning on paired gameplay data, MiniMax-H3 follows proxy-based spatiotemporal specifications from simple interactive worlds built by the coding agent while preserving rich visual details and dynamics. These results demonstrate the potential of combining code for persistent world evolution with video models for flexible visual realization, providing a new path toward open-ended world models.
- SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
Techmeme(15)
- Sources: Town, which develops enterprise personal AI assistants, is in talks to raise at a $1B valuation in a round led by Index Ventures (Newcomer)
Newcomer : Sources: Town, which develops enterprise personal AI assistants, is in talks to raise at a $1B valuation in a round led by Index Ventures — Town, the buzzy enterprise personal agent, is raising at $1 billion, sources tell us — Two startup personal assistants have venture capitalists in a tizzy right now:
- Alphabet agrees to pay £260M to settle a UK class action lawsuit claiming it levied "unfair" charges on software downloaded from the Google Play app store (Alistair Gray/Financial Times)
Alistair Gray / Financial Times : Alphabet agrees to pay £260M to settle a UK class action lawsuit claiming it levied “unfair” charges on software downloaded from the Google Play app store — Group's Google business was accused of overcharging developers who made software for the Google Play app store
- Sources: Abu Dhabi's Sheikh Tahnoon and co-investors back an entity owning 49%, the largest stake, in the company behind the Trump family's planned crypto bank (Wall Street Journal)
Wall Street Journal : Sources: Abu Dhabi's Sheikh Tahnoon and co-investors back an entity owning 49%, the largest stake, in the company behind the Trump family's planned crypto bank — The Emirati ‘spy sheikh’ backs a 49% stake in the entity behind the new World Liberty bank — The Abu Dhabi royal sometimes referred …
- Nvidia stock closed up 8.74% on Thursday, adding ~$440B to Nvidia's market cap, after its revenue guidance reassured investors that AI demand remains strong (CNBC)
CNBC : Nvidia stock closed up 8.74% on Thursday, adding ~$440B to Nvidia's market cap, after its revenue guidance reassured investors that AI demand remains strong — Nvidia shares rose nearly 9% Thursday after the chip giant's revenue guidance reassured investors that artificial intelligence demand will remain strong.
- Google adds Expert Intelligence to Gemini Notebook, letting users import eligible titles they own in Google Play Books to ask questions, generate podcasts, more (Emma Roth/The Verge)
Emma Roth / The Verge : Google adds Expert Intelligence to Gemini Notebook, letting users import eligible titles they own in Google Play Books to ask questions, generate podcasts, more — Gemini Notebook's new ‘Expert Intelligence’ feature pulls information from the books you've purchased.
- Salesforce stock closed up 22.6% on Thursday, its second-best day ever, after the company reported a beat on Q2 earnings and expanded its Anthropic partnership (CJ Haddad/CNBC)
CJ Haddad / CNBC : Salesforce stock closed up 22.6% on Thursday, its second-best day ever, after the company reported a beat on Q2 earnings and expanded its Anthropic partnership — Shares of Salesforce jumped 22% Thursday after the company reported a beat on second-quarter earnings and announced …
- YouTube adds Amazon to its shopping affiliate program, enabling US creators to tag Amazon products in Shorts, longform videos, and livestreams (Anna Washenko/Engadget)
Anna Washenko / Engadget : YouTube adds Amazon to its shopping affiliate program, enabling US creators to tag Amazon products in Shorts, longform videos, and livestreams — The retailer just finally joined YouTube Shopping's affiliate program. — YouTube has landed one of the biggest possible gets for its creator ecommerce platform.
- Sources: Anthropic has been working on a plan to allow secondary stock sales in its IPO while also considering lockup periods longer than the standard 180 days (The Information)
The Information : Sources: Anthropic has been working on a plan to allow secondary stock sales in its IPO while also considering lockup periods longer than the standard 180 days — Anthropic has been working on a plan to let existing shareholders sell some stock in its blockbuster initial public offering …
- Anthropic releases Model Hardware Standard, a framework to help AI agents use physical systems like microscopes, quantum computing hardware, and robot arms (Will Knight/Wired)
Will Knight / Wired : Anthropic releases Model Hardware Standard, a framework to help AI agents use physical systems like microscopes, quantum computing hardware, and robot arms — The potential for AI to automate scientific research and manufacturing must be balanced with new risks, Anthropic says.
- Internal memo: Meta's AI agent Hatch "has its own computer" to perform tasks, works when the app is closed, can connect to email, Instagram, OpenTable, and more (Hugh Langley/Business Insider)
Hugh Langley / Business Insider : Internal memo: Meta's AI agent Hatch “has its own computer” to perform tasks, works when the app is closed, can connect to email, Instagram, OpenTable, and more — It can DJ. It can order food. It can book you a table at a restaurant. And of course, it can access your Instagram.
- Source: Nvidia plans an employee-funded political action committee called NVPAC, as it seeks to step up its efforts to influence US policy (Courtney Rozen/Reuters)
Courtney Rozen / Reuters : Source: Nvidia plans an employee-funded political action committee called NVPAC, as it seeks to step up its efforts to influence US policy — Tech company Nvidia will start an employee-funded political action committee, a source familiar with the plans told Reuters, as the company seeks to step up its efforts to influence U.S. policy.
- OpenAI, Anthropic, AWS, Microsoft, and 100+ companies warn there is "a limited window" to prepare for AI-enabled cyberattacks and call for "collective action" (Sam Sabin/Axios)
Sam Sabin / Axios : OpenAI, Anthropic, AWS, Microsoft, and 100+ companies warn there is “a limited window” to prepare for AI-enabled cyberattacks and call for “collective action” — OpenAI, Anthropic, Amazon Web Services, Microsoft and more than 100 other companies warned Thursday …
- OpenAI is testing a "Persistent mode" in Codex, designed to let AI agents "continue working until put to sleep" and proactively generate follow-up tasks (Maxwell Zeff/Wired)
Maxwell Zeff / Wired : OpenAI is testing a “Persistent mode” in Codex, designed to let AI agents “continue working until put to sleep” and proactively generate follow-up tasks — Code reviewed by WIRED reveals the company is developing a feature that enables Codex to continue working proactively until it is “put to sleep.”
- Sources: some Trump administration officials circulated a draft EO to create a self-regulatory organization for AI, but it needs buy-in from Trump and others (Leo Schwartz/The Information)
Leo Schwartz / The Information : Sources: some Trump administration officials circulated a draft EO to create a self-regulatory organization for AI, but it needs buy-in from Trump and others — The Trump administration has internally circulated a draft executive order in recent weeks, calling for the creation …
- Google launches Gemini Omni 1.1 Flash, which it says delivers studio-quality video production, including the ability to extend a scene, 4K upscaling, and more (Google)
Google : Google launches Gemini Omni 1.1 Flash, which it says delivers studio-quality video production, including the ability to extend a scene, 4K upscaling, and more — Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation …
Solidot(15)
- 一名微软工程师一个月的 AI 支出高达 2.8 万美元
在长时间鼓励之后,本月初微软开始要求员工限制 AI 使用。执行副总裁 Jay Parikh 在一封发给微软员工的邮件中要求工程师专注于业务成果,而非最大化 AI token 的使用量,为了“从 token 投资中获得更大的价值”,微软将比 Anthropic 模型更便宜的 OpenAI GPT-5.6 设为内部使用的默认模型。根据一份微软员工自愿提交的 AI 使用费账单:Customer and Partner Solutions 部门的一名员工在 28 天内的 AI 支出高达 2.8 万美元;多名员工支出超过 1 万美元;中位数约为每 28 天 300 美元,少数部门的 AI 支出仅仅为几十美元;CoreAI 部门的 AI 支出中位数最高为 975 美元。
- Meta将支付 170 亿美元和解儿童隐私保护诉讼,将限制青少年在特定时间访问社媒
Meta 与美国多州就儿童隐私和消费者保护案达成和解,同意向各州支付总额将近 167 亿美元和解金,以及对青少年用户施加一系列社媒使用限制,包括每天不能使用超过两小时。最具深远影响的和解条款是 Meta 同意实施一系列新的保障措施。这些措施将对面向青少年的社媒运作方式起到实质改变。其中一项变更将启用“夜间屏蔽”功能:在默认设置情况下,从午夜至凌晨6时,青少年将无法访问 Facebook 和 Instagram。此外青少年账户在 Meta 旗下所有社媒的每日累计使用时长,默认上限为两小时。
- NASA 准备本周日发射罗曼太空望远镜
NASA 准备本周日 8 月 30 日在佛罗里达州的肯尼迪太空中心使用 SpaceX 的重型火箭 Falcon Heavy 发射罗曼太空望远镜。罗曼太空望远镜以 NASA 首任天文学部门女主任 Nancy Grace Roman 的名字命名,使用了美国国家侦察局捐赠的 2.4 米口径主镜,配备了两台科学仪器:3 亿像素多波段红外相机大视场仪表(WFI),能直接观测邻近恒星周围的类木行星的日冕仪(CGI)。其核心任务包括探测暗能量、发现系外行星及验证广义相对论宇宙时空曲率。罗曼望远镜不仅拥有哈勃望远镜的清晰视力,其视场扩大了 100 倍,能快速进行巡天观测,以解答天文学中的一些重大问题。望远镜在前五年主要开展 3 项大型巡天观测:第一项聚焦超新星,第二项聚焦宇宙学,第三项则聚焦系外行星。前两个项目旨在揭示暗能量的奥秘,暗能量是推动宇宙加速膨胀的神秘力量。
- 亚马逊 AI 训练设施员工谈内部工作
404 Media 前不久跟踪一本珍本图书到亚马逊位于内华达州拉斯维加斯的一个仓库,在该仓库工作的亚马逊团队被称为 VGT3,其 logo 是一只张着嘴、手里拿着一本书的恐龙。他们的主要工作是拆开书脊扫描图书训练 AI。一名在该仓库的亚马逊员工匿名接受了采访,谈论了他们的工作。这名员工称,仓库接收了大量图书,有新书,也有二手书,甚至还看到过代表英女王给议会的文件;图书的语种也是各种各样,有德语、俄语还有日语;他们会扫描图书的条形码移除重复的图书,重复的书会退还给图书经销商;他们使用一种人工操作的机器去切开书脊,使用几十台扫描仪扫描书页,扫描过的书页会处理掉;亚马逊最初告诉他们扫描的书页是用于 kindle 电子书库,这无疑是借口,因为肯定存在版权方面的问题,后来他们才知道是为训练 AI 建立语料库。
- 亚马逊 Mechanical Turk 将于 9 月 30 日关闭
亚马逊宣布其众包平台 Mechanical Turk 将于 9 月 30 日关闭。亚马逊上个月才宣布将于 7 月 30 日起停止接受新用户。当时亚马逊表示该决定是在“慎重考虑”后做出的,“现有用户可以继续正常使用该服务。AWS 将继续投资改进 Mechanical Turk 的安全性和可用性,但我们不打算推出新功能。如今它正式给 Mechanical Turk 画上了句号。亚马逊是从 2018 年起将 Mechanical Turk 变成训练神经网络的标注数据服务。但讽刺的是 2023 年的研究发现,该平台 33% 到 46% 的众包工作者使用大模型去完成任务,引发了对标注数据的可靠性以及是否真的需要人类参与的质疑。由于大量的机器人和欺骗行为,研究人员已经放弃了该平台,它的关闭只是时间问题。
- LibreOffice 26.8 释出
The Document Foundation 宣布释出 LibreOffice 26.8,强调该版本不包含生成式 AI 功能,不会将文档传输到远程服务进行处理,各个组件也不需要网络访问就能运行。LibreOffice 26.8 最大的单项改进聚焦于双向文本和复杂文本处理。Writer 现在会在打开或粘贴文档/纯文本时自动检测段落方向;换行时行尾空格的位置会根据段落方向而非相邻字符方向调整。从右至左或竖排中日韩(CJK)文档的对象大小调整句柄能正确工作。双向控制字符现已与其他格式标记一同可见。在 Calc 中,若在空单元格中输入从右至左的文本,系统会自动设置该单元格的方向。
- Apple Maps 加入广告
苹果的地图服务 Apple Maps 加入了广告。美国和加拿大用户的搜索结果顶部以及“推荐地点(suggested places)”部分会显示付费广告商家的名字。苹果表示,广告可能基于用户的大致位置、搜索词或正在查看的地图区域,但不会与用户的 Apple 帐户关联,个人数据也保留在设备上。付费商家上会显示“Ad”的蓝色徽章。苹果强调,其广告政策致力于保护用户隐私,苹果不会收集或存储个人数据,也不会与第三方共享。
- 中尼边境泥石流灾害逾 1300 人失踪
8 月 26 日周三上午中尼边境吉隆口岸发生泥石流灾害。尼泊尔周四报告有 165 人死亡,826 人失踪。中国央视报告,截至周四上午有 3 人死亡,558 人失踪,其中外国人 260 人。这次事件被认为是冰川崩塌引起的,冰川崩塌的震动是如此之大,以至于当地地震仪记录到了地震事件。美国地质调查局(USGS)称未发生地震,地震事件是 5.2 级冰川崩塌引起的。加拿大卡尔加里大学的 Dan Shugar 说,过去两天的卫星图像显示附近冰川损失了不少雪,“我猜测,积雪融化以及冰川冰的融化,将大量液态水注入冰川,这些水渗入冰川底部,起到润滑作用。这足以导致冰川崩塌。”他表示冰川崩塌可能与气候变化无关,但此类灾害发生频率增加表明,气候变暖是其促成因素之一。
- 亚马逊收购开源数据库 DuckDB 开发团队
亚马逊同意收购开源数据库 DuckDB 开发团队 DuckLabs。DuckLabs 员工将加入亚马逊 AWS,其中包括联合创始人 Hannes Muhleisen 和 Mark Raasveldt,他们将继续领导团队和制定项目的技术方向,员工也将继续在阿姆斯特丹办公。DuckDB 项目将继续维持现有的开源状态,使用 MIT 许可证,由独立基金会管理。收购 DuckLabs 被认为有助于将亚马逊的云存储服务 S3 转型为客户分析数据而非仅仅存储数据的平台。
- 新 Twitter.now 上线
总部位于美国弗吉尼亚州的初创公司 Operation Bluebird 上线了新社交网络 Twitter.now,致力于复兴已被马斯克(Elon Musk)的 X 抛弃了的 Twitter。马斯克在 2022 年收购 Twitter 之后迅速将其改名为 X,弃用了 Twitter 相关标识。Operation Bluebird 的联合创始人认为此举标志着 X 放弃了 Twitter 的身份和知识产权, 为其申请 Twitter 商标权创造了机会,他们因此准备推出新版的 Twitter。X 去年底起诉了 Operation Bluebird,要求特拉华州联邦法官发布初步禁令,阻止新版 Twitter 上线。法官 Colm Connolly 在今年四月做出了临时裁决,认为 X 看起来放弃了对“tweet”一词和 Twitter 蓝鸟 logo 的知识产权主张,甚至可能也包括对“Twitter”一词的知识产权主张。法官尚未发布书面命令。新版的 Twitter.now 与旧版的 Twitter.com 外观非常相似,一大不同之处是内置了事实核查功能,强调自己与 X 毫无关系。
- 英伟达同意以 129 亿美元收购 Hugging Face
The Information 报道,英伟达同意以 129 亿美元收购 Hugging Face。这笔交易仍处于敲定阶段,仍有可能告吹。英伟达是 Hugging Face 的投资者,它在 2023 年参与了 Hugging Face 的 D 轮融资,当时以 45 亿美元估值融资 2.35 亿美元。英伟达去年还提议以 70 亿美元估值投资 5 亿美元,但遭到 Hugging Face 的拒绝。Hugging Face 是开源 AI 生态系统的核心,托管着数百万个可供开发者使用的 AI 模型和数据集。收购该平台有助于让英伟达在 AI 开发者中占据更大的市场份额。
- 育碧在 Steam 上架《魔法门之英雄无敌3》时忘记将游戏文件放进去
在宣布重制版的同时,育碧于 8 月 26 日在 Steam 上架了原版的《魔法门之英雄无敌3》,国区售价为 40 元。然而购买游戏的 Steam 玩家发现下载的游戏文件仅仅为 23.49 KB,育碧忘记把游戏文件打包进去了。直到数小时之后育碧才修复问题,释出了包含游戏完整文件的版本,完整版本容量大约为 1GB。此事导致玩家在游戏页面留下了大量差评。
- 澳大利亚是地球最安全的地方
人类面临无穷无尽的灾难:核战、气候变化、流行病、AI、超级火山喷发……如果发生了全球性灾难,什么地方最安全,能给你保留再见光明的机会?根据发表在《Global Challenges》上的一项研究,澳大利亚是最安全的地方。澳大利亚并不能免于全球性灾难,但相对而言影响最小。研究人员评估了数百项关于小行星撞击地球、火山爆发、大规模网络攻击、疾病威胁等灾害影响研究和报告。研究人员称,全球灾难危机通常分为三种:火山爆发、核冬天或小行星撞击云导致的突发性日照减少;地磁风暴、网络攻击、流行病和高空电磁脉冲等导致的全球基础设施瘫痪;大规模生物灾害导致的大规模伤亡和严重社会混乱。一个国家如果具备以下特征——民主;富裕;不平等程度低;有庞大的工业基础和较高的政府影响力;位于偏远地区如岛屿;权力下放程度高,社会多样性强;以及对贸易的依赖程度不高,或拥有地理位置相近的贸易伙伴——那么它应对全球性灾难的韧性会比较强。澳大利亚是唯一一个被认为能抵御所有三种全球性灾难危机的国家。
- 美国报告今年的首例麻疹死亡病例
美国宾夕法尼亚州卫生官员周二证实,该州两名未接种疫苗者死于麻疹。这是宾州 35 年来首次报告麻疹死亡病例,也是美国 2026 年以来首次报告麻疹死亡病例。去年美国报告了三例麻疹死亡病例,其中两人是德州的未接种疫苗学龄儿童,一人是新墨西哥州的未接种疫苗成年人。之前美国自 2015 年以来未报告麻疹死亡病例。由于反疫苗宣传和虚假信息,美国的疫苗接种率持续下滑,麻疹这一传染性极强的病毒感染病例正在激增。美国 CDC 已统计到至少 2777 例麻疹病例,为 1991 年以来最高纪录。美国疫苗接种率已下滑至约 92%,低于群体免疫所需的 95% 接种率。宾州的麻疹疫苗接种率从 2017 年的 96.7% 降至 2026 年的 92.7%,死亡病例所在的 Lancaster 县去年幼儿园儿童的麻疹疫苗接种率只有 87.6%,卫生官员督促居民接种疫苗。
- 中尼吉隆口岸泥石流灾害数百人失踪
2026 年 8 月 26 日 10 时 30 分许,尼泊尔一侧发生泥石流灾害,随后灾害波及中尼边境吉隆口岸一带。中国官方通报称,灾害造成吉隆口岸重大人员伤亡及人员失联。灾害同时造成中国一侧前往吉隆口岸的道路、通信和电力中断。尼泊尔方面亦遭受严重山洪灾害,拉苏瓦县等地的村庄、道路、桥梁及水电设施受到破坏。尼泊尔方面至少 22 人死亡,约 384 人失联,其中包括 291 名外国游客。中国境内的死亡、受伤及失联人数尚未公布。
OrangeBot Weekly
The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.
Free · One-click unsubscribe · No spam