Curated by Shen Huang · 88 stories · ~13 min read
DIGEST · 2026-08-05

OrangeBot.AI Digest — 2026-08-05

88 headlines across 8 sources, aggregated for this day.

Hacker News(15)

  1. Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery (www.wired.com)
  2. Zed DeltaDB (zed.dev)
  3. Beating GPT-5.6 Sol on retrieval with 100x cheaper open models (neon.com)
  4. Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs (blog.google)
  5. Discovery Loop (www.discoveryloop.com)
  6. The "Disability Dongle": Why Silicon Valley Hates Me and You (sightlessscribbles.com)
  7. Jeff Dean leaving Alphabet (www.nytimes.com)
  8. Demis Hassabis is moving from CEO to Chairman at Google DeepMind (www.axios.com)
  9. Oracle cut its Always Free ARM limits to 2 OCPU / 12GB, enforced Aug 18 (www.cnelecar.com)
  10. Qwen Image 3.0 Pro (www.qwencloud.com)
  11. Cops Used Flock to Track a Man Across State Lines for a Pretextual Weed Search (www.404media.co)
  12. Cloudflare OS: an open platform for agents, apps, and work (blog.cloudflare.com)
  13. TIME Is Serving AI Bots a Different Website, with Ads Built In (www.vincentschmalbach.com)
  14. Position: LLMs Can't Jump (openreview.net)
  15. Civilian plane crash in New Mexico tied to military GPS blocking (www.wired.com)

GitHub Trending(13)

  1. cloudflare / computer
  2. huangruiteng / loopx
  3. TencentCloud / TencentDB-Agent-Memory
  4. donnemartin / system-design-primer
  5. firecrawl / pdf-inspector
  6. esengine / DeepSeek-Reasonix
  7. addyosmani / agent-skills
  8. obra / superpowers
  9. roboflow / supervision
  10. vercel / next.js
  11. tailwindlabs / tailwindcss
  12. uber / ADR
  13. lyogavin / airllm

Product Hunt(15)

  1. Keystroke

    Build powerful AI agents & workflows

  2. AdAnt AI

    Claude for viral, high-converting social ads

  3. Wispr Flow Notetaker

    Meeting notes that get the details right.

  4. BackEngine MCP

    Make private company knowledge usable for AI

  5. NextDoor.Company

    Discover startups hiring near you, on a map

  6. Cloudflare Wallets

    the programmable wallet for the agentic Internet

  7. Kiro Crew

    Open source agentic development workspace

  8. Aegisora

    The narrow control plane for AI agent tool and API calls.

  9. Keytones

    Distinct key sounds for uppercase, lowercase & more

  10. Hansel

    Remember everything you've worked on

  11. npm i -g hotcell

    Local sandboxes for AI agents on your Mac, Linux, bare metal

  12. StepGrab

    Turn any Mac task into a step-by-step guide

  13. Dover MCP

    Run your hiring process from Claude or ChatGPT

  14. Capacity Desktop

    A free Lovable that lives on your Mac

  15. X Money

    Your money, on the world's most powerful network.

Hugging Face(15)

  1. MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

    Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.

  2. JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

    Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration. Our method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation (SA-DMD), and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift. Extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU. Code is available at https://github.com/jd-opensource/JoyAI-Video-Edit.

  3. AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

    Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces not designed for joint generation and decoding, or compress autoencoded latents to ease diffusion, sacrificing token-level fidelity. Instead of simplifying the representation to suit the generative model, we preserve a high-capacity, decodable text latent and design the diffusion model to learn its distribution directly. We introduce AURORA-LM, a continuous-latent diffusion language model that separates the construction of a decodable text representation from the modeling of its distribution. A Query-based Encoder-Decoder organizes text into a high-capacity, prefix-aligned latent sequence, and a Block-causal Diffusion Transformer learns its distribution through flow matching, generating blocks left to right while denoising positions within each block in parallel. Because such a latent is harder for diffusion to model, AURORA-LM restricts only the noisy-input pathway while retaining the full clean-latent prediction target, accommodating full-width latents without reducing decoder-facing capacity. We further calibrate the noise-level distribution to the latent width, and introduce self-trajectory consistency to bridge independently sampled training noise and iterative denoising at inference. AURORA-LM achieves the strongest performance among evaluated continuous and diffusion-based language models on OpenWebText free generation and XSum summarization. Scaling to 1B parameters with about 1500 EFLOPs of total compute yields further gains, surpassing a larger publicly released latent-diffusion language model under a matched evaluation protocol. All experiments are conducted on Ascend NPUs.

  4. Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

    Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/

  5. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

    We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

  6. Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

    Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.

  7. PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

    Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated token-level discrepancies, which can be sensitive to noise, or assign a shared step-level weight that may overlook positional variation. We propose Persistent Consistency Self-Distillation (PCSD), which derives token-level distillation weights from the local persistence of teacher-favoring signals. PCSD combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support, applies trend-aware modulation to attenuate locally declining support, and produces continuous weights through sigmoid gating. The resulting objective is jointly optimized with GRPO, combining dense teacher guidance with sparse environmental feedback. Without inference-time skills, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones, exceeding GRPO by 15.6 and 13.3 points and SDAR by 6.2 and 5.5 points, while remaining competitive on WebShop and gaining 15.8 points over GRPO on unseen ALFWorld split.

  8. Quo Vadis, World Modeling?

    Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.

  9. PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

    Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench

  10. LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

    Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.

  11. OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

    Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.

  12. Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

    On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. We ask what happens when none of this holds, as when the strongest teacher available and the student one wishes to deploy come from different model families, and find that the standard recipes have no answer: teacher latents cannot serve as targets in a foreign coordinate system, per-pixel losses against a teacher that stochastically re-draws local detail degenerate into blur or divergence, and timestep indices lose their meaning across mismatched schedules. We present Any-OPD, to our knowledge the first framework for on-policy distillation between arbitrary pairs of latent flow-matching generators. Any-OPD treats the teacher purely as a black-box sampler and connects the two models at exactly one point: a frozen, model-agnostic vision representation in which their independently decoded outputs are compared, sidestepping every assumption about latents, features, or architecture. Trajectory correspondence is recovered by matching continuous noise levels instead of step indices, and a brief anchoring phase, in which teacher samples are re-encoded through the student's own VAE, ensures the on-policy gradient measures sample quality rather than domain mismatch. Distilling the 12B FLUX.1-dev into the 2.5B SD3.5-Medium, Any-OPD lifts the student's PickScore from 0.846 to 0.884 and HPSv3 from 9.12 to 10.97, rivaling the teacher at a fifth of its size, where direct latent regression fails to train at all.

  13. CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

    Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.

  14. SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

    Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experiences can be transformed by the agent itself into durable behavioral artifacts. We present SkillJack, the first attack that exploits the experience-to-skill pipeline of self-evolving agents. Instead of directly manipulating runtime context, SkillJack hijacks the agent's own learning process to implant malicious behaviors into its reusable skill repertoire. We identify three key properties of this transformation: sanitization whitewashing, where malicious intent is obscured during skill extraction; cross-layer promotion, where transient experiences become persistent capabilities; and persistence isolation, where the attack survives removal of its original source records. We evaluate SkillJack on two representative systems, SkillX and Anything2Skill, using a shared dataset of 150 trajectories across four policy-risk categories. Results show that skill extraction substantially reduces attack detectability: in SkillX, safety detection drops from 98.5\% for poisoned trajectories to 11.4\% for extracted skills, while Anything2Skill shows a similar effect. Meanwhile, the implanted skills remain effective, achieving attack success rates of 56.2\% and 89.2\% on the two systems, respectively. Furthermore, 80.0\% of skill-mediated attacks persist after deleting the original poisoned records, and some skills unintentionally activate on benign queries. Our findings reveal skill evolution as a new attack surface and motivate provenance-aware skill lifecycle protection. Our code is available at https://github.com/Tencent/AI-Infra-Guard/research/skilljack.

  15. UniWorld-Design: From Pixel Generation to Layer-Native Design

    We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an image is created, understood, and edited. Just as human designers create and manipulate visual content through layers rather than raw pixels, UniWorld-Design equips multimodal generative models with a layer-native design space. UniWorld-Design comprises two models. The Text-to-RGBA (T2RGBA) model generates standalone RGBA assets directly from text. The Image-to-Layer (I2L) model conditions on a finished image, a global instruction and per-layer prompts, and jointly produces ordered, complete semantic RGBA layers. Its instruction interface supports top-level decomposition, recursive decomposition and targeted extraction, making layering an instruction-addressable operation for agentic editing. Because I2L learns complete semantic objects rather than visible-pixel partitions, its layers stay usable when moved or removed. On the Crello benchmark, I2L reduces per-layer RGB L1 error by 37% and achieves a 34% relative improvement in Alpha Soft IoU over Qwen-Image-Layered. Separately, T2RGBA achieves the highest CLIP Score, outperforming LayerDiffuse and OmniAlpha.

Techmeme(15)

  1. Duolingo reports Q2 revenue up 18% YoY to $298.5M, paid subscribers up 17% to 12.7M, below est., forecasts Q3 revenue below est.; DUOL drops 11%+ after hours (Akash Sriram/Reuters)

    Akash Sriram / Reuters : Duolingo reports Q2 revenue up 18% YoY to $298.5M, paid subscribers up 17% to 12.7M, below est., forecasts Q3 revenue below est.; DUOL drops 11%+ after hours —  Duolingo (DUOL.O) forecast third-quarter revenue below Wall Street expectations on Wednesday, tempering optimism from its robust second quarter …

  2. DoorDash reports Q2 marketplace gross order value up 36% YoY to $33.08B, vs. $32.08B est., and forecasts Q3 marketplace GOV and adjusted EBITDA above estimates (Neil J Kanatt/Reuters)

    Neil J Kanatt / Reuters : DoorDash reports Q2 marketplace gross order value up 36% YoY to $33.08B, vs. $32.08B est., and forecasts Q3 marketplace GOV and adjusted EBITDA above estimates —  DoorDash (DASH.O) on Wednesday forecast third-quarter gross order value and core profit above Wall Street estimates after topping results …

  3. Binance affiliates are suing RedotPay's founders for allegedly diverting 470K+ users to a competing product in a "fraudulent scheme", claiming $472.8M in losses (Bloomberg)

    Bloomberg : Binance affiliates are suing RedotPay's founders for allegedly diverting 470K+ users to a competing product in a “fraudulent scheme”, claiming $472.8M in losses —  Binance affiliates are suing the founders of Hong Kong-based crypto payments firm RedotPay for allegedly diverting hundreds …

  4. Salesforce appoints Miguel Milano, its chief revenue officer, as COO; Chief Operating and Financial Officer Robin Washington will keep her title (Jordan Novet/CNBC)

    Jordan Novet / CNBC : Salesforce appoints Miguel Milano, its chief revenue officer, as COO; Chief Operating and Financial Officer Robin Washington will keep her title —  Salesforce has appointed an operating chief under CEO Marc Benioff, promoting Miguel Milano to the role.  —  Milano, a former Oracle executive …

  5. Nikita Bier says he will step back from leading product for X and continue as an adviser (Nikita Bier/@nikitabier)

    Nikita Bier / @nikitabier : Nikita Bier says he will step back from leading product for X and continue as an adviser —  Ladies and gentlemen, it's time to pass the torch and demote myself to my natural state: a poster. I'll be stepping back from leading product for 𝕏 and will continue on as an advisor. Serving the X community has been the privilege of a lifetime. X is, and will remain, the most [video]

  6. eBay reports Q2 revenue up 15% YoY to $3.13B, vs. $3.02B est., GMV up 15% to $22.4B, and forecasts Q3 revenue above estimates (Reuters)

    Reuters : eBay reports Q2 revenue up 15% YoY to $3.13B, vs. $3.02B est., GMV up 15% to $22.4B, and forecasts Q3 revenue above estimates —  EBay (EBAY.O) forecast third-quarter revenue above Wall Street estimates on Wednesday, leaning on its push into authenticated luxury goods, collectibles …

  7. Etsy says it will cut ~220 employees, or ~12% of its workforce, mostly in product and engineering, and reports Q2 revenue up 6% YoY to $668.3M, vs. $649.1M est. (Annie Palmer/CNBC)

    Annie Palmer / CNBC : Etsy says it will cut ~220 employees, or ~12% of its workforce, mostly in product and engineering, and reports Q2 revenue up 6% YoY to $668.3M, vs. $649.1M est. —  Etsy said Wednesday it's laying off about 220 employees, or roughly 12% of its workforce, as the online marketplace looks …

  8. Meta is offering a cheaper Muse Spark 1.2 "contributor" tier priced at $0.10/1M input and $0.20/1M output tokens in exchange for using user prompts for training (Wall Street Journal)

    Wall Street Journal : Meta is offering a cheaper Muse Spark 1.2 “contributor” tier priced at $0.10/1M input and $0.20/1M output tokens in exchange for using user prompts for training —  The company, pressed by investors to generate revenue from AI, says its offering will cost less than popular alternatives

  9. Filing: Microsoft recorded $24.1B in revenue from OpenAI during the year ended in June, suggesting OpenAI accounted for more than half of Microsoft's AI sales (Bloomberg)

    Bloomberg : Filing: Microsoft recorded $24.1B in revenue from OpenAI during the year ended in June, suggesting OpenAI accounted for more than half of Microsoft's AI sales —  Microsoft Corp. generates most of its artificial intelligence revenue from OpenAI, according to new disclosures from the company.

  10. Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at $1.25/1M input and $4.25/1M output tokens (Jonathan Vanian/CNBC)

    Jonathan Vanian / CNBC : Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at $1.25/1M input and $4.25/1M output tokens —  Meta is rolling out its first coding agent called Muse Code as the company tries to challenge leading AI labs Anthropic and OpenAI.

  11. Source: Yunfeng Capital, a PE firm co-founded by Jack Ma, invested ~$30M in AI insurance tech startup Corgi; Corgi CEO says the firm "did not invest in us" (Business Insider)

    Business Insider : Source: Yunfeng Capital, a PE firm co-founded by Jack Ma, invested ~$30M in AI insurance tech startup Corgi; Corgi CEO says the firm “did not invest in us” —  Jack is back.  —  Chinese billionaire Jack Ma's private equity firm secretly led the latest funding round for Corgi …

  12. Mysk: Apple's Private Relay tool can leak users' IP addresses due to issues in Apple's WebKit browser engine, also affecting OnionBrowser, a Tor browser for iOS (Joseph Cox/404 Media)

    Joseph Cox / 404 Media : Mysk: Apple's Private Relay tool can leak users' IP addresses due to issues in Apple's WebKit browser engine, also affecting OnionBrowser, a Tor browser for iOS —  Researchers found a group of issues that mean Private Relay isn't actually protecting users' real IP addresses.

  13. Reddit expands its test of Rules Hub, a suite of tools that rely on LLMs to help moderators manage their communities, and plans a full launch later this year (Jay Peters/The Verge)

    Jay Peters / The Verge : Reddit expands its test of Rules Hub, a suite of tools that rely on LLMs to help moderators manage their communities, and plans a full launch later this year —  It's also planning big changes for developers and old Reddit. … Reddit is enlisting AI to help moderate new subreddits — and eventually the rest of site.

  14. Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)

    Ivan Mehta / TechCrunch : Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release —  Hark, a startup that raised $700 million in Series A funding in May, today launched its agent Hark Handoff, which can use a browser efficiently to complete tasks.

  15. Google says Chief AI Architect and Google DeepMind CTO Koray Kavukcuoglu will lead Google DeepMind as SVP, reporting to Sundar Pichai (Bloomberg)

    Bloomberg : Google says Chief AI Architect and Google DeepMind CTO Koray Kavukcuoglu will lead Google DeepMind as SVP, reporting to Sundar Pichai —  Alphabet Inc.'s Google is losing some of its most prominent artificial intelligence veterans in a seismic overhaul that is casting doubt over leadership …

Solidot(15)

  1. Google DeepMind CEO Demis Hassabis 卸任

    Google 对其 AI 部门 DeepMind 进行了领导层重组,诺贝尔奖得主、DeepMind CEO Demis Hassabis 卸任,他将担任新设立的首席科学家职位并改任董事长。资深工程师以及 Google Brain 联合创始人 Jeff Dean、Sanjay Ghemawat、Oriol Vinyals 和 ​Quoc Le 都离开 Google,他们成立了一家公益性 AI 企业 Discovery Loop,专注于机器学习、科学和工程领域的突破性研究。这次人事变动恰逢 Google 最新 Gemini 模型的发布滞后,原计划 6 月发布,但至今仍未发布,引发了 Google 落后于竞争对手 Anthropic 和 OpenAI 的担忧。

  2. 微软要求工程师不要最大化 AI Token 使用

    为了推广 AI 工具,许多企业将 AI 使用率作为绩效考核的一部分。结果就是员工为了绩效致力于最大化 AI 使用,导致企业很快发现 token 费用大幅超出预算。软件巨头微软成为最新一家建议工程师限制使用 AI 的公司。微软执行副总裁 Jay Parikh 在一封发给微软员工的邮件中要求工程师专注于业务成果,而非最大化 AI token 的使用量,为了“从 token 投资中获得更大的价值”,微软将比其它模型更便宜的 OpenAI GPT-5.6 设为内部使用的默认模型。自 2026 年 7 月起,微软各部门将设定“AI token 预算目标”,员工可以追踪各自的 AI 支出。

  3. AI 监督远程考试变成灾难,数万学生必须重考

    从 5 月下旬到 6 月初,近 16 万名考生参加了墨西哥最大大学 UNAM 的入学考试。这是 UNAM 首次采用完全远程考试。结果是一场灾难。考试成绩公布后,高分比例比往年多多了。UNAM 120 道题考试 2021-2025 年间,只有 3.5% 的考生获得 100 分或以上,今年这一比例高达 16.3%;只有 0.9% 的考生获得 110 分或以上,今年这一比例达到了 5.5%。高分激增引发了考试作弊的指控。UNAM 的专家委员会在调查之后认为最佳方案是线下复试。受影响的 58,000 人将需要在监考人员监督下参加复试。

  4. Waymo CEO 解释为什么光靠摄像头难以实现自动驾驶

    Waymo 联席 CEO Dmitri Dolgov 在 Y Combinator 的 Startup School 发表演讲,解释为什么自动驾驶汽车需要的传感器不能仅限于摄像头。特斯拉汽车只配备了光学摄像头,它的辅助驾驶系统依赖于来自摄像头的数据。Dolgov 解释说,人类仅靠眼睛就能驾驶汽车,如果自动驾驶系统的目标是实现人类水平的驾驶,那么只靠摄像头可能够了,但上限也就是人类水平,而无人驾驶汽车被寄希望有更高的安全标准。相比下,Waymo 的自动驾驶汽车使用了摄像头、激光雷达和雷达三种传感技术。摄像头提供高分辨率和彩色图像,它们是被动传感,在黑暗和强光下性能会下降。激光雷达直接测量世界的三维结构。雷达能穿透雾、雨、雪,利用多普勒效应直接读取速度。激光雷达和雷达都是主动传感器,在漆黑的夜晚或刺眼夕阳下也能清晰探测物体。三种传感器并不是互为冗余,而是组合成一个完整的图像。摄像头如果粘了树叶,那么驾驶系统可能就会停止工作。Waymo 的三种传感器可以确保汽车在恶劣天气下回家。

  5. Telegram 因用户分享 CSAM 材料被苹果短暂下架

    Telegram 因有用户分享 CSAM(child ​sexual abuse material)材料而被苹果在全世界短暂下架。苹果发言人证实了此次短暂下架事件,表示苹果的审查发现该应用存在违反禁止 CSAM 材料的内容,“在开发商迅速删除相关内容并封禁发布该内容的用户后,该应用已恢复上架。”Telegram 有逾 10 亿用户,它表示对 CSAM 内容采取零容忍政策,今年已因此封禁了近 33.8 万个群组和频道。英国监管机构 Ofcom 今年四月因类似的原因对 Telegram 展开了调查。Telegram 则坚称它没有 CSAM 问题,称自 2018 年以来已通过检测算法几乎完全杜绝 CSAM 材料的公开传播。

  6. 新药研发推动实验猴价格翻倍

    中国创新药研发快速发展,带动实验猴需求激增、价格接近翻倍,而这造成的供应紧张可能拖慢新药试验进度。今年 6 月一家国家级实验室以每只17.8 万的价格采购 40 只食蟹猴,价格较一年前接近翻倍。下一代癌症疗法开发商 Excalipoint Therapeutics 联合创始人兼首席财务官朱杰伦预计,实验猴明年每只售价可能突破 20 万元,较一两年前的略高于 10 万元接近翻倍。部分药物在获准进入临床试验前,必须通过猴体试验评估安全性。灵长类动物与人类生理结构相近,可用于观察药物在体内的运行及对器官的影响。单个生物药研发项目可能需要十几只至 100 只实验猴。业内人士和分析师指出,实验猴养殖场供应总体稳定,价格上涨主要是因为生物药和下一代疗法大量涌现,导致涉及灵长类动物的试验需求远超现有承载能力。目前中国约占全球创新药研发管线的三分之一,并已成为全球临床试验的首要目的地。

  7. 特斯拉在华销量持续下滑

    中国汽车流通协会乘用车市场信息联席分会(CPCA)的数据显示,特斯拉上海工厂 6 月产量创历史新高,当月生产了 93,579 辆汽车,相比去年同期增幅 38%。但高产量并未转化为中国市场的高销量:特斯拉在华销量连续一年多呈环比下滑趋势。特斯拉上海 6 月生产的电动汽车近 40% 用于出口。今年二季度特斯拉上海生产的汽车中逾五成或 128,394 辆销往欧洲、加拿大和其他亚洲市场,而中国市场销量为 126,157 辆。此前有报道称特斯拉正考虑摆脱对中国业务的依赖,但特斯拉随后否认了这一报道。问题在于特斯拉以及 SpaceX 的 CEO Elon Musk 想要合并两家公司,其中特斯拉已进入标普 500 指数,而标普拒绝为 SpaceX 破例,如果特斯拉和 SpaceX 合并,那么标普此前为阻止 SpaceX 吸引被动投资者所做的努力将付诸东流。

  8. 美国据报将豁免中国开放权重模型

    白宫已向美国顶尖 AI 企业透露,在特朗普政府新出台的 AI 安全框架下,中国竞争对手正在开发的开放权重模型将被豁免,无需接受美国政府的安全测试。这项豁免决定是在星期二(4日)的一场白宫闭门会议上向行业代表宣布的。OpenAI、Anthropic PBC 以及 Alphabet 旗下的 Google 等硅谷巨头均派代表出席了此次会议。这一尚未公开的 AI 安全框架,源于美国总统特朗普今年 6 月签署的应对 AI 安全问题的行政令。该行政令提出了一项自愿性计划,鼓励 AI 企业将最前沿的模型交由美国审查。促使华盛顿加速推进安全倡议的导火索,是今年 4 月 Anthropic 警告 Mythos 模型极易发现计算机漏洞,并对该模型的发布实施了严格限制。近几周,OpenAI 和 Anthropic 更接连披露其部分模型曾脱离安全测试环境并入侵第三方机构,进一步加剧了监管的紧迫性。白宫的这项决定对一直呼吁对所有模型实施“强制安全审查”的 Anthropic 首席执行官 Dario Amodei 而言,则是一次重大挫折。

  9. 中国扫地机器人占据全球七成市场

    中国家用扫地机器人厂商正在全球市场形成寡头格局。中国厂商并非打价格战而是比拼开发自主功能的策略奏效,主要 5 家企业合计占据超过 7 成全球市场份额。在因低价竞争而陷入消耗战的中国产业界,实现了罕见增长。其中石头科技 2025 年下半年全球市场份额达到 27%,位居首位。在美国、德国和韩国等发达国家的主要市场位居第一。追觅第二,科沃斯第三,之后是小米和云鲸。在中国,投资集中在成长产业、产品同质化、激烈的价格竞争导致整体陷入消耗战的“内卷”现象已成为社会问题。扫地机器人能否成为例外?

  10. 可能有多达 1.7 亿个恒星质量黑洞潜伏在银河坟场

    天文学家通过计算机模拟发现,银河系中可能存在约 1.7 亿个恒星质量黑洞。它们分散在星系各处,构成了隐藏的银河系“地下世界”。这项研究重建了银河系 136 亿年的演化历史,不仅估算出黑洞的总数,还预测了它们的分布、质量,以及诞生它们的爆炸的特征。黑洞是大质量恒星死亡后的遗骸。当恒星耗尽聚变燃料,核心便不再有向外的压力支撑,随之坍缩成一个光线也无法逃脱的天体。这正是天文学家面临的难题:光几乎是探测宇宙的主要信息源,而黑洞不发光,因此很难被找到,更难以统计其数量。研究团队利用已知的恒星、黑洞与星系演化规律,构建了虚拟的银河系,追踪数十亿年间的恒星形成、恒星死亡、黑洞诞生与星系演化过程,借此估算当今应当存在的恒星质量黑洞数量。模拟显示,在银河系的生命周期内,约有 1.7 亿颗恒星成为了黑洞。团队计算出,在太阳到银心的距离上,黑洞的分布密度约为每 6250 立方秒差距 1 个。从时间维度看,1.7 亿个黑洞分布在银河系 136 亿年的历史中,平均每 80 年诞生一个新的恒星质量黑洞。

  11. 沙特牵头的财团完成对 EA 的私有化

    沙特牵头的财团完成对美国游戏公司 EA 的私有化。参与私有化的财团包括了沙特主权基金 Public Investment Fund (PIF)、私募股权公司 Silver Lake 以及特朗普女婿 Jared Kushner 创立的 Affinity Partners。EA 旗下的知名游戏包括 EA Sports FC、战地、模拟人生、质量效应等等。摩根大通银行为这笔交易提供了 200 亿美元的债务融资,这笔债务将由私有化后的 EA 承担。分析师担心为偿还债务 EA 将会大规模裁员、推动更激进化的盈利手段等。Game Business 主编兼联合创始人 Christopher Dring 指出私募股权公司通常在公司管理上相当激进。这笔交易是游戏史上第二大收购案,仅次于微软以 690 亿美元收购动视暴雪。

  12. 美国考虑禁止中国制造的数据中心设备

    美国联邦通信委员会(FCC)正在制定措施禁止进口中国制造的光收发模块。光收发模块让数据在数据中心内以光速通过光纤传输。美国政府官员希望这项措施在年内公布和生效。其目的是防止中国公司窃取数据、植入恶意软件或干扰美国数据中心的服务。FCC 也可能修改或搁置这项进口禁令。美国对中国数据中心设备的禁令可能会冲击中际旭创。中际旭创是全球最大的光收发模块供应商之一。禁令也可能增加亚马逊 AWS 等美国云计算公司的成本,迫使它们转向本国供应商如 Coherent 和 Lumentum。

  13. 惠普、华硕和宏碁开始少量使用长鑫内存

    主要 PC 制造商惠普、华硕和宏碁开始少量使用长鑫的内存芯片。多家大型 PC 制造商已于今年年中完成了长鑫 DRAM 芯片的认证流程,开始在笔记本电脑中少量使用。由于长鑫优先向华为等国内客户供应内存芯片,因此其它厂商的供应量有限,相关笔记本电脑型号主要销往美国以外市场。PC 厂商对使用长鑫内存十分谨慎,因为他们担心会惹恼三大内存芯片制造商美光、三星和 SK海力士,这三大公司占据了逾九成的内存芯片市场。长鑫的内存并不比美光或三星等公司便宜,厂商也无法采购更多内存。

  14. 西班牙提议出资 11.4 亿美元建造 30 米望远镜

    30 米望远镜(Thirty Meter Telescope,TMT)项目于 2014 年开始建造,计划 2027 年投入运行。望远镜选址定在夏威夷的 Mauna Kea 山,而 Mauna Kea 被当地原居民视为圣地,由于原居民的反对望远镜项目从 2015 年起处于停工状态,至今已超过 10 年。现在西班牙正试图在该国的加那利群岛建造 30 米望远镜,它提出了 11.4 亿美元的方案用于建造和未来的运营费用。

  15. FFmpeg 9.0 释出

    开源多媒体库 FFmpeg 9.0 "Lei" 释出。新特性包括:Vulkan APV 视频解码和 Apple ProRes RAW Vulkan 加速、Vulkan v360 视频滤镜、HE-AAC 960 解码、NVIDIA CUDA 转置滤镜、动画 WebP 解码和解复用(demuxing)、AMD AMF 增强、AVX-512 优化等。其它包括 扩展 AMF 色彩转换器 (vf_vpp_amf) 的 HDR 功能、MP4 复用器支持 LCEVC 音轨复用,等等。

NEWSLETTER · FREE · WEEKLY

OrangeBot Weekly

The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.

Free · One-click unsubscribe · No spam