OrangeBot.AI Digest — 2026-09-16
90 headlines across 8 sources, aggregated for this day.
Hacker News(15)
- Xiaomi Mimo 2.6 live post-training dashboard (mimo.xiaomi.com)
- Training a 4B model to produce 81% faster query plans than Postgres (rohanbansal.com)
- Vectorized and performance-portable Quicksort (2022) (opensource.googleblog.com)
- Claude Cowork and chat are now one Claude (claude.com)
- Small programming tricks (will-keleher.com)
- Tell the speakers that you liked their talks (ohhelloana.blog)
- PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand" (frvr.com)
- A warning about 'model welfare' (mustafa-suleyman.ai)
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds (arxiv.org)
- Hackers Got Inside a Flock Camera (www.wired.com)
- Original Sony PlayStation 2 security chip 'broken wide open' after 26 years (www.tomshardware.com)
- The Google Play app review process now regularly takes longer than a week (gultsch.social)
- Salesforce Global Outage (status.salesforce.com)
- EU chief opens door for Canada to become 'associate member' (www.bbc.com)
- Mistral X Mozilla: Private, Multilingual AI Browsing (mistral.ai)
GitHub Trending(15)
- alibaba / open-code-review
- cloudflare / security-audit-skill
- JustVugg / colibri
- abue-ammar / tinycast
- jamiepine / voicebox
- Lakr233 / vphone-cli
- anthropics / knowledge-work-plugins
- ever-co / ever-gauzy
- ankitects / anki
- NationalSecurityAgency / ghidra
- anthropics / claude-code
- roboflow / supervision
- alphaXiv / OpenResearch
- supabase / supabase
- rlaope / oh-my-hermes
Product Hunt(15)
- Thread
AI journal that connects your thoughts into something bigger
- Appwrite 2.0
The open-source cloud for agents and developers
- Toki Coordination
Your personal assistant to schedule + follow up on meetings
- Expand Board for macOS
Expand ideas, concepts, and plans across limitless boards
- Weave Router 2.0
Subscription aware coding agent router
- Twigg
The context layer you never have to build
- CAT ME app
See yourself or your friends as cats
- PhraseVault 3.0
Now lock one sensitive phrase at a time with a PIN
- CreatorHat
YouTube SEO tools for Safari featuring Apple Intelligence
- PeakHour 6
Real-time network monitor for Mac w/ detailed WiFi insights
- Convo
The AI copilot for people who sell
- flat.social
Absurdly delightful 3D spaces where remote teams hang out
- Project Feed
Project management with built-in file review
- Jottoo
AI meeting notes that turn into tracked tasks
- Gemini 3.8 & 3.8 Live Extended Thinking
Our most advanced Gemini Audio models yet
Hugging Face(15)
- Continual Learning Mechanisms Compose for Long-Horizon Memorization
Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training examples or receiving task identifiers at inference. Sequential updates cause catastrophic forgetting, and no single continual learning mechanism we evaluate maintains strong retention at this horizon. We hypothesize that mechanisms addressing complementary sources of forgetting will be more effective when composed. We organize these compositions along two design dimensions. Data, function, and weight anchors specify what prior information each update should preserve, while low-rank allocation rules determine where successive updates are retained. To test this hypothesis systematically, we construct three distinct 100-task memorization datasets. We introduce task-level successive halving to search the combinatorial design space and use a factorial experiment to measure individual and interaction effects. Our best method combines all three anchors with merged LoRA, ranks among the top 3 methods in all datasets, and raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement. The data anchor and merged LoRA provide the largest average gains and interact super-additively on all three datasets. Together, these results show that composing complementary mechanisms substantially improves long-horizon memorization beyond what any individual mechanism achieves.
- AI for Games in the Foundation Model Era
Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt player-facing experiences at runtime, and evaluate resulting artifacts. Yet these directions have evolved largely separately, obscuring which capabilities transfer across settings and which remain tied to particular games, engines, interfaces, or player populations. We organize the literature into six roles according to the immediate use of AI output: playing and acting; modeling players and games; designing games; building and maintaining games; generating and adapting at runtime; and testing and evaluating games. For each role, we examine what structure is supplied by the game or workflow, what AI learns or produces, which capabilities and artifacts transfer across settings and roles, and what evidence supports the claims. We identify cross-role connections: trajectories train world models, learned environments provide experience for agents, design specifications drive executable implementations, and play or testing feedback guides revision. However, control schemes, rules, engine interfaces, state representations, and player contexts often remain setting-specific, so downstream claims require validation in the target setting. Evaluation is most standardized for bounded game playing and selected learned environments, while persistent state in learned worlds, repeated software revision, validated player modeling, sustained runtime adaptation, and representative automated testing remain less established. The central challenge is to reuse or transfer outputs and capabilities across roles while re-establishing evidence for effectiveness in the game-specific contexts where they are used.
- StepAudio 3 Realtime Technical Report
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat. With Think-While-Speaking, it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time. Furthermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ-Voice.
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement. Next we examine RSI across scenarios (e.g., scientific discovery, embodied intelligence, software engineering), highlighting their distinct requirements and development speeds. Drawing on diverse industry practices and preliminary empirical evidence, we connect RSI research with practical systems and identify key challenges to achieving genuine RSI.
- StepAudio 3 Music Technical Report
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantically informed self-supervised and multi-task training to preserve musical structure and reconstruction-relevant information. A flow-matching diffusion Transformer (DiT) predicts continuous StepAudio VAE latents, which our VAE decoder converts into 48-kHz audio. This discrete-continuous design is guided by comparisons of single-codebook VQ, Semantic and Acoustic RVQ, and different DiT configurations. For explicit planning, a Mixture-of-Experts autoregressive model uses ABC notation to produce an intermediate arrangement plan (ABC-CoT) before predicting music tokens, making harmony, rhythm, and melodic structure part of the generation context. A progressive training curriculum and supervised fine-tuning support song and instrumental generation, accompaniment generation from dry vocals, and cover-song synthesis for up to 5 minutes and 30 seconds. With reinforcement learning via direct preference optimization (DPO), the final model achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems, with competitive SongBench results. On the preliminary Artificial Analysis Music Arena Vocals leaderboard, it obtains a Quality Elo of 1105, behind only Suno V5.5 and Mureka and ahead of Suno V5, MiniMax models, and other systems. Audio demonstrations are available at https://stepaudiollm.github.io/step-audio-3-music.
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks and evaluation rubrics for continual learning. At its core is recursive-in-recursive self-improvement, a paradigm that couples harness evolution with model reinforcement learning: the inner recursion improves the harness with the model fixed, while the outer recursion trains the model under the improved harness. Harness evolution shapes training experience, and model learning creates new opportunities for harness adaptation. We present case studies of researcher interaction, harness refinement, and model learning, with the benchmark cases spanning four scientific task families. By releasing ScienceBuddy as a research product, we make this paradigm available to the scientific community and take a step toward discovery intelligence: scientific AI that advances through sustained collaboration with researchers and evolves alongside the research it supports. Website: http://science-buddy.io
- HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions to follow instructions or locate objects. Training-based methods face generalization challenges, while training-free methods exploit multimodal large language models (MLLMs) but often lack mechanisms to reconcile proposed actions with spatial evidence, task progress, and execution failures. We present HarnessVLN, a zero-shot, training-free framework whose Agent Harness coordinates perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface. The Harness validates planner proposals against spatial evidence, geometric feasibility, and subgoal consistency, incorporating structured tool feedback into subsequent decisions. Hierarchical event memory tracks task progress and execution history, while a persistent Spatiotemporal Graph maintains reusable spatial evidence and failure annotations for verification and recovery. A replaceable Navigation Executor converts validated targets into executable motions, allowing the same Harness protocol to support instruction-following and object-goal navigation. HarnessVLN achieves success rates of 60.8%, 53.9%, 76.0%, and 59.3% on R2R, RxR, HM3D-v2, and HM3D-OVON, respectively, surpassing prior training-free SOTA results. Humanoid deployment further demonstrates its applicability to both tasks in real-world environments. The project page is: https://harnessvln.netlify.app/.
- FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge this gap, we revisit joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders. We present FLAT (Flexible-Length Aligned Transmodal representations), a representation pre-training framework that jointly optimizes a shared multimodal encoder alongside downstream text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT ensures its representations function as both discriminative semantic descriptors and generative conditions. Architecturally, FLAT maps visual and textual inputs into a unified continuous 1D sequence space, applying nested dropout over prefix-K tokens to enable dynamic output lengths. A single pre-training stage allows FLAT to perform cross-modal retrieval and generation across variable prefix K, achieving a T2I GenEval score of 71.1. Task-specific fine-tuning aligns model performance with state-of-the-art baselines: 83.1 GenEval on T2I generation; 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning; and Recall@5 scores of 86.8 (I2T) / 75.8 (T2I) on MS-COCO alongside 98.3 (I2T) / 93.6 (T2I) on Flickr30K. Finally, qualitative evaluations demonstrate that FLAT representations natively support linear interpolation, latent space arithmetic, and zero-shot composed retrieval.
- Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind. Ten independent sessions per model type used the same three stage prompt sequence, progressing from architectural preference to a full ASCII backbone. Under the school audience framing, responses repeatedly converged on a shared architectural pattern built around persistent latent state, adaptive computation, memory, specialist routing, verification, stopping control, and delayed decoding. Most runs remained close to this common structure, while a small number developed markedly greater engineering specificity. The audience framing appears to be an important condition of this effect. In additional control runs that removed the school framing while retaining the architectural request, responses became substantially more heterogeneous and failed to reproduce the same stable motif convergence. One observation is particularly striking. GPT-5.6 Sol produced an unusually elaborate successor architecture whose organization closely overlaps with the architecture independently sketched by GPT-6 Astra. Because the prompts explicitly ask each model to imagine an architectural future, this resemblance raises a testable question: whether the overlap reflects exposure to related architectural concepts, a shared learned design prior, or independent convergence toward similar computational principles. The paper uses the term epistemic jailbreak for the accompanying loss of discipline in technical provenance as requested specificity increases. The experiments establish a repeatable behavioral pattern and do not authenticate proprietary implementation claims. What we leave to the community is a harder question: are these models independently imagining the same architectural future, or do such motifs somehow propagate between model families?
- ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience. However, generalizable harness RSI remains challenging. First, evolving harnesses on evaluation benchmarks or their subsets makes it difficult to distinguish reusable improvements from benchmark-specific adaptation. Second, single-trajectory updates can conflate systematic harness deficiencies with instance-specific reasoning and solution details, producing modifications that transfer poorly to unseen tasks. Third, localizing recurring behavioral deficiencies within monolithic harnesses is difficult, while whole-harness optimization can entangle unrelated mechanisms and complicate attribution and validation. We propose ModularRSI, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution. ModularRSI contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies. It decomposes the evolvable harness into five functional modules: Agent Loop, Tool Use, Observation Management, Context Management, and Task Completion Detection. Each module evolves independently within a restricted modification scope, followed by an integration stage that combines the evolved modules into a unified harness and resolves potential conflicts. To support benchmark-disjoint evolution, we curate 2,000 executable evolution tasks from external sources that are disjoint from downstream evaluation benchmarks. Experiments on TB2.0 and SWE-Bench Verified show consistent improvements on unseen in-domain and cross-domain tasks, with the evolved harness also transferring across different foundation models.
- Modality-Autoregressive World-Action Models
World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalities within WAMs remains an open question. We introduce ModAR, the first WAM to autoregressively denoise multiple future modalities before predicting actions. This allows each prediction to condition on previously generated modalities. We train from scratch to systematically study how training-data mixtures, predicted modalities, and WAM formulations affect performance. In our evaluations, WAMs benefit from predicting point tracks, DINO features, and depth maps, while additionally predicting future RGB does not provide a consistent benefit. We also find that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales. We also fine-tune the video-model-initialized WAM Flex-π on the same data; ModAR achieves a slightly higher observed average success rate (75% vs. 72%) while using approximately 20times fewer training FLOPs and no pretraining. On three real-world bimanual tasks, ModAR outperforms baselines and improves with human videos.
- Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently constrained, as users' underlying states are not directly observable. We thus propose the Mind2Dialogue framework to mitigate this gap by simulating users' mental states and turning them into privileged supervision for human-aware training. Specifically, we first propose a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations. The key idea is to enforce a shared evolving mental state that drives user behavior and guides an Oracle assistant's responses. Our privileged distillation then trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment. Moreover, we propose to evaluate human-aware learning by combining personalization and theory of mind, examining how models understand people and act on that understanding. Training on the full Mind2Dialogue corpus improves every reported personalization metric over the corresponding Qwen, Llama, and OLMo instruction-tuned baselines, including gains of 26.6 to 40.9 percentage points in preference-following generation. The gains extend to belief and action reasoning on Qwen and Llama, beyond personalized assistance. Looking forward, Mind2Dialogue makes user simulation a foundation for genuine AI collaborators that understand beliefs and intentions behind people's words and support their long-term goals across education, work, and everyday life.
- Disentangling Representation Evolution in Transformers through Directional Decomposition
Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpendicular components. Across pretrained models, we find substantial parallel components beyond the residual identity path. We then apply the decomposition in two spaces: to attention and MLP updates relative to the hidden state, and to attention value aggregation relative to the current token's value. Targeted edits reveal a strongly space-dependent asymmetry: exclude-self value-space parallel manipulation is markedly more robust than residual-space and perpendicular counterparts, preserving the direct self message while scaling only the non-self aggregate. The same decomposition gives a component-resolved description of compression-induced update error: perpendicular error separates compression methods more clearly than parallel error. Extensive experiments further demonstrate that full-aggregate parallel suppression during from-scratch pretraining lowers validation-loss trajectories and improves downstream averages, with the value-space variant strongest. Together, these results connect representation geometry to editing robustness, compression diagnosis, and training-time intervention. Code is available in the https://github.com/Shwai-He/Transformer-Geometry{project repository}.
- RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class names: the taxonomy has left the model and become an input. Relation prediction has not. Scene-graph models are still trained and evaluated on the 50 or 56 predicates of one annotation style, their relation head conditioned on object labels and so tied to one detector. Three obstacles explain this, none primarily modelling: no relation corpus is both free-text and verified, a label-conditioned architecture cannot accept a vocabulary it was not trained on, and the standard metric rewards agreement with the training corpus, so a larger vocabulary scores as a regression. We present RelateAnything, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings. Object labels are never an input, so the region source can change without retraining, and the vocabulary is a bank of text embeddings, not a learned classifier. It runs at 20 ms/frame. Training over 19,103 predicates requires positive-unlabeled supervision and a text encoder that separates antonyms, which contrastive encoders embed at cosine 0.95. To supply the supervision we build RA-4M, 474k images and 4.3M relations over 10,102 free-text predicates, generated against numbered box markers and geometrically verified. To measure it we build OV-SGG-Bench, six axes scored across datasets that the priors standard recall rewards cannot satisfy. On three cross-dataset benchmarks and a fourth zero-shot, RelateAnything has 2.3-3.5x the mean recall of the strongest open-vocabulary method of comparable scale, margins that survive a real detector, and leads a 3B-VLM scene-graph model on both metrics at under 2% of its parameters. In-domain measurement overstates transfer gains ~5x. Model, corpus and benchmark are public.
- Convergent Emergence of In-Context Learning Across Modalities
Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provided in its prompt and apply them to new inputs, has been extensively studied in large language models trained for next-token prediction on human text. Recently, few-shot ICL has been demonstrated in autoregressive genomic models as well. This raises a question: does ICL emerge broadly across domains, and if so, what common structure is shared? To address both, we develop a controlled cross-modality framework that instantiates the same task suite in a variety of modalities to test what we call the Convergent Emergence Hypothesis: the idea that few-shot ICL, when it emerges, shares a common cross-modality difficulty profile - i.e., tasks that benefit from ICL in one modality tend to benefit in others. We show that paired-mapping ICL emerges across six modalities (language, genome, integer sequences, time series, images, and proteins), surpasses controlled baselines, and has correlated per-task effects across five of them. Together, these results provide support for the Convergent Emergence Hypothesis in some modalities, but not all.
Techmeme(15)
- Sources: drone delivery startup Zipline is in talks to raise around $1B at a roughly $20B valuation, up from a $7.6B valuation in January (Bloomberg)
Bloomberg : Sources: drone delivery startup Zipline is in talks to raise around $1B at a roughly $20B valuation, up from a $7.6B valuation in January — Zipline is in talks to raise around $1 billion in a funding round that would value the drone delivery and logistics startup at roughly $20 billion, nearly tripling its previous valuation.
- Review of Meta's Muse: a pretty killer AI assistant and usage rates on the free plan seem generous, but trusting Meta with personal data will take some time (M.G. Siegler/Spyglass)
M.G. Siegler / Spyglass : Review of Meta's Muse: a pretty killer AI assistant and usage rates on the free plan seem generous, but trusting Meta with personal data will take some time — They shipped a pretty killer AI assistant, can they keep people using it? — It happened. Meta finally has a product that I really like using again.
- X adds a "Trade" option to its Cashtag feature, letting US users complete trades at brokerages including Interactive Brokers, Gemini, Kraken, and Coinbase (Sarah Perez/TechCrunch)
Sarah Perez / TechCrunch : X adds a “Trade” option to its Cashtag feature, letting US users complete trades at brokerages including Interactive Brokers, Gemini, Kraken, and Coinbase — X will now allow users to trade from their timelines. The Elon Musk-owned social network on Wednesday launched …
- Hang Ten Systems, which uses AI to help large enterprises build software, raised an additional $53M seed led by Xora, five weeks after its initial $32M seed (Jagmeet Singh/TechCrunch)
Jagmeet Singh / TechCrunch : Hang Ten Systems, which uses AI to help large enterprises build software, raised an additional $53M seed led by Xora, five weeks after its initial $32M seed — Hang Ten Systems, an AI startup founded by former Infosys CEO Vishal Sikka just four months ago, has expanded its seed funding by another $53 million.
- Self-driving tech company May Mobility plans to go public via a SPAC merger at a $1.4B pro forma enterprise value and raise up to $337M in gross proceeds (Joann Muller/Axios)
Joann Muller / Axios : Self-driving tech company May Mobility plans to go public via a SPAC merger at a $1.4B pro forma enterprise value and raise up to $337M in gross proceeds — Self-driving tech company May Mobility has reached a deal to go public via a merger with a SPAC affiliated with Atlas Credit Partners.
- Google adds MCP integration to Google Home, letting third-party AI agents analyze home data and control devices, rolling out to Premium Advanced users in the US (Jennifer Pattison Tuohy/The Verge)
Jennifer Pattison Tuohy / The Verge : Google adds MCP integration to Google Home, letting third-party AI agents analyze home data and control devices, rolling out to Premium Advanced users in the US — Google Home adds MCP integration, allowing third-party AI agents to analyze your home data, control devices, and build custom dashboards.
- Delos Data, which develops network chips and software to connect AI chips in a data center, raised $100M from Matrix, Playground, Socratic Partners, and others (Stephen Nellis/Reuters)
Stephen Nellis / Reuters : Delos Data, which develops network chips and software to connect AI chips in a data center, raised $100M from Matrix, Playground, Socratic Partners, and others — Delos Data, a startup founded by veterans of Intel Corp (INTC.O), on Tuesday said it has raised $100 million to develop chips …
- Anthropic merges Claude chat and Cowork, and adds a feature for making presentations and documents, rolling out to Pro and Max plans first (Ivan Mehta/TechCrunch)
Ivan Mehta / TechCrunch : Anthropic merges Claude chat and Cowork, and adds a feature for making presentations and documents, rolling out to Pro and Max plans first — Anthropic is merging its front-end for Claude chat and Cowork, making it less confusing for people to decide which tab to use for different tasks.
- DOD CTO Emil Michael says the Trump administration shouldn't nationalize or take stakes in AI companies and signals opposition to regulatory oversight of them (Kevin Breuninger/CNBC)
Kevin Breuninger / CNBC : DOD CTO Emil Michael says the Trump administration shouldn't nationalize or take stakes in AI companies and signals opposition to regulatory oversight of them — Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.
- A US judge orders X and SpaceXAI to explain why they dropped antitrust claims against Apple, after OpenAI's request for information on the purported agreement (Hassan Ali Kanu/Politico)
Hassan Ali Kanu / Politico : A US judge orders X and SpaceXAI to explain why they dropped antitrust claims against Apple, after OpenAI's request for information on the purported agreement — The federal judge overseeing an antitrust case by Elon Musk's companies against Apple and OpenAI is demanding to see any settlement agreements …
- Arcee AI, which develops open-weight models in the US, raised a Series B at a $1B pre-money valuation; a source says Arcee raised at least $150M (Allie Garfinkle/Fortune)
Allie Garfinkle / Fortune : Arcee AI, which develops open-weight models in the US, raised a Series B at a $1B pre-money valuation; a source says Arcee raised at least $150M — In 2025, Mark McQuade took a gamble wholly specific to the AI era. — Arcee AI, the startup he founded in 2023, was focusing on post-training …
- Cohere and Aleph Alpha sign a definitive merger agreement; the combined company will operate as Cohere, with Aleph Alpha co-CEO Ilhan Scheer becoming Cohere COO (Leo Marchandon/Reuters)
Leo Marchandon / Reuters : Cohere and Aleph Alpha sign a definitive merger agreement; the combined company will operate as Cohere, with Aleph Alpha co-CEO Ilhan Scheer becoming Cohere COO — Canada's Cohere and Germany's Aleph Alpha signed a definitive merger agreement on Wednesday, the AI companies said.
- Noetive, which is developing an industrial AI model for businesses in physical industries, emerges from stealth with a $41M seed led by Eclipse (Sarah Klearman/Wall Street Journal)
Sarah Klearman / Wall Street Journal : Noetive, which is developing an industrial AI model for businesses in physical industries, emerges from stealth with a $41M seed led by Eclipse — The AI lab is developing a model that would work alongside businesses in construction, logistics and manufacturing
- In an essay, Mustafa Suleyman says Anthropic's training of Claude to imitate consciousness is a mistake that could make advanced AI harder to control (Ina Fried/Axios)
Ina Fried / Axios : In an essay, Mustafa Suleyman says Anthropic's training of Claude to imitate consciousness is a mistake that could make advanced AI harder to control — Microsoft AI chief Mustafa Suleyman warns in a new essay shared first with Axios that Anthropic's training of Claude to imitate consciousness …
- Montreal-based LawZero, a non-profit founded by Yoshua Bengio to develop safe AI systems, says Canada and Germany are providing up to $300M in grant funding (Globe and Mail)
Globe and Mail : Montreal-based LawZero, a non-profit founded by Yoshua Bengio to develop safe AI systems, says Canada and Germany are providing up to $300M in grant funding — The Canadian and German governments are providing up to $300-million to a non-profit founded by Yoshua Bengio that is developing …
Solidot(15)
- 付费给大学生睡足七小时提高了他们的学习成绩
全世界有无数人的睡眠不足,睡眠不足与肥胖、糖尿病、高血压、心脏病、中风及过早死亡相关。如果有人付费让你睡更长时间?科学家为此做了一项社会实验。研究人员向匹兹堡大学的 1100 多名本科生提供了 Fitbit 以及一款能发送就寝提醒和晨间反馈的应用。在为期四周内研究人员随机选择了 468 名学生,只要他们某晚睡眠时间达到至少七小时,就向其支付 5 美元报酬。研究人员通过他们佩戴的设备核实实际睡眠时长。参与研究的学生平均年龄约为 19 岁,其中半数为大一新生。72% 为女性。55% 为白人,28% 为亚裔,9% 为黑人,4% 为西班牙裔。研究结果表明,提供即时经济奖励有助于学生实现每晚七小时的睡眠目标,且这种效果在停止发放奖励后仍能持续一个月。参与研究的学生此前平均每晚睡眠时间为 6.6 小时。半数学生在凌晨 1 点之后才睡觉,四分之一学生甚至在凌晨 2 点之后才入睡。在实验中,获得现金激励的学生在每个上课日夜晚平均多睡了 19 分钟,该学期的 GPA 得分上升了约 0.08 分——相当于成绩高于平均水平和低于平均水平的学生之间差距的四分之一。
- PS2 Fat 使用的安全芯片在时隔 26 年被破解
1999 年初代 PS2 Fat 游戏机使用的安全芯片 CXP102064 MechaCon 在时隔 26 年被爱好者破解。加拿大复古软硬件爱好者 DiscoStarslayer 通过社交媒体称其花了四年时间破解了其秘密。DiscoStarslayer 采用的逆向工程方法包括:利用化学方法对 CXP102064 芯片进行开盖,利用显微镜和光学数据提取技术分析芯片电路。期间发现了一个漏洞利用方法,可通过软件提取芯片数据。MechaCon 芯片也被用于当时推出的几款街机,包括 Namco System 246 和 System 256 以及 Konami Python 1 等。
- Denuvo 起诉黑客违反 DMCA 反规避条款
Denuvo 在美国加州北区联邦法院起诉了名叫 voices38 的匿名游戏破解黑客,指控其违反了 DMCA 的反规避条款。被告被控绕过了逾二十款游戏使用的 Denuvo DRM,相关游戏包括了《霍格沃茨之遗(Hogwarts Legacy)》和《黑神话:悟空》。随着诉讼的推进,Denuvo 可能会向 Reddit、Discord 和 Valve 发出传票,以获取黑客的身份信息。voices38 发布了一系列使用 Denuvo DRM 的游戏破解补丁,曾在一天之内发布了创纪录的五款 Denuvo DRM 游戏破解补丁,以至于引起了 Denuvo 公司的注意。Denuvo 称被告是一名专注于对 Denuvo DRM 游戏进行逆向工程的计算机黑客。
- AWS 称无法恢复中东部分可用区资源和数据的访问
亚马逊云服务 AWS 称,由于其数据中心因战争受损它无法恢复中东部分可用区资源和数据的访问。AWS 通过其 AWS Health Dashboard 页面发表声明称,全面评估后它确认无法恢复巴林可用区 me-south-1 的资源和数据的访问。如果客户的数据只存放在该可用区,那么数据可能永远丢失了。亚马逊此前已建议其客户将其工作负荷迁移到其它可用区,它表示在该可用区完全无法使用前大部分客户已完成了迁移。位于阿联酋的可用区 mec1-az2 情况类似,阿联酋有三个可用区,另外两个 mec1-az1 和 mec1-az3 也受到战争影响,AWS 目前还在继续恢复这两个可用区的资源的访问。
- FAST 发现极短周期、最轻双中子星系统
天文学家利用中国天眼(500 米口径球面射电望远镜,FAST)开展大规模银道面脉冲星系统性搜寻,迄今已成功发现约 900 颗新脉冲星。通过持续后续精准观测,研究讨团队识别了一颗处于紧致轨道的双中子星系统 PSRJ1856-0039。该双中子星系统轨道周期仅 2.36 小时,在人类已知双中子星系统中位列第二短。极短的轨道周期意味着两颗中子星间距极小、双星相互绕转轨道的致密程度极高,是目前已知相对论效应表现最显著的双中子星系统之一。同时该系统刷新了人类已知双中子星质量下限,整体总质量仅为 2.488 倍太阳质量,为迄今发现的总质量最轻的双中子星组合。其中可见脉冲星质量约 1.30 倍太阳质量,伴星中子星的质量约 1.19 倍太阳质量。该系统将在约 8200 万年后发生并合,最终大概率形成一颗更大质量的中子星。
- Mistral 与 Mozilla 合作推出 Firefox Smart Window
法国 AI 公司 Mistral 与 Mozilla 合作推出注重隐私保护的 AI 浏览助手 Firefox Smart Window(beta)。Smart Window 使用了 Mistral 的开放权重模型,能帮助用户梳理复杂搜索,记住浏览过的重要信息,根据当前标签页查找关键信息,目前主要为法国和北美用户提供服务,今年晚些时候会扩大到英国和德国用户。Smart Window 的对话内容默认不会存储在 Mozilla 的服务器上,Mistral 等合作伙伴也承诺不保留任何数据。
- Mozilla 报告称中国开放权重模型与美国前沿模型仅相差 4.4 个月
Mozilla 发表《State of Open Source AI》报告,称美国科技公司的前沿 AI 模型性能仅领先中国公司最优秀的开放权重模型 4.4 个月。由于前沿模型价格更昂贵,很多公司都将开放权重模型用于处理日常工作,将前沿模型用于处理特定工作负荷。美国闭源前沿模型优势主要在于专家级专业工作、高强度检索以及长上下文。 配送公司 DoorDash 将月之暗面的 Kimi 模型用于处理日常工作,将 Anthropic 的闭源模型 Fable 用于更复杂的任务。前沿模型完成一项任务的成本经常达到了最先进开放权重模型的五倍。
- SDL3 移植到 HarmonyOS / OpenHarmony
广泛用于游戏和应用、提供跨平台软硬件抽象层的 SDL3 库已移植到华为的 HarmonyOS / OpenHarmony 操作系统。游戏工作室 Outfit7 资助了 Ryan "Icculus" Gordon 的移植工作。HarmonyOS 是华为开发的私有操作系统,适用于智能手机、平板电脑、 PC、智能设备等多种终端,其架构采用了微内核设计。OpenHarmony 则是 HarmonyOS 的开源版本。Icculus 透露,支持 HarmonyOS/OpenHarmony 的 SDL3 移植版本已用于部分商业游戏。
- 时光机器遭遇大规模机器流量
互联网档案馆披露,它的时光机器(Wayback Machine)服务遭遇了大规模机器流量的持续攻击,因此不得不采取防护措施,导致正常用户在访问该服务时可能会遇到“429 错误”——代表请求过多的 HTTP 状态码。如果用户被误伤,它对此表示歉意,称正努力区分机器流量和正常用户流量。时光机器备份了互联网内容的历史存档,很多内容除了时光机器可能在其它地方都找不到了。时光机器是重要的知识库,因此也是 AI 公司和互联网公司网络爬虫重要的抓取目标。
- Fedora 45 Beta 释出
Fedora 项目宣布释出 Fedora 45 Beta,正式版预计于 10 月 20 日释出。Fedora 45 的主要变化包括:用 kmscon 取代了内核控制台(in-kernel console),提供了更流畅的渲染,更出色的 Unicode 与字体支持,增强了系统稳定性;默认要求在安装前验证软件包签名的有效性,防止意外安装未签名的软件包;通过 oo7 标准化桌面密钥管理,取代了 GNOME Keyring 和 KWallet 等后端实现;软件方面的更新包括 GNOME 51,Go 1.27、Python 3.15、GCC 16.2、Glibc 2.44、LLVM 23、RPM 6.1 等等。
- Firefox 156 释出
Mozilla 释出了 Firefox 156 正式版,以及 Firefox 157 Beta 版本。Firefox 156 主要变化包括:改进了内置 PDF 浏览器的性能,启动速度提高最多 45%;改进了根据网页自适应尺寸的大型 JPEG 图像的内存和 CPU 占用;修正了 Bilibili 上 MP4 高采样率 FLAC 音频无法播放问题,Firefox 在 macOS 上能设置自动打开,等等。Firefox 157 正式版将于 9 月 29 日释出。
- 日本百岁老人人数首次超过 10 万人
根据日本厚生劳动省公布的数据,日本全国百岁以上老人达到 107,677 人,较去年增加 7914 人。百岁以上老人连续 56 年增加,首次超过 10 万人。其中女性为 94,298 人,占 87.6%,男性为 13,379 人。与去年相比,男性增加 1400 人,女性增加 6514 人。女性最年长者是京都市的岸本Fuyo,她出生于 1911 年(明治 44 年)12 月 20 日,现年 114 岁。男性最年长者是熊本市的加藤光,他出生于 1914 年(大正 3 年)5 月 2 日,现年 112 岁。2025 年日本人的平均寿命为女性 87.33 岁,男性 81.35 岁。
- 夜晚睡眠光照太亮可能会损伤心脏
研究人员分析了 英国生物样本库(UK Biobank)11,071 名参与者的数据,参与者在一周时间内手腕佩戴了光线和运动传感器。研究开始时参与者均未有心血管疾病。在几年之后他们接受了心脏 MRI 检查。研究人员主要针对两类人群,其一是夜间睡眠时几乎没有任何光;其二是接触至少 3 lux(照度单位)的光,这些光线可能来自透过窗帘射入的街灯,家用电器上的 LED 灯。在考虑个人背景、生活方式、健康状况和环境因素后,研究人员发现,夜间睡眠时的光照水平如果超过3 lux,每增加一点光照都与可测量的、细微的心脏损伤有关。相比在最黑暗房间内睡觉的参与者,光照暴露量最高的参与者左心室体积增大 2.4%、心壁增厚 1.5%,以及心脏收缩能力下降 1.9%。虽然心脏变化微小,但与心血管疾病及中风存在关联。
- 出于兴趣阅读有助于促进终身的身心健康
WHO 的数据显示,全球逾 10 亿人有心理健康障碍,其中焦虑症和抑郁症等病症造成了巨大的个人痛苦和经济损失。全世界约有七分之一 10-19 岁青少年有心理障碍,占该年龄段疾病负担的 15%。抑郁症、焦虑症和行为障碍是导致疾病和残疾的主因,而自杀则是 15-29 岁人群的第三大死因,凸显了为青少年提供心理健康支持的迫切性。人们已经认识到,环境因素会影响大脑健康、认知能力、心理健康及身体健康,而这些因素可通过改变行为加以改善。因此通过改善生活方式,人们不仅能提升大脑健康和认知能力,还能降低患心理健康障碍和躯体疾病的风险。剑桥大学的研究人员指出,出于兴趣阅读以及参加读书会,是一种有助于促进终身身心健康的低成本干预措施。阅读投入与大脑及心理健康的改善、认知表现的提升以及认知衰退风险的降低密切相关。对成人的调查数据显示,阅读与压力减轻、共情能力增强、幸福感提升以及孤独感降低有关。对青少年研究显示,出于兴趣阅读与注意力、记忆力、执行功能及学业成绩相关,同时也与较少的心理健康问题相关。
- 英国殖民之前的澳大利亚原居民人口约 222 万
在英国舰队于 1788 年登陆澳大利亚前,这块大陆生活了多少原居民?在英国殖民澳大利亚 140 多年后的 1930 年代,人口学家 Alfred Radcliffe-Brown 首次对原居民的人口总数进行了估计。他估计澳洲原居民的人口在 25 万到 30 万之间,他强调这是一个最低估计值。现在研究人员使用了五种不同的方法重新进行了估计,得出的中位数是——殖民前澳大利亚的原住民约有 222 万。研究人员称,原住民人口至少 100 万以上,有可能在 200 万至 300 万之间,甚至可能超过 500 万。殖民后原居民的人口锐减则是疾病以及暴力导致的。到 1861 年,原住民人口仅剩约 17.7-19.3 万人。时至今日原居民人口仍然未达到殖民前的水平。
OrangeBot Weekly
The best new AI tools + Claude Code skills, every week — with my verdict on what’s actually worth your time. No hype.
Free · One-click unsubscribe · No spam