WEEK · 2026-W38

Weekly Digest — 2026-W38

215 unique stories (2026-09-142026-09-20), aggregated across 8 sources.

Hacker News(42)

  1. Dario, Please (pop.rdi.sh)
  2. iOS 27, iPadOS 27, and macOS 27 (www.apple.com)
  3. Steam Frame starts at $1059 (store.steampowered.com)
  4. Pion, an agent designed to run any company autonomously (andonlabs.com)
  5. Distributed Systems Classics (2017) (nvartolomei.com)
  6. Microsoft patches Windows and Excel – breaks audio, remote access, and paste (www.theregister.com)
  7. Introducing System One Models and Jev (typesafe.ai)
  8. Gemini 3.8 Live and 3.8 Live Extended Thinking (blog.google)
  9. An Update on Wayback Machine Access (blog.archive.org)
  10. Most people prefer traditional architecture (www.worksinprogress.news)
  11. America's Driver's License Breach Is a National Security Disaster (www.lawfaremedia.org)
  12. There's a 100% Chance AI Agents Are Ruining the Internet (www.404media.co)

GitHub Trending(24)

  1. JustVugg / colibri
  2. alibaba / open-code-review
  3. multimodal-art-projection / YuE
  4. debpalash / VoiceStudio
  5. 666ghj / MiroFish
  6. Panniantong / Agent-Reach
  7. ever-co / ever-gauzy
  8. Homebrew / BrewUI
  9. melgarafael / DeskcommCRM
  10. cloudflare / security-audit-skill
  11. abue-ammar / tinycast
  12. jamiepine / voicebox

Product Hunt(42)

  1. Aside

    AI browser that actually gets work done for you

  2. Slashy Assistant

    The AI assistant that does email for you

  3. Deplo

    A simple-to-use alternative to cloud deployments

  4. TryCase

    AI tests your PRs. Get a video walkthrough before you merge.

  5. MemoryPet 2.0

    Turn your browsers toolbar into an animated usage monitor

  6. Hello Inbox

    Get more marketing emails into the inbox

  7. Narrative

    AI-first video editor, just describe edits & refine in chat

  8. OpenAI Agents API

    Cloud agents, run on OpenAI's Codex harness

  9. Payflip

    Pay anyone you can name. No IBAN, no wallet address.

  10. Kodro

    Code robots in an offline Python learning simulator

  11. Portfolio Frame

    Frame, annotate, and export screenshots that look designed

  12. siift

    Turn AI noise into better business decisions

Hugging Face(26)

  1. DataFlex-RL: An Evaluation Platform for RLVR Data Policies

    Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

  2. Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

    Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks through five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into trainable reasoning. Our data engine constructs resettable coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed environments. Candidate trajectories are retained only after execution verification and evidence auditing, yielding 164,269 trajectories for long-context supervised fine-tuning. The three checkpoints improve over their starting models by an average of 23.76% on the full CyberGym suite and 10.49% across the pooled CTF suites. As of September 1, 2026, Feyospace-s1 achieves a verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard, while all three checkpoints rank 1st among models at comparable parameter scales. To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability.

  3. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.

  4. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.

  5. SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

    Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods usually use a lightweight selector to score context units, followed by hard Top-K selection that blocks gradients from the language modeling loss. Consequently, these methods commonly distill layer-wise dense attention distributions. Although this encourages the selector to rank context units by dense attention weights in the original model, the ranking is not directly aligned with their impact on predictions under a fixed attention budget (i.e., the number of attended context units per query), potentially wasting the limited budget on less useful units. To address this misalignment, we propose Simple Attention Sparsification (SAS), a gated sparse attention mechanism that optimizes context ranking end-to-end with the language modeling loss. The key idea is to inject the selector's continuous scores into attention logits during training, allowing the loss to update the selector through standard backpropagation. We identify several choices crucial for this simple design to work well in practice: placing the gate inside the attention softmax in log form, using normalized softmax gates to calibrate historical context against the always-retained current block, and preserving continuous selector scores so the model learns relative priorities rather than only hard selections. To support long-sequence training, we implement a memory-efficient Triton kernel that integrates SAS into FlashAttention-style computation. Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.

  6. COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

    Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce COBRA-Skills, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.

  7. Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

    We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.

  8. Atria Dawn: The Dawn of Agentic Superintelligence

    As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.

  9. ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

    In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.

  10. Dream-RSI: Recursive Self-Improvement through Evolving Worlds

    Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce Dream-RSI, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, Dream-RSI secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, Dream-RSI achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.

  11. PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

    We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.

  12. Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

    The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with sequence length. Grouped-query attention (GQA) reduces this cost by sharing key-value heads, but still stores both a key and a value at every step. We introduce Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map. At inference, the map can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path. A small shared decoupled RoPE channel retains positional information through a separately cached positional key. For the configurations studied, this representation reduces persistent cache scalars by approximately 45-47% relative to matched GQA. At the 350M-parameter scale with 30B FineWeb-Edu tokens, the 16-dimensional positional variant reaches 44.18 average accuracy across five tasks, compared with 44.36 for GQA and 43.88 for MLA. These results demonstrate near-GQA benchmark accuracy with a more compact cache representation. To translate this compact representation into faster autoregressive inference, we have developed custom decoding kernels and are currently evaluating their end-to-end inference performance with an open-source release planned soon.

Techmeme(42)

  1. Cornelis, spun off from Intel in 2020 to build networking tech that helps AI chips communicate more effectively, raised $205M led by IAG Capital (Dominic-Madori Davis/TechCrunch)

    Dominic-Madori Davis / TechCrunch : Cornelis, spun off from Intel in 2020 to build networking tech that helps AI chips communicate more effectively, raised $205M led by IAG Capital —  Cornelis, a company creating networking technology to help AI chips communicate more effectively, announced Monday that it has raised $205 million …

  2. Source: defense tech startup Shield AI is in talks to raise new funds at a valuation of at least $20B; Shield raised $2B at a valuation of $12.7B in March (The Information)

    The Information : Source: defense tech startup Shield AI is in talks to raise new funds at a valuation of at least $20B; Shield raised $2B at a valuation of $12.7B in March —  Shield AI, a startup building drones and AI-powered software for the military, is in talks to raise new funds at a valuation of at least $20 billion …

  3. Nuance Labs, which builds low-latency AI avatars that can have face-to-face conversations, raised a $50M Series A led by Lightspeed, with Nvidia participating (Shubhangi Goel/Business Insider)

    Shubhangi Goel / Business Insider : Nuance Labs, which builds low-latency AI avatars that can have face-to-face conversations, raised a $50M Series A led by Lightspeed, with Nvidia participating —  A startup that's trying to make AI models better conversationalists with more emotional intelligence has just raised $50 million.

  4. Sources: Trump met privately with Sam Altman backstage at the GOP midterm convention, where they discussed AI and its growing power, at Altman's request (MS NOW)

    MS NOW : Sources: Trump met privately with Sam Altman backstage at the GOP midterm convention, where they discussed AI and its growing power, at Altman's request —  Altman, Elon Musk and Anthropic chief Dario Amodei all urged an artificial intelligence slowdown over the weekend.

  5. Anthropic debuts Claude for Financial Advisors, with connectors to investment analytics and wealth-management tools from BlackRock, Addepar, Schwab, and others (Harshita Mary Varghese/Reuters)

    Harshita Mary Varghese / Reuters : Anthropic debuts Claude for Financial Advisors, with connectors to investment analytics and wealth-management tools from BlackRock, Addepar, Schwab, and others —  AI lab Anthropic on Monday launched a set of tools for financial advisers, connecting its Claude chatbot to investment analytics …

  6. Sources: OpenAI bought Glass Imaging, which is developing AI-powered smartphone camera tech, in a deal valuing it at $300M+; it was valued at ~$100M last year (Wall Street Journal)

    Wall Street Journal : Sources: OpenAI bought Glass Imaging, which is developing AI-powered smartphone camera tech, in a deal valuing it at $300M+; it was valued at ~$100M last year —  Glass Imaging, valued above $300 million in deal, was founded by former Apple employees  —  OpenAI quietly bought …

  7. At a US House hearing, Treasury Secretary Scott Bessent said AI labs should get no liability exemptions and called for more open-source models built in the US (Matt Bracken/FedScoop)

    Matt Bracken / FedScoop : At a US House hearing, Treasury Secretary Scott Bessent said AI labs should get no liability exemptions and called for more open-source models built in the US —  The secretary told House Financial Services Committee lawmakers that the “best way to guarantee safety” is for AI creators to be held …

  8. OpenRouter users spent more on OpenAI's models than on Anthropic's in the week of September 7, the first time that happened since the week of February 26, 2024 (@openrouter)

    @openrouter : OpenRouter users spent more on OpenAI's models than on Anthropic's in the week of September 7, the first time that happened since the week of February 26, 2024 —  OpenRouter users spent more on OpenAI models than on Anthropic models last week. This hasn't happened for more than 2.5 years

  9. Pulley, which offers cap table management software, says it will cease operations after December 8; it had raised $50M+ from investors, including Founders Fund (Melia Robinson/Business Insider)

    Melia Robinson / Business Insider : Pulley, which offers cap table management software, says it will cease operations after December 8; it had raised $50M+ from investors, including Founders Fund —  Pulley, a software that startups rely on to track their funding, is shutting down after seven years.

  10. At the Future of Life Institute's Pro-Human Assembly, Bernie Sanders, Steve Bannon, and others called for tighter restrictions on AI and denounced tech CEOs (New York Times)

    New York Times : At the Future of Life Institute's Pro-Human Assembly, Bernie Sanders, Steve Bannon, and others called for tighter restrictions on AI and denounced tech CEOs —  At an event in Washington, partisanship took a back seat as elected officials, religious leaders, parents and artists called for reining in artificial intelligence.

  11. During a Salesforce event, Jensen Huang says the AI industry doesn't need any new laws or regulations and market forces will help companies safely innovate (Brody Ford/Bloomberg)

    Brody Ford / Bloomberg : During a Salesforce event, Jensen Huang says the AI industry doesn't need any new laws or regulations and market forces will help companies safely innovate —  Nvidia Corp. Chief Executive Officer Jensen Huang dismissed the need for new artificial intelligence security regulations on Tuesday …

  12. Crypto exchange CoinEx says it is closing, citing a lengthy downturn and rising compliance costs; a report said it moved $3B+ for Iran-linked wallets since 2019 (Dylan Tokar/Wall Street Journal)

    Dylan Tokar / Wall Street Journal : Crypto exchange CoinEx says it is closing, citing a lengthy downturn and rising compliance costs; a report said it moved $3B+ for Iran-linked wallets since 2019 —  CoinEx says it is ceasing operations less than three months after a Wall Street Journal article spotlighted its use in Iran

Solidot(39)

  1. 非洲野犬完成了横跨大陆的 4000 公里之旅

    根据发表在《Ecology》期刊上的一项研究,一群非洲野犬完成了横跨大陆、创纪录的 4000 公里之旅。科学家表示这是有记录以来非洲陆生哺乳动物为寻找配偶而行进的最远距离。三只雄犬行进的直线距离大约为 418 公里,但为了绕过人类活动区域它们迂回走了 4000 公里路。非洲野犬是非洲最稀有的捕食者之一,目前野外仅存约 6000 只。它们生活在高度社会化的家族群中,集体狩猎,四处游荡、寻找新领地以及与其它群体进行繁殖机会而闻名。它们无法在自己出生的家族群内繁衍,因此要么等待可能最终继承该家族群,要么在两三岁时出发寻找配偶。在这次寻找配偶而进行的迁徙中,三只雌性犬因落入人类陷阱而有两只死亡。

  2. 越南关联服务器泄漏了 2.2 亿条旅客信息

    Kinryū Labs 发现了一个因错误配置而能被访问的数据库,该数据库 Advance Passenger Information 记录了过去九年进出越南的几乎所有旅客和机组人员的信息。在接到通知之后该数据库的访问于 2026 年 6 月关闭。Kinryu Labs 是在 6 月 3 日发现了名为 pax-info 的 Elasticsearch 集群,该数据库可使用默认凭证登陆,运营者没有改变默认的用户名和密码,它包含了 29 个索引和约 107 GB 的数据。其中两个主要索引分别存储了 210,318,069 条乘客记录和 10,465,631 条机组人员记录,总计 220,783,700 条记录,时间是从 2017 年 1 月 7 日至 2026 年 4 月 30 日。泄露的信息包括乘客和机组人员的姓名、出生日期、性别、国籍、护照或旅行证件号码、证件有效期及签发国。相关的旅行数据则包括航班号与日期、航空公司、出发地、目的地及中转机场、座位信息、行李编号,以及计划、预计和实际飞行时间。涉及的旅客国籍包括韩国、中国、加拿大、新西兰等。

  3. 中国地震局与苹果公司沟通推进地震预警信息接入 iOS

    中国地震局监测司上周五表示,中国地震台网中心正在与苹果公司沟通,力争加快推进地震预警信息接入 iOS 系统。苹果手机用户目前可通过微信小程序获取该局统一发布的地震预警信息。今年 8  月 24 日,四川宜宾长宁发生 4.7 级地震,但成都高新减灾研究所用自己的系统生成了一个“7.7级”的地震预警,并以“中国地震预警网”的名义,通过荣耀、vivo、魅族手机以及小天才手表等终端向用户推送。中国地震局后来把这种行为定性为“擅自生成”“违规推送”。 成都高新减灾研究所对此提出异议,称“中国地震预警网”是它与中国地震局此前合作建设的,否认是“冒用”。

  4. 养狗有助于降低老人患认知症风险

    日本国立环境研究所等机构从 2016 年起,历时 7 年半对约 1.1 万名老年人开展了调查。他们在学术期刊上发表了研究成果。养狗的老年人因认知症需要接受护理的风险比从未养狗的人群低 48%。研究认为,遛狗带来的身体活动以及社交往来起到了积极作用。曾经养过狗的人患认知症的风险也低于从未养过狗的人群。虽然该差异在统计学上并不显著,但推测养狗时期建立的人际联系等因素可能带来了积极影响。研究还表明,养狗能拉动经济。若养狗人群增加,宠物食品、宠物保险、宠物寄养等相关商品与服务的需求预计随之上涨。

  5. 律师在谋杀案中捏造了证词,他将此归咎于 ChatGPT

    律师在法律文件中使用 AI 工具捏造不存在的信息不是什么大新闻,AI 捏造的通常是不存在的案例,然而本案的特殊之处在于 AI 捏造了证词。律师 Stephen Aaron 在一起谋杀案中代表其客户提起上诉,在递交的法律文件中包含了捏造的警方证词以及虚构的证人。Aaron 声称他将一份由计算机生成的庭审记录及其它案卷材料输入了 ChatGPT,想当然地认为它会生成一份“无懈可击的摘要”。他不清楚 AI 工具会产生“幻觉”——即虚构信息。 法官对此难以置信,反问他没看新闻吗?法官对他处以 5000 美元罚款,将把他移交至律师纪律委员会进行调查。

  6. 日本无意结婚的男女比例都超两成

    日本国立社会保障与人口问题研究所公布了 2025 年出生动向基本调查。18-34 岁未婚人群“终生不打算结婚”的男女受访者比例首次都超过 2 成,其中男性为 24.0%,女性为 21.5%。表示“打算将来结婚”的人群中男性占 75.1%,女性占 77.8%。均首次跌破 8 成。回答结婚有好处的人群男性占 56.3%,女性占 63.4%,均创历史最低水平。夫妻理想中的子女数量比 2021 年上一次调查的平均 2.25 人减少 0.07 人至 2.18 人。计划生育的子女数量为 1.95人,自统计开始以来首次跌破 2 人。减少生育的原因回答“育儿和教育花费太高”的受访者达到 52.9%,比例最高。回答“不想高龄生育”(35.0%)和“无法再承受育儿带来的心理及身体负担”(27.8%)紧随其后。

  7. 夜晚睡眠光照太亮可能会损伤心脏

    研究人员分析了 英国生物样本库(UK Biobank)11,071 名参与者的数据,参与者在一周时间内手腕佩戴了光线和运动传感器。研究开始时参与者均未有心血管疾病。在几年之后他们接受了心脏 MRI 检查。研究人员主要针对两类人群,其一是夜间睡眠时几乎没有任何光;其二是接触至少 3 lux(照度单位)的光,这些光线可能来自透过窗帘射入的街灯,家用电器上的 LED 灯。在考虑个人背景、生活方式、健康状况和环境因素后,研究人员发现,夜间睡眠时的光照水平如果超过3 lux,每增加一点光照都与可测量的、细微的心脏损伤有关。相比在最黑暗房间内睡觉的参与者,光照暴露量最高的参与者左心室体积增大 2.4%、心壁增厚 1.5%,以及心脏收缩能力下降 1.9%。虽然心脏变化微小,但与心血管疾病及中风存在关联。

  8. 出于兴趣阅读有助于促进终身的身心健康

    WHO 的数据显示,全球逾 10 亿人有心理健康障碍,其中焦虑症和抑郁症等病症造成了巨大的个人痛苦和经济损失。全世界约有七分之一 10-19 岁青少年有心理障碍,占该年龄段疾病负担的 15%。抑郁症、焦虑症和行为障碍是导致疾病和残疾的主因,而自杀则是 15-29 岁人群的第三大死因,凸显了为青少年提供心理健康支持的迫切性。人们已经认识到,环境因素会影响大脑健康、认知能力、心理健康及身体健康,而这些因素可通过改变行为加以改善。因此通过改善生活方式,人们不仅能提升大脑健康和认知能力,还能降低患心理健康障碍和躯体疾病的风险。剑桥大学的研究人员指出,出于兴趣阅读以及参加读书会,是一种有助于促进终身身心健康的低成本干预措施。阅读投入与大脑及心理健康的改善、认知表现的提升以及认知衰退风险的降低密切相关。对成人的调查数据显示,阅读与压力减轻、共情能力增强、幸福感提升以及孤独感降低有关。对青少年研究显示,出于兴趣阅读与注意力、记忆力、执行功能及学业成绩相关,同时也与较少的心理健康问题相关。

  9. 英国殖民之前的澳大利亚原居民人口约 222 万

    在英国舰队于 1788 年登陆澳大利亚前,这块大陆生活了多少原居民?在英国殖民澳大利亚 140 多年后的 1930 年代,人口学家 Alfred Radcliffe-Brown 首次对原居民的人口总数进行了估计。他估计澳洲原居民的人口在 25 万到 30 万之间,他强调这是一个最低估计值。现在研究人员使用了五种不同的方法重新进行了估计,得出的中位数是——殖民前澳大利亚的原住民约有 222 万。研究人员称,原住民人口至少 100 万以上,有可能在 200 万至 300 万之间,甚至可能超过 500 万。殖民后原居民的人口锐减则是疾病以及暴力导致的。到 1861 年,原住民人口仅剩约 17.7-19.3 万人。时至今日原居民人口仍然未达到殖民前的水平。

  10. F-Droid 上的应用有多少是在 AI 帮助下编写的?

    今天有无数开发者在 LLM 帮助下编写程序,其中包括了开源开发者。那么 Android FOSS 应用商店 F-Droid 中 AI 辅助开发应用的比例有多高?一位 FOSS 维护者对 9 月 12 日 F-Droid 推送更新的 102 款应用及其代码库进行了分析,发现其中 74 款应用(72.5%)主要是 AI 编写的,10 款应用难以明确归类(9.8%), 18 款应用几乎没有 AI 参与的迹象(17.6%)。有 4 个托管在 Codeberg 上的应用主要是 AI 编写的,而 Codeberg 最近宣布了 AI 政策,禁止了此类 AI 应用,但要清除此类应用显然需要更多时间。

  11. 廉价太阳能改变世界能源格局

    巴基斯坦水泥公司 Bestway Cement 正在扩建其太阳能发电设施,计划年底前在现有 26MW 装机容量的基础上增加 6.34MW 装机容量。太阳能满足了该公司逾四分之一的电力需求。受益于中国制造的廉价太阳能组件,Bestway 及其竞争对手加入了全球数百万企业和家庭的行列,在屋顶、庭院、花园等空地上安装太阳能电池板。截至 2025 年底,全球太阳能装机容量已接近 1.2TW。由廉价中国光伏板推动的太阳能革命——以及个人发电模式的兴起——正在改变发展中国家乃至工业化国家的能源格局。在较贫穷国家,数以百万计的人们如今获得了更可靠的电力供应,而这是通过他们自身努力实现的,而非依赖于大规模的基础设施建设。标普全球太阳能与储能研究经理 Josefin Berg 表示,太阳能的增长正在彻底改变电力系统,使其从集中式结构转变为一种任何人都能发电的模式。本世纪初,太阳能电池板的成本约为每瓦发电容量 5-6 美元。如今已降至每瓦约 12 美分。Ember 预计非洲今年将新增约 17 GW 的太阳能装机容量。菲律宾电力分销商 Meralco 表示,今年上半年屋顶太阳能发电量达到了 372 GWh,该国的家用太阳能电池板只需三年多时间即可收回成本。南非国有电力公司 Eskom 估计,截至今年 3 月的一年内,其售电量减少了 11.7 TWh,约 7% 的降幅归因于屋顶太阳能电池板和电池系统的普及。太阳能在阴雨天气发电量会大幅下降,未来的电网系统将需要考虑这一情况。

  12. 一款在浏览器里运行、部署在自己服务器上的 SQL 客户端

    Yusuf Gundogdu 写道:LibreDB Studio 是一个 MIT 协议的 SQL 客户端,不装在本地而是跑在服务器上,浏览器打开就能用,一条 docker run 就起来。16 个驱动覆盖 42 种数据库,PostgreSQL、MySQL、MongoDB、Redis、ClickHouse 这些都在内。9 月 8 日发布了 0.15.0 版本。我觉得值得一提的是他们把 AI 那部分做了实测:28 个模型跑同一套六项数据库任务,27 个通过 Ollama 完全在本地运行,最快的 qwen2.5:7b 只有 4.7 GB,一次完整运行中位数 6 秒,最小的 2.5 GB。数据逐个模型公开,包括没通过的和卡在哪一步。另外只读不是靠解析 SQL 挡的,是数据库自己挡的:PostgreSQL 上开只读事务,SQLite 上每条语句前重设 query_only。