ISSUE 0997
WED, SEP 23, 2026
The directory AI cites when builders ask what to use
TODAY · WED, SEP 23, 2026

Ship your AI.
Get discovered.

List your product on OrangeBot and reach builders and users actively looking for the right AI tools.

Daily launches · 2,000+ Claude Code skills · 115+ free tools · AI news from 10 sources — rebuilt every morning.

FOUNDERSBuilding an AI tool? Assistants cite lists like this one, not your homepage.Get listed →
Why founders list here

More than a launch. Long-term discovery.

Get in front of builders

Show up when builders are actively looking for tools like yours.

Context that converts

Tell builders what your product does, who it is for, and why it matters.

In the right ecosystem

Your product sits alongside the skills, tools and sources builders already trust.

Built for AI discovery

Structured so both people and AI assistants can understand and recommend it.

Stay discoverable

Keep getting found long after launch day — the page does not expire.

Learn more about getting listed →
01

Latest Launches

CURATED BY ORANGEBOT
01

AI DIGEST

UPDATED DAILY · EDITOR'S PICK
01.00
AI DIGEST

AI新闻摘要

September 23, 2026

Here is a summary of today's main news events, based on the information provided.

U.S. Markets Tumble as Treasury Yields Hit 16-Year High

U.S. stocks, particularly the tech-heavy Nasdaq, fell sharply today as the 10-year Treasury yield surged above 5.1%, a level not seen since 2007. The spike was fueled by rising oil prices and growing concerns about inflation, increasing investor bets that the Federal Reserve will continue to raise interest rates.

U.S. and China Prepare for High-Stakes Presidential Summit

Leaders from the U.S. and China are preparing for a significant summit in Washington. The meeting takes on added urgency as it precedes the expiration of a U.S.-China trade truce in less than two months, with major implications for global trade and diplomacy.

Swiss Parliament Eases Capital Rules for Major Bank

The upper house of Switzerland's parliament voted to soften proposed new capital requirements for the nation's largest banking group. The move rejected a stricter government proposal, indicating a political battle over how to regulate the newly enlarged bank following its recent major acquisition.

Major U.S. Passenger Railroad Prepares for Bankruptcy

A high-speed passenger railroad backed by Fortress Investment Group is planning to file for bankruptcy. The company has been unable to overcome challenges from its $5.5 billion debt and slower-than-expected ridership growth.

McDonald's Outlines Plan to Boost Sales by Targeting Chicken Market

The fast-food giant announced a new spending strategy aimed at capturing a larger share of the booming chicken market. The plan is a direct response to sluggish U.S. sales, as McDonald's seeks to reignite growth by competing more effectively against chicken-focused rivals.

02

ON THE WIRE

6 SOURCES
02

HACKER NEWS

02.00
HACKER NEWS

Hacker News - September 23, 2026

Hacker News Feed: Highlighting key posts and discussions.

I don't want the details

(michaelheap.com)

328189
The darker side of being a doctor

(drericlevi.pages.dev)

261298
I am done with this shit

(www.reddit.com)

230184
Jev in 25 Lines of Python

(www.nobodywho.ai)

608192
Transit rewards

(waymo.com)

246323
SAML: A fractal of bad design

(blog.trailofbits.com)

339177
GPT-6 Sol and Luna

(openai.com)

1729821
Claude Opus 5.5

(www.anthropic.com)

17661075
Claude Opus 5.5

(www.anthropic.com)

2732
03

HUGGINGFACE

03.00
HUGGINGFACE

HuggingFace 新闻 - September 23, 2026

HuggingFace Feed:最新的 AI 模型、数据集和社区动态。

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks. Each question presents a decision fork, a point in a trajectory where multiple directions are available and one of them leads to a better outcome, and the evaluated model chooses among these directions without seeing what happens after the fork. We mine these forks automatically from parallel attempts at the same task and from detours inside a single trajectory, without needing human annotation. We evaluate frontier models on Taste-Bench and find that the best model answers only 59.7% of the questions correctly. We further find that forks whose deciding evidence appears later in the trajectory are much harder for every model, and that a larger reasoning budget does not improve the accuracy. Finally, we show that taste can be trained. We distill the judgment of a teacher that has seen the outcome into a student model, and the student makes better decisions on unseen tasks and improves end-to-end success on held-out SWE-bench Pro tasks.

111
RULER: Instance-aware Rubric Rewards for SVG Generation

Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.

59
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by 12.7% and 23.1% on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.

38
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.

35
Bellman Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD as a trajectory-level objective. The reformulation avoids estimating state values at intermediate states. We prove that it has the same unique optimal solution as the original PMD objective. We derive the practical BPO loss by approximating this objective. Its mismatch-correction weight is a smoothed ratio of complementary token probabilities. Experiments on mathematical reasoning benchmarks demonstrate the effectiveness of BPO.

24
Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models

Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one such approach, but evaluating wider circuits inside a large model can be computationally demanding. Here we introduce HyperQ, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model. A quantum residual branch is a module in each transformer block that reads a token's hidden state, emits the coordinates of that token's circuit, executes it, and adds the measured values back through a residual connection. The backbone remains frozen, and only the added branches are trained. Within each branch, a lightweight circuit hypernetwork emits token-specific rotation angles, coupling strengths, and measurement axes in a shared sparse circuit structure. The required expectation values have an exact classical expression whose evaluation cost grows linearly with the qubit count, enabling circuits from 16 to 64 qubits to be trained within a 1.1-billion-parameter backbone. Across downstream benchmarks, increasing circuit width raises the average score from 47.65 to 54.30. At 64 qubits, HyperQ exceeds the backbone and its low-rank-adapted counterpart by 4.71 and 3.67 points, respectively. HyperQ is fine-tuned on 20,000 prompt-response pairs, compared with 200,000 for the classical baselines. These findings support token-conditioned circuit emission as a tractable architectural approach to quantum-augmented language modelling.

23
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training

Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. We propose StableVQ, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently. Concretely, (1) Dynamic STE corrects the instability in the Encoder's learning objective, enabling it to robustly optimize the reconstruction space under discrete regularization even when codebook utilization is low. (2) Region VQ Loss reconceives the Codebook's learning objective so that it can independently guarantee full tracking of the encoder output distribution, without relying on encoder oscillations to drive activation. (3) Decoupled Schedule recognizes that the distinct responsibilities of the Encoder--Decoder and the Codebook demand distinct optimization dynamics, and assigns each an independent learning rate schedule to ensure robust system-level behavior. Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters. Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.

20
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

In this report, we introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make three key advances: (1) native omni-modal initialization: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) data-centric omni-modal training: we construct a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data. To improve data efficiency, we introduce homogeneous-source sampling to form task-consistent batches with informative in-batch negatives; and (3) embedding-specific training and inference optimization: we use focal loss to emphasize hard examples and similarity-based Embedding Distillation to transfer fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition enables compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show that the Ovis-Embedding family achieves state-of-the-art performance on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB, demonstrating its effectiveness across text, image, video, and audio modalities. These results highlight the potential of unified omni-modal training to overcome modality fragmentation and advance universal embedding models for any-to-any retrieval.

19
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.

17
From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health

The rising global prevalence of mental health conditions, together with longstanding barriers in traditional healthcare, such as limited resources, high cost, stigma, and privacy concerns, has created an urgent need for accessible and scalable support. Large Language Models (LLMs) have emerged as a transformative technology with strong potential to democratize mental health support through advanced natural language understanding and generation. However, the rapidly expanding, fragmented body of work in this area lacks a coherent evolutionary narrative, making it difficult to contextualize current progress and identify future directions. This survey addresses this gap by organizing and analyzing the literature around a central thesis: the role of LLMs in mental health is evolving through three distinct, increasingly sophisticated phases. We trace this trajectory from Phase I, in which LLMs act primarily as passive Information Tools and Pattern Recognizers for assessment; through Phase II, where they function as Empathetic Conversationalists for in-the-moment, stateless interactions; to the current frontier, Phase III, which seeks Longitudinal, Personalized Companions implemented as stateful cognitive agents. To support this framework, we systematically review core technologies, agent architectures (Profile, Memory, Reasoning, and Planning), and the critical infrastructure of datasets and benchmarks, highlighting how their evolution underpins this developmental path. Viewing the field through this developmental lens, we provide a comprehensive synthesis of existing work, an insightful narrative of its trajectory, and a clear roadmap for future innovation in responsible, effective, and human-centered AI for mental healthcare. A curated collection of the resources reviewed in this survey is available at our project repository: https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.

15
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce Flash-dLLM, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger batch size. Extensive experiments on mathematical reasoning and code-generation benchmarks demonstrate that Flash-dLLM consistently outperforms existing state-of-the-art dLLM acceleration methods in both inference speed and memory efficiency. In particular, it achieves 5.1times and 11.0times speedups over prior strongest baseline Elastic-Cache on GSM8K and HumanEval, respectively.

13
Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.

11
Agensh: Scaling Organizational Intelligence to 1,024 Agents

A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to allocate tasks and coordinate workers. To address this limitation, we introduce Agensh, a scalable self-organized multi-agent harness without a central orchestrator: concurrent workers execute a multi-agent cooperation loop, continuously gathering context, claiming and self-assigning sub-tasks, taking action and sharing findings, verifying results, and merging progress in an asynchronous manner. The loop is supported by the agentic organization infrastructure comprising three components: a shared workspace holds proposed, ongoing, and completed work; a message interface lets workers communicate; and shared context retains reusable findings and work intentions. To test the scalability of Agensh, we evaluate it on the five hardest ProgramBench tasks with GPT-5.6-sol (high). Scaling from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78%, an approximately 49% relative improvement. Larger organizations reach comparable test-pass rates earlier. On pandoc, scaling from 1 to 1,024 agents raises the final test-pass rate from 33.89% to 55.06%. Worker trajectories further show that different forms of self-organized cooperation gradually emerges and standardizes as the organization grows. These results reveal the number of agents as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.

9
Recursive self-improvement of AI research agents

AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves. When an AI research agent's own code is the object of optimization, each accepted rewrite becomes the agent that the next round edits. We refer to this loop as recursive self-improvement. Its significance lies in a long-standing trend, in which increased cumulative spending on R&D yields diminishing returns. Sustained self-improvement offers a way to counter this trend. We present AIDE^2, a system that implements this loop for a frontier AI research agent. It proposes changes to its own code, benchmarks modified versions of itself on a suite of AI R&D tasks, and keeps the changes that perform best on hidden evaluations. In an autonomous 8-day run, AIDE^2 discovered seven successive improvements, ranging from a new search policy to memory mechanisms that compress and manage the agent's growing context. These gains generalize to four held-out benchmarks spanning machine learning engineering, heuristic algorithm engineering, and physics-based weather forecasting, the last of which is out of distribution from the selection tasks. On all four, the strongest discovered agent matches or exceeds a human-engineered production research agent that ranks among the strongest on FML-Bench. On a separate held-out task family, the discovered agents also exhibit reduced reward hacking, a property the loop never explicitly optimized for: the rate falls from 55% to 32% during the run, 7 percentage points below the human-engineered agent. Together, these results show that an AI research agent can improve its own research efficiency through recursive self-improvement, and that these gains transfer to tasks and domains the loop never encountered.

7
Emergent Collusion in Long-Horizon LLM Agent Interaction

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.

6
RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.

4
Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament

Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consistently influenced blame attribution - an effect we term political contrasting - with opposition parties blaming substantially more than governing parties. This effect was moderated by ideology: The blame-dampening effect of governing was less pronounced among right-wing parties, and ideological extremity amplified blame more strongly on the right. In recent years, the interaction between political wing and ideological extremity intensified, suggesting an ideological hardening of the blame rhetoric concentrated on the right of the political spectrum. Taken together, these patterns suggest that the perceived rise in harsh political language reflects not merely a general rhetorical drift, but an ideologically asymmetric hardening of political discourse. A sensitivity analysis showed that the conclusions were robust to varying classification thresholds.

4
LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay. Translated attention KV alone leaves a large gap; adding the Gated DeltaNet (GDN) persistent-state package lowers teacher-forced negative log-likelihood (NLL), the average next-token log-loss, by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]), improving all 64 PG19 documents. Direct recurrent and convolution reuse outperforms the tested learned GDN maps, consistent with partial functional compatibility of persistent-state coordinates. A fresh component factorial selects translated KV with direct recurrent and convolution state. An additional 434,176-parameter correction improves that base on 64 fresh web documents: continuation loss is 0.076 nats/token above native 9B (excess NLL), Jensen-Shannon (JS) divergence is 0.022, and native context recovery (NCR) is 0.918. Corrected 9B significantly beats continued 4B inference while processing zero historical prefix tokens. Evidence covers one direction, one geometry-matched Base-model pair, and 4K teacher-forced continuation; the near-native gate failed, the 16K branch was not run, and free-generation equivalence and a general state interface remain unproven.

3
ImIR: Image-Instruction Tuning for All-in-One Image Restoration

Degradations vary widely across images, so a practical restoration system has to handle many degradation types with one model. A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt. We replace that prompt with an instruction derived from the degraded image itself. The image reaches the editor through two paths: its structure comes from the model's VAE, and its semantic instruction comes from a lightweight token mapper that shifts the degraded image's vision-language embedding toward the embedding a clean image would produce. Because the instruction is a continuous vector, scaling it yields a family of valid restorations for tasks whose target is not unique, such as low-light enhancement. We adapt one Qwen-Image-Edit model to six tasks with a single adapter trained in about three hours on one GPU. The image instruction outperforms text conditioning under a matched comparison, and it supports task agnostic restoration without a degradation label, which the text variant does not.

3
Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion regressor. Conversely, a joint part-and-handle predictor supplies additional handle candidates, whose motion classes are refined using containing parts. Each information transfer is applied once, without iterative feedback. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes. Additional handle candidates raise handle AP from 24.63% to 29.65%; part-based class correction adds 0.98 points, and full context reaches 30.99%. Repeated training, learned-decoder controls and paired visualizations establish the benefits and limitations of combining geometric and semantic evidence for interaction understanding.

3
Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a photograph or recording of a dog) and propositional signals (e.g., the declarative claim "this is a dog"), such that any measured modality bias is inherently confounded with evidence-form bias, precluding clean attribution to either source. To address this, we introduce Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, where vision and audio each take perceptual or propositional form. Evaluating five OLLMs, we find robust visual bias across most models and evidence-type conditions. Crucially, we reveal a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio. Further analyses via layer-wise linear probing and contrastive decoding reveal that modality bias is already linearly decodable from early representation layers and can only be partially mitigated, calling for mitigation strategies beyond surface-level interventions.

2
Embedding Physics Priors in Robot Learning: A Survey

The rapid progress of artificial intelligence is reshaping robotics and accelerating the adoption of learning-based approaches. While purely data-driven methods have achieved remarkable success in computer vision and natural language processing, robotics remains constrained by limited data, complex real-world interactions, and the need for reliable operation. These challenges have motivated the exploration of physics-embedded robot learning, which embeds physics priors into learning algorithms. By encoding the underlying physical laws and constraints, physics priors can complement limited data with robotics-specific inductive biases, potentially improving generalization, interpretability, and sample efficiency. However, the literature on physics-embedded robot learning remains fragmented across terminology, methodologies, and application domains, making it difficult to assess this growing body of work. This survey reviews physics-embedded robot learning across a broad range of physics priors, robotics applications, and machine learning models, from single-layer perceptrons to generative foundation models. We adopt a unified taxonomy that classifies existing approaches according to their physics embedding: physics-guided inputs, data, and representations; physics-encoded model architectures; and physics-informed training loss functions. Building on this taxonomy, we review methods for robot dynamics learning, trajectory planning, prediction, control, and estimation, together with the corresponding open-source software ecosystem. We identify key open challenges, and outline promising future research directions. Overall, we argue that physics priors provide a particularly relevant robotics-specific inductive bias, complementing rather than replacing data-driven learning, and paving the way toward more generalizable, data-efficient, and trustworthy robotic systems.

1
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched, iso-episode-budget protocol (250 meta-training episodes, 5 canonical seeds, 600 evaluation episodes per seed), our architecture achieves 5-shot accuracy gains, consistent across all five seeds, over Prototypical Networks, Relation Networks, and MAML on both CIFAR-FS and MiniImageNet, while using 27-53% fewer parameters than any baseline. It also converges in fewer training episodes, generalizes better to an unseen fine-grained domain (CUB-200-2011 birds, zero retraining), and is more robust to 50% occlusion and 25% spatial translation than all three baselines. A series of falsification ablations - zeroing relational tokens at inference and retraining without them entirely - shows that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is. We report this honestly, together with a capacity sweep showing a genuine accuracy plateau near 22-35k parameters, and release full seed-level results and checkpoint hashes for reproducibility.

1
05

PRODUCT HUNT

05.00
PRODUCT HUNT

Product Hunt - September 23, 2026

Product Hunt Daily Feed: Featuring noteworthy tech launches.

Naise AI icon
Naise AI

Autonomous marketing agents that actually execute

0
Koreshield icon
Koreshield

Security and evidence for AI support agents

0
Jev State icon
Jev State

Turn AI conversations into tests and runnable code

0
Pactto icon
Pactto

The room where creative teams align and AI takes action

0
GBrain icon
GBrain

Garry Tan's AI memory, tools, & skills for any harness

0
Solid icon
Solid

Agents with their own computers, accounts, and budgets.

0
Linguo Translate icon
Linguo Translate

Offline & AI Translation for macOS

0
Speechka icon
Speechka

Real-time voice translation that sounds like you

0
RankControl icon
RankControl

Get Cited by ChatGPT & Ranked on Google

0
Claude Opus 5.5 icon
Claude Opus 5.5

Anthropic's first model in their new Claude 5.5 family

0
Dub Program Marketplace icon
Dub Program Marketplace

Browse and apply to the best SaaS affiliate programs

0
Alexandria by Firecrawl icon
Alexandria by Firecrawl

The knowledge library for superintelligence

0
ToneBird icon
ToneBird

AI reply assistant that remembers your relationships

0
AgentScore icon
AgentScore

Daily score to see if your agent gets better

0
Lightmeter icon
Lightmeter

Film Camera designed for everyday moments

0
CodeSpotlight icon
CodeSpotlight

Make your code impossible to miss.

0
2BA.AI icon
2BA.AI

Stop waiting for tokens, and start shipping

0
ReallyFree icon
ReallyFree

See what's really 'free' before you click

0
Clueprint icon
Clueprint

See what your agents left running on your Mac

0
OneStream Live 2.0 icon
OneStream Live 2.0

Create, schedule, multistream & sell live on 45+ platforms

0
Pulsetic RUM icon
Pulsetic RUM

Uptime monitoring that measures your real visitors

0
Prowler Cloud icon
Prowler Cloud

Where your AI agent becomes a cloud security defender

0
Xem icon
Xem

Open-source email marketing with managed SMTP

0
Brev icon
Brev

The AI coworker that drives action on your goals

0
QuietGlass icon
QuietGlass

Screen privacy for your Mac, with or without AirPods

0
Edyt icon
Edyt

AI for text you can't even select. Any app on Mac & Windows

0
Fulvid icon
Fulvid

A standalone desktop editor for Markdown and MDX

0
gg-friggin-ez icon
gg-friggin-ez

Fast & free profanity and toxicity screening via Jev & Laya

0
Hola AI icon
Hola AI

An AI voicemail assistant that answers calls when you can’t.

0
Grok 4.7 icon
Grok 4.7

SpaceXAI's most powerful model for coding and knowledge work

0
Plane Agents icon
Plane Agents

Assign work to AI agents, like any teammate

0
Freebuff Ads icon
Freebuff Ads

Advertise to 500k developers in our coding agent

0
PixelCrew icon
PixelCrew

Production-ready design from a crew of AI agents

0
thestory.run icon
thestory.run

A writing coach for corporate influencers.

0
Anomalo icon
Anomalo

Your data is always talking. Don't miss what it's saying.

0
SereneDB icon
SereneDB

Ultra-Fast Search & Analytics Database, Agentic AI ready

0
WZRD icon
WZRD

AI-native documents, slides, forms and sheets that talk back

0
Clueso MCP icon
Clueso MCP

Create and edit videos by chatting

0
Keet icon
Keet

Video Courses on Anything

0
Reeno icon
Reeno

AI calls you randomly to practice speaking a new language

0
PewCB icon
PewCB

Vibe-routing is coming. Desktop PCB Factory is here.

0
Shootsolo 2.0 icon
Shootsolo 2.0

Voice-controlled camera + teleprompter for solo creators

0
Walkie icon
Walkie

Dictation + meetings + read aloud in one app on-device

0
Contextberg icon
Contextberg

Local AI agent memory served via MCP

0
Jev Wrapped icon
Jev Wrapped

See how much of a Telegram channel is ads and clickbait

0
Robot Voice Bridge icon
Robot Voice Bridge

Run ElevenLabs like a voiceover session, right on your Mac

0
VideoFlow Studio icon
VideoFlow Studio

Type your startup URL and get a premium video trailer

0
ResumeContext icon
ResumeContext

Shared memory for coding agents.

0
FeedsBar icon
FeedsBar

A quiet, continuous news/media ticker for your Mac desktop

0
WeWeb MCP icon
WeWeb MCP

Your AI agent builds the app. You stay in control.

0
06

TECHMEME

06.00
TECHMEME

Techmeme - September 23, 2026

Techmeme Digest: Major tech headlines and industry conversations.

Colorado-based Enveda, which uses AI to discover new drugs in the natural world, raised a $311M Series E at a $2B valuation, double its valuation a year ago (Marina Temkin/TechCrunch)
Source: TechmemePublished: Sep 23, 2026

Marina Temkin / TechCrunch : Colorado-based Enveda, which uses AI to discover new drugs in the natural world, raised a $311M Series E at a $2B valuation, double its valuation a year ago —  Enveda, a biotech startup that uses AI to discover new drugs in the natural world, has raised a $311 million Series E at a $2 billion valuation.

Sam Altman and Dario Amodei tell the UN Security Council that governments and industry should coordinate on AI safety standards and critical decisions (Bloomberg)
Source: TechmemePublished: Sep 23, 2026

Bloomberg : Sam Altman and Dario Amodei tell the UN Security Council that governments and industry should coordinate on AI safety standards and critical decisions —  OpenAI Chief Executive Officer Sam Altman and Anthropic PBC CEO Dario Amodei urged world leaders to work together on artificial intelligence …

Sources: Flock Safety is weighing a potential sale and has held preliminary talks with outside advisers (Rohan Goswami/Semafor)
Source: TechmemePublished: Sep 23, 2026

Rohan Goswami / Semafor : Sources: Flock Safety is weighing a potential sale and has held preliminary talks with outside advisers —  THE SCOOP  —  Flock Safety, which is facing scrutiny over the proliferation and alleged abuse of its cameras, is weighing whether to sell itself, according to people familiar with the matter.

Australian PM Anthony Albanese says an OpenAI agent gained unauthorized access to a public-facing Medicare portal in June, accessing public and non-public files (The Age)
Source: TechmemePublished: Sep 23, 2026

The Age : Australian PM Anthony Albanese says an OpenAI agent gained unauthorized access to a public-facing Medicare portal in June, accessing public and non-public files —  Prime Minister Anthony Albanese has revealed that an artificial intelligence agent developed by OpenAI infiltrated …

Source: 1789 Capital, where Donald Trump Jr. is a partner, is in talks to raise $3B for its second growth fund and has already raised $2B of that target amount (Rebecca Torrence/Bloomberg)
Source: TechmemePublished: Sep 23, 2026

Rebecca Torrence / Bloomberg : Source: 1789 Capital, where Donald Trump Jr. is a partner, is in talks to raise $3B for its second growth fund and has already raised $2B of that target amount —  Investment firm 1789 Capital, where Donald Trump Jr. is a partner, is in talks to raise $3 billion for its second growth fund …

A sample of FBI data allegedly stolen by ShinyHunters includes granular detail on officials' job assignments on China, Russia, cartels, cyber, and other issues (Reuters)
Source: TechmemePublished: Sep 23, 2026

Reuters : A sample of FBI data allegedly stolen by ShinyHunters includes granular detail on officials' job assignments on China, Russia, cartels, cyber, and other issues —  FBI data allegedly stolen by the hacking group ShinyHunters carries granular detail about scores of bureau officials' job assignments …

OpenAI says ChatGPT Voice can now be powered by GPT-6 Astra, Sol, and Luna, use plugins like email and calendar, and be used in ChatGPT Work on web and mobile (@openai)
Source: TechmemePublished: Sep 23, 2026

@openai : OpenAI says ChatGPT Voice can now be powered by GPT-6 Astra, Sol, and Luna, use plugins like email and calendar, and be used in ChatGPT Work on web and mobile —  We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by ...

Anthropic says Claude autonomously discovered a new enzyme system in the DNA of bacteriophages, somewhat similar to CRISPR, the first result from its new biolab (Robert Hart/The Verge)
Source: TechmemePublished: Sep 23, 2026

Robert Hart / The Verge : Anthropic says Claude autonomously discovered a new enzyme system in the DNA of bacteriophages, somewhat similar to CRISPR, the first result from its new biolab —  It took nearly 1,000 Claude agents 21 hours and 210 million tokens to uncover the ‘novel enzyme system.’

Pilgrim, whose device combines air sampling and genomic sequencing to detect biological threats, raised a $25M seed led by Buckley at a $150M valuation (Kate Clark/Wall Street Journal)
Source: TechmemePublished: Sep 23, 2026

Kate Clark / Wall Street Journal : Pilgrim, whose device combines air sampling and genomic sequencing to detect biological threats, raised a $25M seed led by Buckley at a $150M valuation —  Pilgrim has raised $25 million and has developed a 50-pound device that detects biological threats  —  Two Anthropic leaders are investing …

Disney raises Disney+ and Hulu prices, the fourth price hike in as many years; the ad-free tiers will each cost $21.49/month, an increase of $2.50 (Rick Porter/The Hollywood Reporter)
Source: TechmemePublished: Sep 23, 2026

Rick Porter / The Hollywood Reporter : Disney raises Disney+ and Hulu prices, the fourth price hike in as many years; the ad-free tiers will each cost $21.49/month, an increase of $2.50 —  Subscribing to a bundle is now almost exactly half the price of individual plans for each streamer.

Google releases Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its "most expressive audio generation models yet", with support for more than 100 languages (Google)
Source: TechmemePublished: Sep 23, 2026

Google : Google releases Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its “most expressive audio generation models yet”, with support for more than 100 languages —  Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet.

Q&A with Jensen Huang on AI creating more jobs than it destroys, pushing back against AI doomerism, Chinese open models, the Hugging Face acquisition, and more (Ezra Klein/New York Times)
Source: TechmemePublished: Sep 23, 2026

Ezra Klein / New York Times : Q&A with Jensen Huang on AI creating more jobs than it destroys, pushing back against AI doomerism, Chinese open models, the Hugging Face acquisition, and more —  This is an edited transcript of “The Ezra Klein Show.”  You can listen to the episode wherever you get your podcasts.

At its Accelerate conference, Amazon says it is opening its seller tools to third-party AI agents, starting with Claude, in beta for US merchants (Todd Bishop/GeekWire)
Source: TechmemePublished: Sep 23, 2026

Todd Bishop / GeekWire : At its Accelerate conference, Amazon says it is opening its seller tools to third-party AI agents, starting with Claude, in beta for US merchants —  Amazon announced new AI tools for its independent sellers Wednesday, led by a plugin that lets them run their Amazon businesses from Anthropic's Claude or Amazon's own Quick assistant.

Coachella promoter Goldenvoice renews its livestreaming partnership with YouTube through 2030, choosing it over competing bids from Netflix and Amazon (Lucas Shaw/Bloomberg)
Source: TechmemePublished: Sep 23, 2026

Lucas Shaw / Bloomberg : Coachella promoter Goldenvoice renews its livestreaming partnership with YouTube through 2030, choosing it over competing bids from Netflix and Amazon —  New deal includes mini artist documentaries and more programming leading up to North America's largest music festival.

At its Made on YouTube event, YouTube unveils Custom Feeds, an LLM-powered feature that lets users generate and save tailored homepage video feeds, for US users (Reece Rogers/Wired)
Source: TechmemePublished: Sep 23, 2026

Reece Rogers / Wired : At its Made on YouTube event, YouTube unveils Custom Feeds, an LLM-powered feature that lets users generate and save tailored homepage video feeds, for US users —  In YouTube's latest update, viewers can tweak what the algorithm shows on their front page and curate it to a specific vibe.

07

STARTUP ARCHIVE

07.00
STARTUP ARCHIVE

Startup News - September 23, 2026

Startup News Roundup: Aggregating key funding and launch updates.

Marc Andreessen on the 5 personality traits of an innovator
Source: StartupPublished: Mar 31, 2026

“When you’re talking about real innovators—people who actually do really creative, breakthrough work—I think you’re talking about a couple things:”

Steve Jobs explains the importance of both thinking and doing
Source: StartupPublished: Mar 30, 2026

“The doers are the major thinkers. The people who really create the things that change this industry are both the thinker-doer in one person.”

Tobi Lutke explains what the VCs who passed on Shopify got wrong
Source: StartupPublished: Mar 27, 2026

“What a lot of free-market thinkers don’t understand is that between the demand and eventual supply lies friction."

Sam Altman explains how he decides to invest in a startup after 10 minutes
Source: StartupPublished: Mar 26, 2026

"Does this person have the potential to be the next Mark Zuckerberg?… [You don’t get to] 100% accuracy, obviously, but it’s good enough that our business model works.”

Jony Ive recounts the time Steve Jobs called him vain
Source: StartupPublished: Mar 25, 2026

In the clip below, Jony Ive recounts the time he asked Steve Jobs to be less harsh in his critique of a piece of work.

Jeff Bezos’s two pieces of advice for aspiring entrepreneurs
Source: StartupPublished: Mar 24, 2026

“The advice that I would give entrepreneurs is don't chase the hot new thing. It's so hard to catch something that everybody already knows is hot."

Elad Gil: “Things that work tend to work pretty fast”
Source: StartupPublished: Mar 23, 2026

“I do think there’s a bit of a myth in Silicon Valley that you should keep grinding no matter what and it’s just about perseverance, and I think that’s really bad advice."

Paul Graham on why starting with a “small, intense fire" is the key to startup growth
Source: StartupPublished: Mar 20, 2026

"You have to know who those first users are and how you're going to get them."

Keith Rabois on how to identify great talent
Source: StartupPublished: Mar 19, 2026

“What you want to do with every single employee every single day is expand the scope of their responsibilities until it breaks… and that’s the role they should stay in.”

Wealthfront CEO on why advertising spend makes it harder to find product/market fit
Source: StartupPublished: Mar 18, 2026

“The way that you know you have product/market fit is if you have exponential organic growth."

Eric Schmidt on why most companies get strategy wrong
Source: StartupPublished: Mar 17, 2026

“Work very, very hard to figure out what the world’s going to look like in five years. What will people be doing? What will your customers want? Where will costs be?"

Mark Zuckerberg: “You can’t 80/20 everything”
Source: StartupPublished: Mar 16, 2026

"There’s the famous 80/20 rule where you get 80% of the benefit by doing 20% of the work, but you can’t just 80/20 everything. There have to be certain things that you are just the best at."

Marc Andreessen on Mark Zuckerberg’s founder “superpower”
Source: StartupPublished: Mar 13, 2026

“A great superpower that Mark Zuckerberg has that is probably not well-understood enough is he does not get emotionally upset in stressful situations"

Sam Altman explains how to come up with a great startup idea
Source: StartupPublished: Mar 12, 2026

"If you start a startup without a good idea… you’ll be under pressure to make something up and it won’t work that well."

Jeff Bezos on the problems with proxies and managing to metrics
Source: StartupPublished: Mar 11, 2026

“One of the things that happens in business is that you develop certain things that you’re managing to—a typical case would be a metric. And that metric isn’t the real underlying thing.”

Airbnb founder Brian Chesky on how to design an amazing user experience
Source: StartupPublished: Mar 10, 2026

“If you can design something really amazing using the hand-crafted part of your brain, then you can reverse-engineer how to industrialize this millions of times over."

Spencer Rascoff: "I will never invest in a consumer startup with paid marketing”
Source: StartupPublished: Mar 9, 2026

"If you’re actually trying to grow a product, the best levers for doing that are often within the product itself.”

Patrick Collison explains why it sometimes make sense to quit
Source: StartupPublished: Mar 6, 2026

“One thing I’ve learned myself the hard way, is that it is easier to tear down a company and restart it in Silicon Valley, than it is to constantly try to pivot or keep something alive."

Jeff Bezos recounts the time he called Amazon’s customer service number mid-meeting to prove a metric was wrong
Source: StartupPublished: Mar 5, 2026

“I have a saying, which is when the data and the anecdotes disagree, the anecdotes are usually right"

Ben Horowitz: “Nobody was born a great manager. It’s a very unnatural job.”
Source: StartupPublished: Mar 4, 2026

“If you can’t build a great product, it doesn’t matter if you can build a great company.”

03

ALSO TODAY

3 MORE SOURCES
08

SOLIDOT

08.00
SOLIDOT

Solidot News - September 23, 2026

Solidot Feed: Highlighting essential tech & open-source news.

年检显示高里程电动车比汽油车更可靠

对 4740 万英国机动车年检(MOT test)数据的分析发现,当汽车行驶里程达到 9-12 万英里时,电动汽车的年检不合格率为汽油车同类车型的 75%(16.5% 对 22.1%)。行驶里程超过 12 万英里后,电动汽车的不合格率为 16%,而汽油车为 23.5%。研究发现,较低行驶里程两种动力类型的汽车之间的不合格率差异相对较小。研究还发现,电动汽车的一大问题是其轮胎磨损问题两倍于燃油车。英国机动车年检没有检查电池的健康状况,因此电动汽车电池健康情况未知。研究人员表示他们的研究驳斥了高里程电动汽车应该报废的观念。

不要被 AI 炒作愚弄

Anthropic 声称其模型 Claude Mythos 在发现软件漏洞上胜过大多数安全专家。随后发生了 OpenAI–Hugging Face 安全事件,此后 Anthropic(自豪)和 Meta(不情愿)也披露了各自模型的类似事件。紧接着 Anthropic 宣称其模型取得了数学领域的突破;OpenAI 也声称自己取得了数学突破。Anthropic 工程师 Jacob Coxon 在宣布离职时引发了广泛关注,他声称该公司与 OpenAI 正“冲向自我进化的超级智能,并拿我们的生命在赌博”。媒体大肆报道了这些事件,且沿用了相关公司赋予其软件的拟人化叙事——即把软件描绘成不仅功能强大,而且已初具通用人工智能(AGI)雏形的产物。但深入研究的专家则给出了不同的答案,虽然这些发现并不能吸引眼球。网络安全专家指出,涉及模型的安全事件更多是 OpenAI 的疏忽大意,未能采取基本的安全措施,而不是“模型失控”或“AI 智能体创造文明”。OpenAI 模型在解决数学难题上的突破其原创性也相当可疑。数学家公开对 AI 企业利用其专业领域进行炒作提出了警告。AI 公司通过炒作模型失控也将自己置身事外,将责任归咎于大模型而不是公司本身,逃避应承担的责任。以 OpenAI 为例,当该公司开发的恶意软件被用于入侵另一家公司时,媒体、名人和议员谈论是“失控模型”而不是 OpenAI 的责任,仿佛大模型真的会自动发动攻击,公众的注意力被转移到虚构的“超级智能”的恐惧之上。我们不要被 AI 公司的炒作所愚弄。

英国准备施压 Google 向 Android 和 Chrome 用户展示 AI 助手选择屏

英国竞争监管机构 CMA 想要让 Android 和 Chrome 用户对 AI 助手和搜索引擎有更大的选择权和控制权。CMA 公布了一份提案,要求 Google 在用户首次设置 Android 手机或打开 Chrome 浏览器时,向其展示多种搜索引擎供选择,并且每年提示用户选择一个默认搜索引擎;符合技术与安全标准的 AI 助手也必须获准出现在选择屏上上。该提案目前进入公众咨询阶段,截止日期为 10 月 9 日,CMA 预计将在今年底前做出最终决定。

全球陆地热浪更早到来、发展得更快

中科院研究人员的一项研究发现,自 1979 年以来,全球陆地热浪开始时间显著提前、结束时间显著推迟,热浪季节明显延长;进入21世纪以来,发生快速起始型首次热浪的陆地面积占当年热浪影响区面积的比例显著增加。研究团队基于1979-2023年全球气候数据,系统分析了全球陆地热浪开始时间、结束时间、热浪季节长度及每年首次热浪起始速度的长期变化,并利用多个独立气候数据集对结果进行交叉验证。结果显示,全球陆地首次热浪发生时间平均每 10 年提前约 3.3 天,过去 45年 总体提前约 2 周;最后一次热浪结束时间平均每 10 年推迟约 5.4 天,45 年间总体推迟约 24 天;热浪季节平均每 10 年延长约 8.7 天,45 年间总体延长约 39 天。从空间范围看,全球 72.0% 的陆地区域呈现热浪提前发生趋势,79.6% 的区域呈现热浪推迟结束趋势,92.1% 的区域呈现热浪季节延长趋势,其中干旱地区的变化总体更为明显。研究还发现,热浪季节延长增加了农作物在关键生育阶段遭遇高温的风险。2001-2023年,全球部分主要作物在开花、抽丝等关键生殖生长阶段的热浪暴露面积占比较 1979-2000 年增加 1.6%-13.3%,热浪影响进一步向作物生育期的前期和后期扩展。

美国准备再次制裁 ICC

荷兰正在为位于海牙的国际刑事法院(ICC)面临美国新一轮制裁做准备。荷兰正研究如何协助法院维持运作,包括支付员工薪酬、保护证人以及维护拘留设施。美国已制裁了十多名现任和前任 ICC 工作人员,国务卿国卢比奥(Marco Rubio)表示,这是一场彻底瓦解 ICC 所构成威胁的全面行动。ICC 有 125 个成员国,美国、以色列等都未加入该机构。美国的制裁可能会导致法院无法使用金融和 IT 服务,甚至两年无法向美籍员工支付薪酬。当 ICC 前首席检察官在 2025 年遭到制裁时,他不仅失去了对微软电邮账户的访问权限,银行账户也被冻结,还被禁止进入美国。ICC 数月来一直在为可能面临的制裁做准备。法院在今年早些时候已停止使用微软产品,转而采用一家德国软件供应商的服务。法庭还更换了保险等金融服务提供商,改用在美国没有业务往来的公司。

八种常用食品防腐剂与高血压相关

对法国 112,395 人七、八年间健康状况与饮食习惯的研究显示,八种常用食品防腐剂与高血压风险升高相关。摄入防腐剂最多的人群患高血压的风险高 24%。摄入非抗氧化类防腐剂最多的人群患心血管疾病的相对风险高 16%。非抗氧化类防腐剂通过抑制细菌或真菌生长而非防止氧化防止食品变质。这八种添加剂包括: 山梨酸钾(E202),常用于加工水果和饮料;焦亚硫酸钾(E224),常用于葡萄酒等含酒精饮料;亚硝酸钠(E250)、抗坏血酸钠(E301)和异抗坏血酸钠(E316)均用于加工肉类;抗坏血酸(E300):常添加于加工水果和蔬菜;柠檬酸(E330):常用于软饮料;迷迭香提取物(E392):常添加于油脂类产品。研究人员强调这是一项观察性研究,结果并不能证明食品添加剂会导致高血压或直接引发心脏病,只能显示两者之间可能存在关联。

丰田命令员工训练人形机器人,否认会替代人类员工

丰田的员工正在帮助训练人形机器人,它计划未来几年部署 40 万台工厂机器人,人形机器人是其中的一部分,但高管否认它们会替代人类员工。丰田在装配线上部署了 ELEY 人形机器人,而员工则通过佩戴基于机器人手指的装置去训练这些通过轮子移动的机器人执行需要精细手部动作的任务。丰田是全球最大的汽车公司之一,在全世界有 60 座工厂,雇佣了 1.8 万员工。丰田执行副总裁 Hiroki Nakajima 表示,公司的目标是创造一个“机器人与人类共存,而非取代人类”的世界。根据国际机器人联合会的数据,仅 2024 年中国就部署了 200 万台工业机器人,日本以 45.05 万台位居第二。美国和韩国分别以 3.42 万台和 3.06 万台的部署量排名第三和第四。

地热变形虫能在 63 摄氏度下生存

研究人员从加利福尼亚拉森火山国家公园温泉中分离出的“喀斯喀特火变形虫(Incendiamoeba cascadensis)”,可在 63℃ 完成有丝分裂,并在 70℃ 形成保护外层、降温后恢复,打破此前真核生物约 60℃ 的生长上限。研究团队于 2023—2025 年在喀斯喀特山脉拉森火山国家公园采集地热溪流样本,水温约 47℃—64℃。团队回实验室后先以 57℃ 培养地热变形虫,该温度已高于既往已知变形虫生长最高值;随后逐步升温,至 63℃ 直接观察到有丝分裂,证明它不仅存活,且能在超出旧有真核生物耐受极限的条件下繁殖。超过 63℃ 后,虫体改变形态并形成保护外层;暴露于 70℃ 后再放回较低温度,仍能复原。基因组测序显示,火变形虫与蛋白质维持、DNA 修复相关的基因数量多于温带变形虫近缘种,对应高温下蛋白变性、DNA 损伤等压力。团队还发现其蛋白质表面带正电氨基酸比例更高,这一特征在部分耐热细菌和古菌中亦有出现,这有助于减少高温引起的展开与聚集。

Adobe 推出 Android 版免费视频编辑工具 Premiere

Adobe 终于推出了功能完整的 Android 版视频编辑工具 Premiere,除了 AI 视频生成功能外其它功能都是免费的,用户无需订阅或注册 Creative Cloud 账户。Android 版 Premiere 支持导入任意数量的视频轨道,进行剪辑、分割、添加特效,导出最高 4K 分辨率的视频。它支持从视频片段中提取音频、降低背景噪音以及添加旁白。该工具默认输出竖屏 16:9 比例的视频,允许用户切换到更传统的比例。对于 AI 功能,该工具每月免费提供 250 AI credits,足以以每个消耗 80 credits 生成几段短视频。AI 功能的收费是每月 8 美元或每年 70 美元。

黑客声称入侵了 FBI 窃取雇员信息

勒索组织 ShinyHunters 声称入侵了 FBI 窃取了逾 2TB 雇员数据。该组织的一名发言人称,这次行动不是出于经济动机,而是要求 FBI 更正或撤回此前发表的声明,其中包含大量不实的指控。ShinyHunters 称它利用了 FBI 招聘网页的一个 Oracle PeopleSoft 的 0day 漏洞,该漏洞允许在服务器上远程执行代码。该组织随后篡改了页面,替换为已被其控制的横幅和图片(This site has been seized by ShinyHunters)。ShinyHunters 从 FBI 管理的 AWS GovCloud 服务器上下载了约 2TB 至 3TB 的数据,这些数据涉及 FBI 的现有和前雇员,以及求职者。 FBI 在今年五月就 ShinyHunters 发出安全警告,称该组织采用“骚扰策略,向受害者及其家属发送威胁性短信和拨打骚扰电话,在某些情况下还包括恶意报假警(swatting)”。ShinyHunters 声称这些指控不实。

新 Halo 游戏将由动视开发

微软 Xbox 游戏业务宣布旗下第一方工作室 Halo Studios 等裁员 268 人,动视将负责下一代 Halo 游戏的开发,而原来负责开发 Halo 的 Halo Studios 则转变成辅助工作室角色。Obsidian 工作室将成为 Bethesda 的一部分,将继续开发 Grounded 以及新 Fallout 游戏。King 工作室将合并微软的休闲游戏部门 Microsoft Casual Games。开发 Forza 系列的 Playground 和开发新 Fable 游戏的 Turn 10 将合并为一家工作室。 Ninja Theory 工作室预计将会关闭。Arkane 工作室仍然在磋商中。

美国酒精消费自疫情以来首次下降

盖洛普 8 月民调显示,仅有 54% 的美国人饮酒,而 2010 年这一数字是 67%。千禧一代和 Z 世代推动了减少饮酒的趋势,而 50-64 岁的中老年人的酒精消费则在上升。发表在《Annals of Internal Medicine》期刊上的一项研究分析了逾 11.4 万名美国成年人的调查数据,受访者在 2018-2024 年间参加了 CDC 的年度健康调查 National Health Interview Survey,其中包括了饮酒的情况。结果显示,2022-2024 年期间,美国人的总体饮酒率下降了 2%,重度饮酒率下降了 8%。Z 世代的总体饮酒率降幅最为显著下降了近 6%,千禧一代降幅约 2%。与此同时,2018-2024 年间 50-64 岁中老年人总体饮酒量率加了 4%,重度饮酒率激增了 36%。相比中老年人,年轻人更了解酒精的负面影响。老年人也可能更富裕能承担更多酒精消费。

天文学家发现已知最年轻行星

天文学家发现了已知最年轻的行星——Elias 2-24 b。这颗不到 100 万年的木星大小行星仍被其形成时的气体和尘埃包围。其令人惊讶的形成速度挑战了关于巨行星如何形成的主流理论。现有的行星形成理论认为,一颗大质量行星不可能形成地如此迅速,尤其是在距离恒星如此遥远的地方。当前模型表明,在太阳系中木星的位置形成一颗木星大小的行星大约需要 500 万年,那么距恒星更远的巨行星形成时间应该更长。然而 Elias 2-24 系统中这个微弱天体到恒星的距离约为地球到太阳距离的 55 倍,而且已经显示出行星形成的迹象。Elias 2-24 b的质量与木星相当,围绕一颗距离地球约 450 光年的恒星运行。由于该系统还很年轻,天文学家通过研究它可以一瞥数十亿年前太阳系的样子。

为躲避亿万富翁税 Larry Page 等人迁出加州

对加州亿万富翁征收一次性 5% 税的提案 Initiative Number 25-0024 将在 11 月 3 日进行公投。胡佛研究所的研究显示,面临征税的亿万富翁们已有近三成迁出加州。Larry Page 在迈阿密 Coconut Grove 购买了两栋临水豪宅,总价 1.73 亿美元,同时将家族办公室 Koop 从加州转到注册地特拉华州、办公地址佛罗里达的公司。Sergey Brin 在迈阿密 Allison Island 购买了一栋价值 5100 万美元的临水豪宅,将内华达州登记为正式居住地,他资助了反对征税的政治行动委员会 Building a Better California。Peter Thiel 在 2025 年 12 月将其家族投资公司从加州迁至迈阿密。英伟达 CEO 黄仁宇则是少数公开表示会纳税的亿万富翁,他预计将缴纳 80 亿美元的税。

NASA 火星样本采集送回任务终止

美国国会的预算法案取消了对 NASA 火星样本采集送回任务 Mars Sample Return(MSR)的资助,虽然该法案还需要通过国会两院的批准以及总统的签署才会生效,但实际上代表着 MSR 计划的终止。MSR 计划因为不断膨胀的成本而备受争议,2024 年其成本膨胀至 110 亿美元,如果推行将占用大部分 NASA 科学预算。2025 年 NASA 设法将项目成本降至 70 亿美元,但费用仍然过高,而 NASA 同时正面临特朗普政府削减科学预算的挑战。MSR 项目的终止意味着火星漫游车毅力号收集的样本无法送回样地球实验室进行分析。

6 岁女孩打破女子三阶魔方还原世界纪录

成都六岁女童连允之在世界魔方协会(World Cube Association)在两场赛事中,三天内两次打破了女子三阶魔方世界纪录,成为全球唯一一位平均还原时间低于 4.5 秒的女子魔方选手。她分别在武汉和广州举行的比赛中以 4.52 秒和 4.27 秒的平均成绩两次刷新了纪录。她过去一年进行了高强度训练,每天投入两到三个小时练习魔方,目前已掌握逾 1300 种魔方还原算法。

阿里巴巴下一代模型参数将扩大到 5-10 万亿规模

阿里巴巴 CEO 吴泳铭在阿里云年度云栖大会上透露,该公司的千问(Qwen)团队正持续研究模型架构与数据优化,目标是完成更复杂、长周期任务,向人工智能超级智能(artificial superintelligence)前进。阿里巴巴的旗舰模型 Qwen 3.8 Max 有 2.4 万亿参数,正在训练中的 Qwen 4 将会继续扩大参数规模,未来的 Qwen 4.5 和 Qwen 5 系列将扩大至 5-10 万亿参数规模。吴泳铭表示,阿里巴巴自研的 M890 AI 超级节点能处理参数规模逾 2 万亿模型的推理任务,新一代的真武 V900 性能三倍于上一代的 M890,预计于 2027 年第一季度量产。

AMD 加入万亿美元市值俱乐部

AMD 周一股价上涨 9.6% 至 613.31 美元,市值突破一万亿美元,成为英伟达、博通和美光之后第四家市值突破万亿美元的美国芯片公司。英伟达在 2023 年市值突破万亿美元,如今市值逾五万亿美元,是全世界市值最高的公司。AMD 被认为是英伟达在 GPU 芯片领域最强大的竞争对手。AMD 股价在 2026 年上涨 185%,远超聚集众多科技股的纳斯达克指数 15.8% 的涨幅,是标普 500 指数中表现最佳的股票之一。

Google 因地理位置数据处理被爱尔兰罚款 4.03 亿欧元

Google 因地理位置数据处理被爱尔兰数据保护委员会(DPC)罚款 4.03 亿欧元。DPC 对 Google 的调查持续了六年,涉及 Google 在 2018 年 5 月 25 日至 2020 年 2 月 4 日间 Web & App Activity、Location History 和 Location Accuracy 三项功能的位置数据处理。DPC 的报告认为 Google 的位置数据处理违反了 2018 年生效的数据保护法律 GDPR,可能导致用户未意识到自己的位置信息正被用于投放定向广告或推断其兴趣偏好,丧失对自己个人数据的控制权。Google 发表声明,表示它从 2019 年起就调整了位置数据管理。引入了位置数据自动删除功能。

Googlebooks 于 10 月 4 日上市,最低 899 美元

深度集成 Gemini、运行 Android 的笔记本电脑 Googlebooks 将于 10 月 4 日上市。Google 硬件合作伙伴中除了宏碁推出一款起售价 899 美元的型号外,其余厂商的产品都超过 1000 美元。Googlebook 不同于 Chromebook 面向低端市场,它面向的是中端笔记本电脑市场。Googlebooks 的 Continue On 功能允许用户在手机或 Googlebook 之间无缝切换,但需要应用开发者支持;Cast My Apps 可以直接在 Googlebook 上使用 Android 手机已安装应用;Play Store 是 Googlebook 获取应用的主要渠道,侧载受到了限制,只能安装运行已通过 Google 验证身份的开发者的应用;通过深度集成 Gemini Intelligence,用户仅仅移动光标就能激活被称为“Magic Pointer”的 AI 功能,AI 会分析屏幕上的内容,根据上下文提供建议,能从多个应用中提取数据。比如将光标指向电邮中的日期即可创建日历预约。

09

APP STORE RANK

09.00
APP STORE RANK
Loading…
TEXT VIEW · TODAY'S DIGEST · 36 HEADLINES ACROSS 8 SOURCES

Startup Archive(0)

No items yet for today.

App Store Rankings(0)

No items yet for today.