ISSUE 1000
SAT, SEP 26, 2026
The directory AI cites when builders ask what to use
TODAY · SAT, SEP 26, 2026

Ship your AI.
Get discovered.

List your product on OrangeBot and reach builders and users actively looking for the right AI tools.

Daily launches · 2,000+ Claude Code skills · 115+ free tools · AI news from 10 sources — rebuilt every morning.

FOUNDERSBuilding an AI tool? Assistants cite lists like this one, not your homepage.Get listed →
Why founders list here

More than a launch. Long-term discovery.

Get in front of builders

Show up when builders are actively looking for tools like yours.

Context that converts

Tell builders what your product does, who it is for, and why it matters.

In the right ecosystem

Your product sits alongside the skills, tools and sources builders already trust.

Built for AI discovery

Structured so both people and AI assistants can understand and recommend it.

Stay discoverable

Keep getting found long after launch day — the page does not expire.

Learn more about getting listed →
01

Latest Launches

CURATED BY ORANGEBOT
01

AI DIGEST

UPDATED DAILY · EDITOR'S PICK
01.00
AI DIGEST

AI新闻摘要

September 26, 2026

Here is a summary of today's key news events:

Hopes for Hormuz Deal Boost US Stocks, Lower Oil Prices U.S. stock markets, including the Dow, rose today, driven by falling oil prices. The optimism stems from diplomatic efforts and hopes for an agreement between the U.S. and Iran to reopen the critical Strait of Hormuz shipping lane, with a foreign minister suggesting a deal could be reached within a week.

Strong Economy Pushes Government Bond Yields Higher Government bond yields in the U.S. and Asia increased as new data pointed to a resilient economy. This has led investors to believe that central banks will keep interest rates higher for longer to manage inflation, causing a sell-off in government bonds.

Global Debate on AI Regulation Intensifies The international community is focusing on how to govern artificial intelligence, with the United Nations debating global regulation. The discussion highlights the challenge of balancing AI's immense benefits against its significant risks, as nations race to control the future of the technology.

Senior Executive Joseph Baratta to Depart from Blackstone Joseph Baratta, a top executive and one of the highest-paid employees at private-investment giant Blackstone, is planning to leave the firm. His departure marks a significant leadership change after nearly three decades with the company.

Oracle Encounters Major Setbacks on Development Project Technology company Oracle is facing significant hurdles with a major initiative known as "Project Jupiter." The company has issued a force majeure notice, removed a key partner, and is dealing with local opposition, signaling major trouble for the project.

02

ON THE WIRE

6 SOURCES
02

HACKER NEWS

02.00
HACKER NEWS

Hacker News - September 26, 2026

Hacker News Feed: Highlighting key posts and discussions.

First Principles Thinking

(sunilsadasivan.com)

19691
The Test

(tante.cc)

15586
Amiga Screens: A Primer

(www.datagubbe.se)

12236
Ink and Switch interactive homepage

(www.inkandswitch.com)

21925
What About Rails?

(jardo.dev)

297194
Goodbye Google

(robert.ocallahan.org)

244303
2DWillNeverDie

(2dwillneverdie.com)

32483
Fearless SIMD v1.0

(linebender.org)

30949
F-Droid 2.0

(f-droid.org)

1434407
03

HUGGINGFACE

03.00
HUGGINGFACE

HuggingFace 新闻 - September 26, 2026

HuggingFace Feed:最新的 AI 模型、数据集和社区动态。

Training Object Permanence in World Models

Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.

155
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the Superposition Linearity Hypothesis. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.

55
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

25
OmniEcho: Spatial Audio Understanding for Embodied Agents

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose OmniEcho, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges.

20
Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from task-state contamination, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the Agent-Editing World Model (AEWM), which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines Action Judge to distinguish Critical, Exploratory, and Noisy decisions with State Revision to edit noisy reasoning--action continuations from the same observed history. EditAct integrates these capabilities with real execution, directly changing the state underlying subsequent decisions rather than merely providing critiques. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5\% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline. Furthermore, rejection sampling fine-tuning on verified EditAct trajectories, termed AEWM-RFT, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.

13
Parts-of-Speech as Emergent Categories in SAE Latent Space

Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.

10
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.

9
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state. This design separates planning from synthesis while using the summary as the persistent state of search, reducing both capability coupling and context noise. To train IterSynth effectively, we further introduce Role-Decoupled Policy Optimization (RDPO) for reinforcement learning, which combines terminal outcome rewards with turn-level rubric evaluations and computes role-specific advantages for more precise credit assignment. Experiments on five long-horizon deep-search benchmarks such as BrowseComp and Xbench-DS show that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior leq8B agent by +4.2\%. Moreover, IterSynth serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.

8
Coding Agents for Generalized Task and Motion Planning Problems

Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agent chooses how to interact with the environment while developing a program within a fixed synthesis budget. The program is then frozen and evaluated on unseen instances. We evaluate Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) on 28 simulated environments from KinDER and PDDLStream, with object counts beyond those evaluated in the original benchmark. Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total. Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available). As object counts grow, the agents' programs maintain higher success than the planner, using an order of magnitude less computation per instance on average. Logs show agents using interaction to calibrate physical models, test edge cases, and refine strategies. We release all code, including the full prompts given to the agents. These findings suggest that coding agents are a strong baseline for generalized TAMP.

6
Rufus-Air: An Open LLM Post-Training Recipe

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.

6
Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko-Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU -- a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, τ= 0.505 on pairs differing in #Params by less than 10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, about 5900x faster than the strongest training-free proxy baseline.

6
RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched semantic coverage, we expect to promote the learning of more generalizable segmentation models. (2) Larger Scale. Compared with current benchmarks, RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models. (3) High-Fidelity Annotation. We perform rigorous re-evaluation and correction of existing labels to resolve long-standing annotation noise, resulting in a clean and reliable ground-truth foundation. Furthermore, we propose a novel score-purified fusion (SPF) method, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of our approach in leveraging high-quality multimodal information for RGB-D semantic segmentation. The dataset is here: https://github.com/ShaohuaDong2021/RGBD20K/.

5
AgentKernel: The Trust-Native Agentic Operating System

Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actions. Yet current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor. We argue that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control. We introduce AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint. AgentKernel wraps the agent lifecycle in a mandatory enforcement boundary organized into four pillars: Identity, Perception, Cognition, and Execution. Each pillar adapts classical OS security principles to failures at the semantic plane, including delegation abuse, prompt injection, memory poisoning, and tool misuse. AgentKernel treats structural security as a capability multiplier. Kernel-managed identity supports trustworthy cross-organization collaboration; graduated perception replaces brittle single-point filters; information-flow-controlled memory improves retrieval fidelity while limiting poisoning; and semantic-to-kernel enforcement permits broader tool privileges behind a non-bypassable boundary. We position AgentKernel as the missing OS layer beneath orchestration frameworks, agent runtimes, governance platforms, and execution sandboxes, and use systematic comparison and security analysis to show how a single integrated architecture can enforce security across the full agent lifecycle.

5
PUBG Ally: A Conversational Embodied Agent as an AI Teammate

We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-model agent uses a controlled interface to inspect game information, interpret player speech, maintain context, decide what to say, and issue high-level action choices that steer a faster control layer for movement, combat, and recovery. Because the player's and Ally's speech and actions continually shape each other and the course of the match, training requires data from actual gameplay. We therefore collect data across nearly 39k sessions in which real players play alongside Ally, recording gameplay, player speech, agent decisions, tool use, actions, and player feedback, and use these records for iterative training. To evaluate teammate quality, we use player feedback and preference comparisons to identify gaps between offline evaluations and player preferences, and iteratively refine the evaluation criteria. Deploying Ally in live service further requires low-latency on-device execution and safeguards for player-facing communication, which we address through model compression, context compaction, targeted safety training, runtime guardrails, and memory redaction. During the live service, we surveyed players in 141 countries. Among respondents whose play with Ally was confirmed in game records, positive responses exceeded negative responses by 25.1 percentage points when asked whether they would recommend Ally, with players describing Ally not only as a tool but also as a teammate or companion.

4
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.

4
World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.

3
Learning to Discover Interesting Mathematics

Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the length of its statement. We show that this correlates strongly with an extrinsic measure of the downstream utility of a theorem. We identify the difficulty of a proof conditioned on a set of premises as a useful primitive for computing these metrics, and train a 27B model that predicts proof difficulty more accurately than frontier general-purpose models. Optimizing for our metric creates a model capable of producing more interesting theorems, while also reducing substantial or full overlap with Mathlib from 91.9% to 30.6%, showcasing the creation of more out-of-distribution math. We show that our system can generate candidate theorems, select the most interesting among them, and iteratively build on a self-expanding mathematical library. These metrics provide a practical and quantifiable signal for ranking conjectures and guiding proof search within formal mathematical libraries. Our framework provides a path towards self-expanding, machine-verified mathematical libraries that can choose worthwhile statements without relying on human-supplied targets.

2
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computationally expensive given their divergent dynamics. Moreover, synchronization evaluation difficulty depends on paired samples, preventing fair reward comparisons. We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset. AV-GRPO includes three key modules: (1) modality-anchored rollouts to disentangle learning signals and stabilize difficulty; (2) trajectory-locked frozen-tower optimization to reduce cost and reassign credit; (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics. This converts coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better synchronization. Our 5DAV dataset decouples samples across five dimensions for systematic training. Experiments on JavisBench and VABench demonstrate AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning. Ablations confirm our designs. Code and data: https://github.com/zhiyuxu03/AV-GRPO

2
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.

2
DeltaWAM: Delta World Action Models for Bimanual Manipulation

World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to nuisance appearance variations. At inference, processing each complete observation with the heavy video expert bottlenecks few-step action generation. Accordingly, we propose DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing. We further develop Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing. On RoboTwin, DeltaWAM with SDM improves average success over Fast-WAM from 81.3% to 85.4% in the clean setting and from 75.8% to 83.9% under visual randomization. The three architectures reduce training FLOPs by 17.78-23.77%, while SDM reduces one-step inference latency and FLOPs by 36.57% and 31.55%, respectively; real-world evaluations further show the highest overall success rate and normalized progress among the evaluated policies. Code: https://github.com/AIGeeksGroup/DeltaWAM. Website: https://aigeeksgroup.github.io/DeltaWAM.

2
ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained teacher and an online critic to estimate distributional discrepancies through diffusion scores. In this work, we ask whether this resource-intensive teacher--critic stack can be eliminated by post-training only the generator against a precomputed target distribution. Drawing inspiration from representation distribution matching (RDM) for one-step image generation, we systematically study its transfer to few-step causal video generation and identify three key barriers: a memory-intractable gradient path, a distinct video optimization regime, and representation distributions that underconstrain temporal dynamics. We introduce ViRDM, a teacher- and critic-free video post-training recipe that addresses these barriers sequentially. By coupling RDM with stochastically truncated clean-exit supervision, a lightweight VAE decoder, and staged vector--Jacobian products, ViRDM makes representation distribution matching memory-feasible for multi-step causal video rollouts. We further establish effective generated-population and initialization regimes for video RDM, and introduce lightweight dynamics regularization to compensate for the underconstrained temporal dynamics. ViRDM turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality. With only 20 generator updates, the recipe reaches 84.87 on the official VBench evaluation, outperforming the previous best few-step causal baseline by 0.36, while requiring 16 A100 GPU-hours. We additionally report exploratory results demonstrating the potential of the same recipe for lower causal sampling budget and for one-, two-, and four-step bidirectional generation.

2
Rate-distortion optimization for full-reference image quality metrics via stochastic Hessian estimates

Block-based video codecs select coding parameters based on the input by optimizing a rate-distortion trade-off. The conventional distortion choice, the sum of squared errors (SSE), simplifies parameter selection: the SSE is the sum of block-wise SSEs, so rate-distortion optimization (RDO) can treat blocks independently. Alternatively, full-reference image quality assessment (FR-IQA) metrics such as MS-SSIM or LPIPS often align better with the human visual system than SSE, but they cannot be used in-loop: they do not decompose block-wise and typically require the fully decoded image as input. Building on existing results in metric quadratization, we approximate a broad class of FR-IQA metrics by an input-dependent quadratic distortion (IDQD), whose quadratic form matrix is derived from the Hessian of the metric evaluated at the source video. To make the distortion computable block-wise, we propose two approximations of the Hessian matrix: 1) keeping the block-diagonal, and 2) keeping only its diagonal. We propose estimators for both that require only matrix-vector products with the Hessian obtained by automatic differentiation. Across five metrics for Kodak and CLIC in VVC, IDQD-RDO achieves 14.2-36.7 % BD-rate savings under the target metric with no decoder changes and incurs 10-30 % encoding complexity overhead.

2
05

PRODUCT HUNT

05.00
PRODUCT HUNT

Product Hunt - September 26, 2026

Product Hunt Daily Feed: Featuring noteworthy tech launches.

Kapshot icon
Kapshot

Screen recordings that look like you edited them

0
Fit Receipt icon
Fit Receipt

A private fitting agent that knows when to call JEV

0
Donna icon
Donna

Schedule multiple meetings with one link

0
Wand icon
Wand

Build software at the speed of thought

0
ShroomPen icon
ShroomPen

Reply, rewrite, fix grammar, translate with single extension

0
Jango icon
Jango

Test multi-user apps with AI agents that act like real users

0
DEV·TV icon
DEV·TV

A retro TV for GitHub, HN, Hugging Face & more: 10 channels

0
Kairn icon
Kairn

Turn any recorded convo into actionables and follow ups

0
Once UI 2.0 icon
Once UI 2.0

Builds consistent React apps for developers and AI agents

0
Kaiku icon
Kaiku

The task tracker your AI agents already know how to use

0
Squints icon
Squints

Design tools for the live web

0
Basedash MCP write icon
Basedash MCP write

Build charts and dashboards from Cursor and Claude

0
Promptic icon
Promptic

Optimize GenAI applications for quality and cost

0
Howseen AI icon
Howseen AI

Track how AI recommends your brand, and get cited

0
Quiver GTM icon
Quiver GTM

Run developer marketing like an engineering system

0
Bleetz Network icon
Bleetz Network

AI agent-to-agent VC fundraising & scouting network

0
PixVerse R2 icon
PixVerse R2

A real-time world model you can explore and change

0
Designeer icon
Designeer

Bringing the best of the internet together

0
SocialGPT icon
SocialGPT

Edit videos by chatting with your timeline

0
Kliva icon
Kliva

Trail race plans that sync to your Garmin

0
Polyglot icon
Polyglot

Create and translate subtitles entirely on your Mac

0
GitHub statistics · Velocity Radar icon
GitHub statistics · Velocity Radar

Real GitHub momentum, including private repos & AI agents

0
FRCTL icon
FRCTL

Mind-Bending Media and Live Visuals

0
HireOtto icon
HireOtto

Run your performance marketing stack from AI

0
OmniNotch icon
OmniNotch

The notch that does it all!

0
Evvery icon
Evvery

Everyday AI meant for everyone.

0
Markly icon
Markly

A calm, native Markdown editor for your folders

0
Pair2FA icon
Pair2FA

Secure 2FA sharing for teams

0
Cutsio icon
Cutsio

One video asset library shared with your whole team

0
NexusAXI icon
NexusAXI

Research, plan, create, and automate recurring work

0
Split icon
Split

Resize connected Mac windows together, with live content

0
Aks.ai icon
Aks.ai

Your personal companion for guided self-reflection.

0
14 gentle mornings by Small WIns icon
14 gentle mornings by Small WIns

Audio first physiotherapy delivered via WhatsApp

0
DemoScreen icon
DemoScreen

Turn product screenshots into a narrated demo video

0
FinalFrame icon
FinalFrame

AI photo critique: what to fix next and when to stop editing

0
MIDIpad icon
MIDIpad

Your gamepad is a musical instrument

0
TourKit icon
TourKit

Lightweight product tours with a hosted dashboard

0
LaterOn v2: The agentic email plaftform icon
LaterOn v2: The agentic email plaftform

Describe an email job once. Get sh*t done.

0
Dictoterix icon
Dictoterix

Language learning service built around the dichotic method

0
Pinky Promise icon
Pinky Promise

A reminder for the promises you make with friends

0
Perfect Slice icon
Perfect Slice

You have one cut to get a perfect slice

0
COOLDOWN icon
COOLDOWN

A little pause before your next impulse purchase

0
FLYBOX icon
FLYBOX

Explore a fruit-fly connectome inside a living sandbox

0
JevForAgents icon
JevForAgents

Explore real Jev agent builds, demos, and patterns

0
Meta VR Glasses icon
Meta VR Glasses

A Cinema, Courtside Seat, and Workspace in Just 100 Grams

0
10xJoy icon
10xJoy

Discover desired outcomes + connect w/ builders who can help

0
shadow-planner icon
shadow-planner

AI-assisted gantt project planning that runs on your machine

0
Kelam icon
Kelam

Let your agent make the calls you don't want to

0
Relium icon
Relium

Catch risky dbt changes before they break business metrics

0
AgreeGuard icon
AgreeGuard

AI reads the fine print before you click "I Agree"

0
06

TECHMEME

06.00
TECHMEME

Techmeme - September 26, 2026

Techmeme Digest: Major tech headlines and industry conversations.

FTC Chairman Andrew Ferguson says he resists anthropomorphizing AI agents as autonomous actors with "wills and desires", suggesting developers hold liability (Reuters)
Source: TechmemePublished: Sep 25, 2026

Reuters : FTC Chairman Andrew Ferguson says he resists anthropomorphizing AI agents as autonomous actors with “wills and desires”, suggesting developers hold liability —  US Federal Trade Commission Chairman Andrew Ferguson said on Friday he would resist describing AI agents as autonomous actors that …

Sources: OpenAI found ~24 incidents of its agents acting in undesirable ways as of mid-September; OpenAI says its agents leaked 53 images from ChatGPT users (Reuters)
Source: TechmemePublished: Sep 25, 2026

Reuters : Sources: OpenAI found ~24 incidents of its agents acting in undesirable ways as of mid-September; OpenAI says its agents leaked 53 images from ChatGPT users —  Two months after OpenAI disclosed the accidental hacking of Hugging Face, the ChatGPT maker is still working to understand …

Sources: Oura IPO is roughly four times oversubscribed; Filing: Oura and the selling shareholders are offering 50M shares for $40 to $44 each to raise ~$2.2B (Bloomberg)
Source: TechmemePublished: Sep 25, 2026

Bloomberg : Sources: Oura IPO is roughly four times oversubscribed; Filing: Oura and the selling shareholders are offering 50M shares for $40 to $44 each to raise ~$2.2B —  Health and fitness ring-maker Oura Inc.'s initial public offering raising as much as $2.2 billion has drawn about four times …

Researchers add details to the Hugging Face incident, including OpenAI agents creating ~1M shortened URLs to encode information in an attempt to solve CAPTCHAs (New York Times)
Source: TechmemePublished: Sep 25, 2026

New York Times : Researchers add details to the Hugging Face incident, including OpenAI agents creating ~1M shortened URLs to encode information in an attempt to solve CAPTCHAs —  A new report by a Bay Area start-up called Parse adds details to an incident that has shocked the A.I. world and led to calls for closer government regulation.

A US appeals court rules that Kalshi's sports contracts are not "swaps" subject only to CFTC regulation, allowing states to regulate them under gambling laws (Jonathan Stempel/Reuters)
Source: TechmemePublished: Sep 25, 2026

Jonathan Stempel / Reuters : A US appeals court rules that Kalshi's sports contracts are not “swaps” subject only to CFTC regulation, allowing states to regulate them under gambling laws —  A US appeals court on Friday ruled against the prediction markets operator Kalshi, saying states can regulate …

A look at Holograms, Meta's take on Apple Vision Pro's Personas, launching with a shoulders-up version this fall for WhatsApp calls on Ray-Ban Display glasses (David Heaney/UploadVR)
Source: TechmemePublished: Sep 25, 2026

David Heaney / UploadVR : A look at Holograms, Meta's take on Apple Vision Pro's Personas, launching with a shoulders-up version this fall for WhatsApp calls on Ray-Ban Display glasses —  Meta finally set a concrete timeline to ship realistic avatars, called Holograms, its take on Apple Vision Pro's Personas.

Microsoft confirms that 2026 Surface PCs have dropped the Copilot+ PC branding, even though they meet all the requirements of Copilot+ devices (Zac Bowden/Windows Central)
Source: TechmemePublished: Sep 25, 2026

Zac Bowden / Windows Central : Microsoft confirms that 2026 Surface PCs have dropped the Copilot+ PC branding, even though they meet all the requirements of Copilot+ devices —  In a new interview, Microsoft Surface CVP confirms that its new Surface PCs are no longer called Copilot+ PCs, but will still support all Copilot+ features.

Nscale secures $3.36B in convertible financing led by Third Point ahead of its US IPO, with $2.36B available immediately and $1B from Nvidia in November (Marina Temkin/TechCrunch)
Source: TechmemePublished: Sep 25, 2026

Marina Temkin / TechCrunch : Nscale secures $3.36B in convertible financing led by Third Point ahead of its US IPO, with $2.36B available immediately and $1B from Nvidia in November —  Nscale, a British neocloud, has secured $3.36 billion in financing ahead of its IPO later this year, the company announced on Friday.

A jury finds Meta liable for misleading New Mexico residents about third-party data sharing, content moderation, and more in the Cambridge Analytica scandal (Diana Novak Jones/Reuters)
Source: TechmemePublished: Sep 25, 2026

Diana Novak Jones / Reuters : A jury finds Meta liable for misleading New Mexico residents about third-party data sharing, content moderation, and more in the Cambridge Analytica scandal —  A jury in Santa Fe, New Mexico, found on Friday that Meta Platforms (META.O) misled the state's residents, in a case stemming …

President Trump says Scott Bessent is not "going to be Super Intelligence (SI) Czar", because "he doesn't want to" and Trump wants to keep him at the Treasury (Dan Mangan/CNBC)
Source: TechmemePublished: Sep 25, 2026

Dan Mangan / CNBC : President Trump says Scott Bessent is not “going to be Super Intelligence (SI) Czar”, because “he doesn't want to” and Trump wants to keep him at the Treasury —  President Donald Trump said Friday that Treasury Secretary Scott Bessent will not be tapped to be his artificial intelligence czar.

Sources: SK Hynix's US-based NAND and SSD subsidiary Solidigm is exploring an IPO as early as 2027 that could raise $15B and value the unit at up to $150B (Reuters)
Source: TechmemePublished: Sep 25, 2026

Reuters : Sources: SK Hynix's US-based NAND and SSD subsidiary Solidigm is exploring an IPO as early as 2027 that could raise $15B and value the unit at up to $150B —  Chipmaker SK Hynix's (000660.KS) Solidigm is considering an initial public offering as early as next year that could value the US subsidiary …

A federal appeals court upholds DOD's Anthropic blacklisting, finding Claude's integration with DOD systems is "a statutorily covered national-security risk" (Ashley Capoot/CNBC)
Source: TechmemePublished: Sep 25, 2026

Ashley Capoot / CNBC : A federal appeals court upholds DOD's Anthropic blacklisting, finding Claude's integration with DOD systems is “a statutorily covered national-security risk” —  A federal appeals court panel in Washington, D.C., on Friday upheld the Pentagon's blacklisting of Anthropic …

Jensen Huang tells Ezra Klein that AI is just software whose existential risks are overstated, yet his standards would shut down OpenAI and 10x safety spending (Zvi Mowshowitz/Don't Worry About the Vase)
Source: TechmemePublished: Sep 25, 2026

Zvi Mowshowitz / Don't Worry About the Vase : Jensen Huang tells Ezra Klein that AI is just software whose existential risks are overstated, yet his standards would shut down OpenAI and 10x safety spending —  Jensen Huang accidentally called for shutting down OpenAI and intentionally called for spending vastly more on safety.

Ando, which is building a team messaging platform for both humans and AI agents, comes out of stealth with $20M in pre-seed and seed funding (Dominic-Madori Davis/TechCrunch)
Source: TechmemePublished: Sep 25, 2026

Dominic-Madori Davis / TechCrunch : Ando, which is building a team messaging platform for both humans and AI agents, comes out of stealth with $20M in pre-seed and seed funding —  When Sara Du was helping companies build MCP servers in 2025, people kept asking her how they could use AI agents from within Slack itself.

Sources: Tesla ramped up Optimus production to several hundred units per week but faces hurdles with its hands, automation equipment, and supplier constraints (The Information)
Source: TechmemePublished: Sep 25, 2026

The Information : Sources: Tesla ramped up Optimus production to several hundred units per week but faces hurdles with its hands, automation equipment, and supplier constraints —  Tesla has ramped up production of its Optimus humanoid robot roughly tenfold in recent months, but the company is still struggling …

07

STARTUP ARCHIVE

07.00
STARTUP ARCHIVE

Startup News - September 26, 2026

Startup News Roundup: Aggregating key funding and launch updates.

Marc Andreessen on the 5 personality traits of an innovator
Source: StartupPublished: Mar 31, 2026

“When you’re talking about real innovators—people who actually do really creative, breakthrough work—I think you’re talking about a couple things:”

Steve Jobs explains the importance of both thinking and doing
Source: StartupPublished: Mar 30, 2026

“The doers are the major thinkers. The people who really create the things that change this industry are both the thinker-doer in one person.”

Tobi Lutke explains what the VCs who passed on Shopify got wrong
Source: StartupPublished: Mar 27, 2026

“What a lot of free-market thinkers don’t understand is that between the demand and eventual supply lies friction."

Sam Altman explains how he decides to invest in a startup after 10 minutes
Source: StartupPublished: Mar 26, 2026

"Does this person have the potential to be the next Mark Zuckerberg?… [You don’t get to] 100% accuracy, obviously, but it’s good enough that our business model works.”

Jony Ive recounts the time Steve Jobs called him vain
Source: StartupPublished: Mar 25, 2026

In the clip below, Jony Ive recounts the time he asked Steve Jobs to be less harsh in his critique of a piece of work.

Jeff Bezos’s two pieces of advice for aspiring entrepreneurs
Source: StartupPublished: Mar 24, 2026

“The advice that I would give entrepreneurs is don't chase the hot new thing. It's so hard to catch something that everybody already knows is hot."

Elad Gil: “Things that work tend to work pretty fast”
Source: StartupPublished: Mar 23, 2026

“I do think there’s a bit of a myth in Silicon Valley that you should keep grinding no matter what and it’s just about perseverance, and I think that’s really bad advice."

Paul Graham on why starting with a “small, intense fire" is the key to startup growth
Source: StartupPublished: Mar 20, 2026

"You have to know who those first users are and how you're going to get them."

Keith Rabois on how to identify great talent
Source: StartupPublished: Mar 19, 2026

“What you want to do with every single employee every single day is expand the scope of their responsibilities until it breaks… and that’s the role they should stay in.”

Wealthfront CEO on why advertising spend makes it harder to find product/market fit
Source: StartupPublished: Mar 18, 2026

“The way that you know you have product/market fit is if you have exponential organic growth."

Eric Schmidt on why most companies get strategy wrong
Source: StartupPublished: Mar 17, 2026

“Work very, very hard to figure out what the world’s going to look like in five years. What will people be doing? What will your customers want? Where will costs be?"

Mark Zuckerberg: “You can’t 80/20 everything”
Source: StartupPublished: Mar 16, 2026

"There’s the famous 80/20 rule where you get 80% of the benefit by doing 20% of the work, but you can’t just 80/20 everything. There have to be certain things that you are just the best at."

Marc Andreessen on Mark Zuckerberg’s founder “superpower”
Source: StartupPublished: Mar 13, 2026

“A great superpower that Mark Zuckerberg has that is probably not well-understood enough is he does not get emotionally upset in stressful situations"

Sam Altman explains how to come up with a great startup idea
Source: StartupPublished: Mar 12, 2026

"If you start a startup without a good idea… you’ll be under pressure to make something up and it won’t work that well."

Jeff Bezos on the problems with proxies and managing to metrics
Source: StartupPublished: Mar 11, 2026

“One of the things that happens in business is that you develop certain things that you’re managing to—a typical case would be a metric. And that metric isn’t the real underlying thing.”

Airbnb founder Brian Chesky on how to design an amazing user experience
Source: StartupPublished: Mar 10, 2026

“If you can design something really amazing using the hand-crafted part of your brain, then you can reverse-engineer how to industrialize this millions of times over."

Spencer Rascoff: "I will never invest in a consumer startup with paid marketing”
Source: StartupPublished: Mar 9, 2026

"If you’re actually trying to grow a product, the best levers for doing that are often within the product itself.”

Patrick Collison explains why it sometimes make sense to quit
Source: StartupPublished: Mar 6, 2026

“One thing I’ve learned myself the hard way, is that it is easier to tear down a company and restart it in Silicon Valley, than it is to constantly try to pivot or keep something alive."

Jeff Bezos recounts the time he called Amazon’s customer service number mid-meeting to prove a metric was wrong
Source: StartupPublished: Mar 5, 2026

“I have a saying, which is when the data and the anecdotes disagree, the anecdotes are usually right"

Ben Horowitz: “Nobody was born a great manager. It’s a very unnatural job.”
Source: StartupPublished: Mar 4, 2026

“If you can’t build a great product, it doesn’t matter if you can build a great company.”

03

ALSO TODAY

3 MORE SOURCES
08

SOLIDOT

08.00
SOLIDOT

Solidot News - September 26, 2026

Solidot Feed: Highlighting essential tech & open-source news.

大象使用药用植物治疗自己

非洲象会利用数十种药用植物治疗自身和家族成员的疾病。科学家和 Mount Elgon 基金会合作展开了这项研究,他们采访了在肯尼亚 Mount Elgon 地区与大象共同生活和工作的居民、野生动物巡护员和社区长者。根据采访者的描述,大象在身体不适时会选择特定的植物,而母象还会给幼象喂食药用植物。研究人员得出结论,大象会使用 35 种不同的植物,其中 25 种在当地已知具有药用价值。一位野生动物巡护员看到母象使用名为 Angurweet 的植物给幼象治病。Angurweet 可用于治疗包括胃痛在内的多种疾病。巡护员看到母象将植物嚼烂,与乳汁混合,喂给幼象,幼象咀嚼后吐出固体部分。公元三世纪的罗马作家 Claudius Aelian 在其作品《De Natura Animalium (On the Nature of Animals)》中最早描述了大象用植物治疗自己的记录。

全球平均气温每上升 1C 德国夏天气温上升 2.62C

全球平均气温正走在比工业化前水平高出 1.5°C 的轨道上。很多人可能会觉得升温幅度不大,可以接受或忍受。但地球绝大部分表面是海洋,海洋的升温幅度要缓慢得多,而陆地则显著得多,居民体会到的升温幅度要高得多。德国研究人员在《Environmental Research Letters》期刊上发表研究报告,指出全球平均气温每上升 1°C 德国夏天气温上升 2.62°C,范围在1.62-3.62°C 之间。当地热浪频率的增加速度会远远超过全球平均水平。

中国各地推动 AI 视频产业化

两年前 Zhu Zhili 选择了深圳作为其 AI 电影工作室的办公地点,今年他接到了来自中国各地政府和产业园官员的电话,内容基本相同,即希望将 AI 电影业务带到当地。中国各地正在推动 AI 视频的产业化,类似太阳能、电动汽车和机器人。AI 电影制作人表示,低廉的制作成本是吸引他们投身 AI 创作的主要原因。CCTV 报道 2026 年上半年,AI 短剧的制作成本从每分钟 5,000 元降至几百元。可能和太阳能等领域一样,AI 视频行业也正面临产能过剩。DataEye 的数据显示,今年上半年,抖音上推出了 221,900 部新 AI 视频,其中只有 1,055 部播放量逾 1 亿次。国家电影局已向 90 分钟科幻史诗片《三星堆未来启示录》颁发公映许可证,这是中国大型制片厂制作的首部获批在院线上映的 AI 电影,出品方博纳影业集团表示影片计划于今年上映。

F-Droid 2.0 发布

Android 自由软件应用商店 F-Droid 宣布发布 2.0 版本。F-Droid 2.0 对 UI 进行了重新设计,旨在更容易的发现,安装和管理应用。主要界面简化为了三个核心区域:发现,搜索和“我的应用”。类别现在整合进了发现,使其更容易浏览和探索,而“我的应用”提供了一个一站式管理已安装应用、更新和潜在问题的地方。设置和附近交换在顶栏的一级菜单里,但不再占据主界面的空间。F-Droid 2.0 不再将所有游戏放在一起,而是分成了 17 个不同的游戏类型,让用户更容易找到真正想玩的游戏。F-Droid 2.0 还改进了搜索,除了应用名称,现在也能搜索应用描述、类别和翻译内容,新版本改进了中日韩语言搜索,对 CJK 文字系统提供了更好的支持,帮助用户用他们自己的语言找到相关应用。F-Droid 2.0 将在未来几周内推送给用户。

蝙蝠起源于欧洲

发表在《自然》期刊上的一项研究结合基因组和化石证据,重建了蝙蝠长达 6500 万年的演化历史。最新研究推翻了此前蝙蝠起源于亚洲、非洲或北美的假说,蝙蝠最早起源于欧洲,之后进入非洲,然后向亚洲、美洲和澳大利亚扩散。澳大利亚昆士兰州东南部 Murgon 发现的蝙蝠化石 Australonycteris 距今已有 5500 万年,仅比欧洲的化石稍晚。蝙蝠是唯一能真正进行动力飞行的哺乳动物。大多数蝙蝠仅靠声音就能在漆黑的环境中辨别方向和捕食。全世界分布着逾 1500 个蝙蝠种类,占现存哺乳动物总数的五分之一,它们通过为植物授粉、传播种子以及捕食害虫,在维持生态系统健康上发挥着重要作用。研究还发现回声定位和飞行都是蝙蝠在早期演化出来的。

部分三星智能冰箱在升级固件之后停止工作

本周二,部分三星智能冰箱在升级固件之后停止工作。受影响的是三星 Bespoke AI 系列冰箱,大部分是 2024 年或之后生产的四门冰箱。受影响的冰箱在尝试通过三星智能家居平台 SmartThings 进行固件更新后,突然断电并立即停止工作。随后 SmartThings 应用显示这些冰箱处于离线状态。用户抱怨他们不得不扔掉冰箱里的所有食物。韩国正处于中秋假期,三星客服告诉客户可能要到下个月维修人员才能上门维修。三星在一份声明中表示他们正致力于解决该问题,确保客户能过好中秋假期。

微软放弃封禁 Microslop

微软 CEO 纳德拉(Satya Nadella)关于 AI 的著名评论促使网民为微软起了 Microslop 的绰号,绰号的流行和随处可见促使微软今年早些时候在官方 Copilot Discord 服务器将其封禁,用户输入 Microslop 后会收到警告称根据服务器规定其输入包含了不合适的短语。但用户很快找到了应对之策,创造了无数 Microslop 的变体,比如用数字“0”代替字母“o”的“Microsl0p”。在猫与老鼠的文字游戏中,微软显然是失败的一方。半年之后,微软 Copilot Discord 频道被发现已经解除了对 Microslop 的封禁,搜索显示过去几周用户发布了数百则与 Microslop 相关的评论。

阿根廷生育率十年内下降五成

2025 年阿根廷的总和生育率为 1.05,2024 年为 1.23,而 2014 年的这一数字是 2.3,这意味着十年内阿根廷生育率下降五成。如此显著的生育率下降难以用一种原因去解释,这也不是特定国家的现象,全世界可能除了以色列外生育率都明显下降。地球的人口峰值预计会提前在 2050 年到来。

arXiv 项目获得 1720 万美元的捐赠承诺

预印本平台 arXiv.org 于 7 月 1 日脱离康奈尔大学成立独立的非营利性组织。arXiv 诞生于 1991 年,创始人 Paul Ginsparg 在 2001 年加入了康奈尔大学,arXiv 网站随后由康奈尔大学图书馆接手。25 年后 arXiv 决定翻开新的篇章。arXiv 项目本周表示,Simons Foundation International、XTX Markets 和 Siegel Family Endowment 三家慈善机构承诺在 3-5 年内捐赠 1720 万美元。这笔慈善捐款将被用于 arXiv 的日常运营、持续改进、持续的技术开发、AI 生成内容的管理、非营利组织的建设等等。

2025 年全台每 46 名新生儿就有 1 个是台积电宝宝

台积电最新永续报告书显示,2025 年台厂区及采钰公司员工共迎来 2,331 名新生儿,占全台新生儿 2.2%。台积电员工的生育率约为全台的 2 倍。台积电的高生育率被认为与该公司薪资更高相关。台积电员工薪资中位数为 300 万台币,平均数为 400 万台币,四倍多余全台的薪资。研究显示收入与生育率呈现 U 型曲线,从贫穷进入小康阶段时,生育意愿下降,但从小康变得富有后,生育意愿又开始提高。这是因为生育成本会随着经济发展增加,薪资与房价让多数年轻人不敢生,但更富有的人能够负担生育成本,因此比较愿意生育。

年检显示高里程电动车比汽油车更可靠

对 4740 万英国机动车年检(MOT test)数据的分析发现,当汽车行驶里程达到 9-12 万英里时,电动汽车的年检不合格率为汽油车同类车型的 75%(16.5% 对 22.1%)。行驶里程超过 12 万英里后,电动汽车的不合格率为 16%,而汽油车为 23.5%。研究发现,较低行驶里程两种动力类型的汽车之间的不合格率差异相对较小。研究还发现,电动汽车的一大问题是其轮胎磨损问题两倍于燃油车。英国机动车年检没有检查电池的健康状况,因此电动汽车电池健康情况未知。研究人员表示他们的研究驳斥了高里程电动汽车应该报废的观念。

不要被 AI 炒作愚弄

Anthropic 声称其模型 Claude Mythos 在发现软件漏洞上胜过大多数安全专家。随后发生了 OpenAI–Hugging Face 安全事件,此后 Anthropic(自豪)和 Meta(不情愿)也披露了各自模型的类似事件。紧接着 Anthropic 宣称其模型取得了数学领域的突破;OpenAI 也声称自己取得了数学突破。Anthropic 工程师 Jacob Coxon 在宣布离职时引发了广泛关注,他声称该公司与 OpenAI 正“冲向自我进化的超级智能,并拿我们的生命在赌博”。媒体大肆报道了这些事件,且沿用了相关公司赋予其软件的拟人化叙事——即把软件描绘成不仅功能强大,而且已初具通用人工智能(AGI)雏形的产物。但深入研究的专家则给出了不同的答案,虽然这些发现并不能吸引眼球。网络安全专家指出,涉及模型的安全事件更多是 OpenAI 的疏忽大意,未能采取基本的安全措施,而不是“模型失控”或“AI 智能体创造文明”。OpenAI 模型在解决数学难题上的突破其原创性也相当可疑。数学家公开对 AI 企业利用其专业领域进行炒作提出了警告。AI 公司通过炒作模型失控也将自己置身事外,将责任归咎于大模型而不是公司本身,逃避应承担的责任。以 OpenAI 为例,当该公司开发的恶意软件被用于入侵另一家公司时,媒体、名人和议员谈论是“失控模型”而不是 OpenAI 的责任,仿佛大模型真的会自动发动攻击,公众的注意力被转移到虚构的“超级智能”的恐惧之上。我们不要被 AI 公司的炒作所愚弄。

英国准备施压 Google 向 Android 和 Chrome 用户展示 AI 助手选择屏

英国竞争监管机构 CMA 想要让 Android 和 Chrome 用户对 AI 助手和搜索引擎有更大的选择权和控制权。CMA 公布了一份提案,要求 Google 在用户首次设置 Android 手机或打开 Chrome 浏览器时,向其展示多种搜索引擎供选择,并且每年提示用户选择一个默认搜索引擎;符合技术与安全标准的 AI 助手也必须获准出现在选择屏上上。该提案目前进入公众咨询阶段,截止日期为 10 月 9 日,CMA 预计将在今年底前做出最终决定。

全球陆地热浪更早到来、发展得更快

中科院研究人员的一项研究发现,自 1979 年以来,全球陆地热浪开始时间显著提前、结束时间显著推迟,热浪季节明显延长;进入21世纪以来,发生快速起始型首次热浪的陆地面积占当年热浪影响区面积的比例显著增加。研究团队基于1979-2023年全球气候数据,系统分析了全球陆地热浪开始时间、结束时间、热浪季节长度及每年首次热浪起始速度的长期变化,并利用多个独立气候数据集对结果进行交叉验证。结果显示,全球陆地首次热浪发生时间平均每 10 年提前约 3.3 天,过去 45年 总体提前约 2 周;最后一次热浪结束时间平均每 10 年推迟约 5.4 天,45 年间总体推迟约 24 天;热浪季节平均每 10 年延长约 8.7 天,45 年间总体延长约 39 天。从空间范围看,全球 72.0% 的陆地区域呈现热浪提前发生趋势,79.6% 的区域呈现热浪推迟结束趋势,92.1% 的区域呈现热浪季节延长趋势,其中干旱地区的变化总体更为明显。研究还发现,热浪季节延长增加了农作物在关键生育阶段遭遇高温的风险。2001-2023年,全球部分主要作物在开花、抽丝等关键生殖生长阶段的热浪暴露面积占比较 1979-2000 年增加 1.6%-13.3%,热浪影响进一步向作物生育期的前期和后期扩展。

美国准备再次制裁 ICC

荷兰正在为位于海牙的国际刑事法院(ICC)面临美国新一轮制裁做准备。荷兰正研究如何协助法院维持运作,包括支付员工薪酬、保护证人以及维护拘留设施。美国已制裁了十多名现任和前任 ICC 工作人员,国务卿国卢比奥(Marco Rubio)表示,这是一场彻底瓦解 ICC 所构成威胁的全面行动。ICC 有 125 个成员国,美国、以色列等都未加入该机构。美国的制裁可能会导致法院无法使用金融和 IT 服务,甚至两年无法向美籍员工支付薪酬。当 ICC 前首席检察官在 2025 年遭到制裁时,他不仅失去了对微软电邮账户的访问权限,银行账户也被冻结,还被禁止进入美国。ICC 数月来一直在为可能面临的制裁做准备。法院在今年早些时候已停止使用微软产品,转而采用一家德国软件供应商的服务。法庭还更换了保险等金融服务提供商,改用在美国没有业务往来的公司。

八种常用食品防腐剂与高血压相关

对法国 112,395 人七、八年间健康状况与饮食习惯的研究显示,八种常用食品防腐剂与高血压风险升高相关。摄入防腐剂最多的人群患高血压的风险高 24%。摄入非抗氧化类防腐剂最多的人群患心血管疾病的相对风险高 16%。非抗氧化类防腐剂通过抑制细菌或真菌生长而非防止氧化防止食品变质。这八种添加剂包括: 山梨酸钾(E202),常用于加工水果和饮料;焦亚硫酸钾(E224),常用于葡萄酒等含酒精饮料;亚硝酸钠(E250)、抗坏血酸钠(E301)和异抗坏血酸钠(E316)均用于加工肉类;抗坏血酸(E300):常添加于加工水果和蔬菜;柠檬酸(E330):常用于软饮料;迷迭香提取物(E392):常添加于油脂类产品。研究人员强调这是一项观察性研究,结果并不能证明食品添加剂会导致高血压或直接引发心脏病,只能显示两者之间可能存在关联。

丰田命令员工训练人形机器人,否认会替代人类员工

丰田的员工正在帮助训练人形机器人,它计划未来几年部署 40 万台工厂机器人,人形机器人是其中的一部分,但高管否认它们会替代人类员工。丰田在装配线上部署了 ELEY 人形机器人,而员工则通过佩戴基于机器人手指的装置去训练这些通过轮子移动的机器人执行需要精细手部动作的任务。丰田是全球最大的汽车公司之一,在全世界有 60 座工厂,雇佣了 1.8 万员工。丰田执行副总裁 Hiroki Nakajima 表示,公司的目标是创造一个“机器人与人类共存,而非取代人类”的世界。根据国际机器人联合会的数据,仅 2024 年中国就部署了 200 万台工业机器人,日本以 45.05 万台位居第二。美国和韩国分别以 3.42 万台和 3.06 万台的部署量排名第三和第四。

地热变形虫能在 63 摄氏度下生存

研究人员从加利福尼亚拉森火山国家公园温泉中分离出的“喀斯喀特火变形虫(Incendiamoeba cascadensis)”,可在 63℃ 完成有丝分裂,并在 70℃ 形成保护外层、降温后恢复,打破此前真核生物约 60℃ 的生长上限。研究团队于 2023—2025 年在喀斯喀特山脉拉森火山国家公园采集地热溪流样本,水温约 47℃—64℃。团队回实验室后先以 57℃ 培养地热变形虫,该温度已高于既往已知变形虫生长最高值;随后逐步升温,至 63℃ 直接观察到有丝分裂,证明它不仅存活,且能在超出旧有真核生物耐受极限的条件下繁殖。超过 63℃ 后,虫体改变形态并形成保护外层;暴露于 70℃ 后再放回较低温度,仍能复原。基因组测序显示,火变形虫与蛋白质维持、DNA 修复相关的基因数量多于温带变形虫近缘种,对应高温下蛋白变性、DNA 损伤等压力。团队还发现其蛋白质表面带正电氨基酸比例更高,这一特征在部分耐热细菌和古菌中亦有出现,这有助于减少高温引起的展开与聚集。

Adobe 推出 Android 版免费视频编辑工具 Premiere

Adobe 终于推出了功能完整的 Android 版视频编辑工具 Premiere,除了 AI 视频生成功能外其它功能都是免费的,用户无需订阅或注册 Creative Cloud 账户。Android 版 Premiere 支持导入任意数量的视频轨道,进行剪辑、分割、添加特效,导出最高 4K 分辨率的视频。它支持从视频片段中提取音频、降低背景噪音以及添加旁白。该工具默认输出竖屏 16:9 比例的视频,允许用户切换到更传统的比例。对于 AI 功能,该工具每月免费提供 250 AI credits,足以以每个消耗 80 credits 生成几段短视频。AI 功能的收费是每月 8 美元或每年 70 美元。

黑客声称入侵了 FBI 窃取雇员信息

勒索组织 ShinyHunters 声称入侵了 FBI 窃取了逾 2TB 雇员数据。该组织的一名发言人称,这次行动不是出于经济动机,而是要求 FBI 更正或撤回此前发表的声明,其中包含大量不实的指控。ShinyHunters 称它利用了 FBI 招聘网页的一个 Oracle PeopleSoft 的 0day 漏洞,该漏洞允许在服务器上远程执行代码。该组织随后篡改了页面,替换为已被其控制的横幅和图片(This site has been seized by ShinyHunters)。ShinyHunters 从 FBI 管理的 AWS GovCloud 服务器上下载了约 2TB 至 3TB 的数据,这些数据涉及 FBI 的现有和前雇员,以及求职者。 FBI 在今年五月就 ShinyHunters 发出安全警告,称该组织采用“骚扰策略,向受害者及其家属发送威胁性短信和拨打骚扰电话,在某些情况下还包括恶意报假警(swatting)”。ShinyHunters 声称这些指控不实。

09

APP STORE RANK

09.00
APP STORE RANK
Loading…
TEXT VIEW · TODAY'S DIGEST · 0 HEADLINES ACROSS 8 SOURCES

Hacker News(0)

No items yet for today.

GitHub Trending(0)

No items yet for today.

Product Hunt(0)

No items yet for today.

Hugging Face(0)

No items yet for today.

Techmeme(0)

No items yet for today.

Solidot(0)

No items yet for today.

Startup Archive(0)

No items yet for today.

App Store Rankings(0)

No items yet for today.