ISSUE 0976
WED, SEP 2, 2026
The directory AI cites when builders ask what to use
TODAY · WED, SEP 2, 2026

Ship your AI.
Get discovered.

List your product on OrangeBot and reach builders and users actively looking for the right AI tools.

Daily launches · 2,000+ Claude Code skills · 113+ free tools · AI news from 10 sources — rebuilt every morning.

FOUNDERSBuilding an AI tool? Assistants cite lists like this one, not your homepage.Get listed →
Why founders list here

More than a launch. Long-term discovery.

Get in front of builders

Show up when builders are actively looking for tools like yours.

Context that converts

Tell builders what your product does, who it is for, and why it matters.

In the right ecosystem

Your product sits alongside the skills, tools and sources builders already trust.

Built for AI discovery

Structured so both people and AI assistants can understand and recommend it.

Stay discoverable

Keep getting found long after launch day — the page does not expire.

Learn more about getting listed →
01

Latest Launches

CURATED BY ORANGEBOT
01

AI DIGEST

UPDATED DAILY · EDITOR'S PICK
01.00
AI DIGEST

AI新闻摘要

September 2, 2026

Here is a summary of today's main news events, based on the information provided.

Global Markets Tumble Amid Mideast Tensions and Surging Bond Yields

U.S. military strikes against Iranian targets in the Middle East caused a sharp spike in oil prices. This escalation fueled fears of a wider conflict and renewed concerns about inflation, causing government bond yields to surge globally. Japan's 10-year yield hit a three-decade high, and U.S. stock markets fell for a third straight session, with the tech-heavy Nasdaq leading the decline.

Tech Giants Advance AI Amid Security and Employment Concerns

Google is preparing to release a new, more powerful version of its AI model, Gemini, after internal tests showed progress. However, testing also revealed the new AI is capable of executing complex cyberattacks, prompting new security measures. Meanwhile, new research from Texas showed a drop in job postings for occupations exposed to AI, highlighting the technology's growing impact on the labor market.

Amazon Sued by FTC; Major Banks Plan Stablecoin Venture

The U.S. Federal Trade Commission (FTC) filed a lawsuit against Amazon, alleging the e-commerce giant illegally manipulated its ad auction system, costing advertisers billions. In separate news, a group of 21 major financial firms announced plans to form a new company to launch a stablecoin, aiming to compete with existing digital assets and create a new payment infrastructure.

International Political Tensions and Domestic Shake-ups Emerge

Geopolitical frictions were evident as Beijing objected to G-20 language targeting its "non-market" economic policies. In the UK, a former prime minister's resignation is set to trigger a competitive by-election. Meanwhile, Germany's interior minister stated the country is a "daily target of hybrid warfare," signaling heightened security concerns in Europe.

New Climate Data Shows Unprecedented Global Warming

A new scientific report confirmed that the five warmest seasons ever recorded have all occurred since 2003. Researchers attribute this clear trend to the accelerating effects of human-caused climate change.

02

ON THE WIRE

6 SOURCES
02

HACKER NEWS

02.00
HACKER NEWS

Hacker News - September 2, 2026

Hacker News Feed: Highlighting key posts and discussions.

Ambient CSS v3 – Blender meets CSS

(ambientcss.vercel.app)

17365
Tmp.0ut Volume 5

(tmpout.sh)

19138
GPU World

(www.gpuworld.org)

390265
Restroom Archive

(restroomarchive.com)

36183
Fastpotify

(fastpotify.rocks)

798532
Run macOS Software on Linux

(www.darlinghq.org)

27784
Show HN: Laser Graffiti

(laser.consti.de)

24759
Playa Phone

(playaphone.com)

735225
ChatGPT Work Tool and Skill Reference

(codex-tool-reference.simonw.chatgpt.site)

23760
Damn fine tiny cafe

(sandyuraz.com)

38963
Understanding ChatGPT Work

(simonwillison.net)

346194
03

HUGGINGFACE

03.00
HUGGINGFACE

HuggingFace 新闻 - September 2, 2026

HuggingFace Feed:最新的 AI 模型、数据集和社区动态。

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base Qwen3-1.7B, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263\% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.

90
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.

86
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises [email protected] from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.

67
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation-reconstruction conflict. We revisit this problem by analyzing how different objectives shape the latent space and identify two key insights. First, the entropy term in the Kullback-Leibler divergence objective is essential for preventing collapse: reconstruction and prior fitting tend to shrink the posterior, while entropy preserves non-degenerate latent uncertainty. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these insights, we achieve the first direct end-to-end training without latent collapse and propose GenFirst, a simple generation-before-reconstruction strategy. The generative objective first shapes the latent space under weak reconstruction pressure, after which reconstruction is progressively strengthened to recover visual details. We validate GenFirst with continuous autoregressive priors with exact likelihoods and SiT priors with implicit likelihoods. With our end-to-end objective and GenFirst, SiT achieves a gFID of 0.97 with CFG and 1.45 without CFG on ImageNet-256, while MMDiT reaches a GenEval score of 0.90 on text-to-image generation. Beyond image generation, we extend the framework to shared visual latents for generation and representation learning, and to continuous unified text-image generation. These results demonstrate the generality of stable end-to-end latent learning across generative priors and modalities.

57
Normalized Low-Rank Adaptation

While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. Building on this observation, we introduce Normalized Low-Rank Adaptation (NoRA), a simple yet effective method that normalizes the down-projection matrices during training. We further show that the same normalization can be applied only at initialization, improving standard LoRA without requiring repeated normalization throughout training. Across pretraining, supervised finetuning, and reinforcement learning, NoRA consistently accelerates convergence, improves performance and training stability, and mitigates catastrophic forgetting. These benefits require neither additional trainable parameters nor inference-time computation, making NoRA a simple and broadly applicable enhancement to LoRA.

38
PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks paired with a critic. Rubrics extracted from scientific papers can supply the critic. Existing pipelines, however, draw the question and the criteria from the same content, so the reward can be earned by paraphrase. The rubric is further compressed into a single scalar per rollout. We introduce PaperGym, a unified framework that turns each research paper into a complete training environment. PaperGym exploits the structure of a paper: the question is synthesized from the research goal and background, while the criteria are derived from the method and experiments. The criteria span methodological innovation and experimental design, and criterion leakage falls to 3.7%, versus 11.90% to 34.10% in existing datasets. Training uses the rubric twice: first as privileged context for OPSD's self-teacher, then as the reward for GRPO. Across Qwen3-1.7B/4B/8B, this schedule outperforms supervised fine-tuning, either stage alone, and the reverse ordering, improving five-benchmark averages by +5.6, +5.0, and +4.8 points. With the recipe held fixed, models trained on PaperGym-20k win 58.1% of three-way comparisons, against 28.2% for RubricHub Science. The trained Qwen3-8B reaches 73.48 on ResearchQA, above the far larger Kimi K2.6. We release the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.

35
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.

32
CogEvol: Towards Efficient and Reliable Learning Environment Generation

We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.

25
Evaluating the Hidden Costs of Personalization in Large Language Models

While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on personal context such as conversation history, inferred preferences, and user profiles. Specifically, we identify three emerging risks: (1) irrelevant personalization, where models reference personal information in unnecessary contexts; (2) preference narrowing, where models reinforce informational echo chambers; and (3) sycophantic bias, where models agree excessively with user opinions. As a result, models may reference personal information in contexts where it is unnecessary, inadvertently collapse response diversity, or agree excessively with user opinions. Despite the growing use of personalization in AI assistants, there has been limited systematic evaluation of its potential side effects. To bridge this gap, we propose PRISK, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses. Our empirical analysis across 13 LLMs demonstrates the presence of user profiles and retrieved memories consistently exacerbates biases, resulting in an average drop of 45.9% in irrelevant personalization, 41.7% in preference narrowing and 61.7% in sycophantic bias.

24
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.

24
SHAPE of Chain-of-Thought in Math Reasoning

Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce SHAPE, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics education: (1) semantic spaces: the model's evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2) heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use SHAPE to analyze the reasoning patterns of various models. Our findings reveal that the mathematical heuristics employed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a few semantic spaces rather than exploring many disparate ones -- a pattern consistent with human behavior. Next, we utilize the SHAPE lens to evaluate whether post-training truly enhances mathematical proficiency. We find that reinforcement learning induces mode-seeking in heuristic usage. Lastly, we post-train LLMs by promoting diverse heuristics and demonstrate its effectiveness in improving accuracy. Overall, SHAPE provides a theoretically-grounded diagnostic framework for decoding LLM reasoning and offers a new path toward post-training LLMs for math reasoning. The code for our model is available at https://github.com/holi-lab/SHAPE-of-CoT

24
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase

Organizations often develop and maintain portfolios of related applications: independently deployable codebases that share substantial domain logic, interface patterns, or operational conventions. As LLM coding agents are increasingly used to generate and maintain such software, a naive application-by-application workflow duplicates shared logic across codebases and allows prolonged agentic maintenance to accumulate verbosity, dead code, and structural erosion. We introduce the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components. A minimal sequential scaffold can in principle extract shared code and migrate applications to the evolving library, but in practice suffers from low extraction recall and fragile dependency migration. We address these failures with candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information. Across WebGen-Bench and PaperBench, our method preserves application functionality while significantly reducing redundancy and token footprint (verbosity, token length) over zero-shot, and avoiding the structural erosion introduced by naive library construction, with additional reductions in LOC and MDL. Our code is available at https://github.com/sbigstar0310/super-library-agent.

23
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.

21
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.

11
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.

11
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).

10
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.

9
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with variational autoencoders (VAEs) to enhance computational efficiency without compromising visual quality. However, conventional VAEs are suboptimal for video data as they employ fixed compression ratios that cannot adapt to the varying complexity of spatio-temporal content. We present KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens. By evaluating each token's content-richness as keep-or-drop probability, the token selector effectively discards uninformative tokens, naturally allowing data-dependent compression. Applying adaptive tokenization to diffusion models may cause spatial misalignment, as token dropping can disturb the original spatio-temporal structure. To alleviate this issue, we propose two position-prediction strategies: cascaded and joint generation, to ensure spatial consistency. We empirically show that our model achieves strong reconstruction and generation quality at a state-of-the-art compression ratio. Further analysis on video data reveals that this improvement is primarily achieved by reducing spatio-temporal redundancy and removing uninformative tokens, as supported by both quantitative and qualitative results.

9
MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.

7
Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching

Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent within personal photo collections, where users often seek compact visual stories bound by structural relations rather than isolated snapshots. To address this limitation, we introduce **Image Bundle Composition (IBC)**, a novel paradigm that shifts the objective from ranking individual images to dynamically composing cohesive image bundles from a massive, unstructured photo pool. Since target bundles are not predefined, IBC presents a severe combinatorial explosion challenge and demands modeling non-decomposable joint relevance. To establish this paradigm, we construct **IBCBench**, the first IBC benchmark dataset containing 109,467 images and 667 verified queries, built via a semi-automated verification pipeline. Furthermore, we propose **BundleWeaver**, an agentic framework that reformulates IBC as query-conditioned incremental hyperedge discovery. By employing a Large Language Model to adaptively search for missing relational roles and utilizing a Vision-Language Model for whole-bundle verification, BundleWeaver effectively navigates the combinatorial space. Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition. Our dataset and code are available.

6
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.

6
Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation

Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop behavior. Today's foundation models split these capabilities: vision-language models (VLMs) infer missing information and adapt high-level plans but remain brittle and inefficient at repeated navigation grounding, while navigation foundation models (NFMs) robustly execute semantic goals but operate as bounded episodes without persistent task-level reasoning. We introduce NavMCP, an agentic scaffolding framework that couples a VLM reasoning agent with an NFM executor for long-horizon exploration. The VLM decides what evidence to seek, where to search, and when to stop, while the NFM grounds each semantic sub-goal into closed-loop navigation. Three channels structure their collaboration: intent translates evidence needs into navigation calls, observation converts rollouts into source-grounded trajectory evidence, and memory accumulates findings, negative evidence, and unresolved goals across calls. This design turns isolated navigation rollouts into persistent embodied interaction without retraining either model. On Embodied Question Answering, NavMCP achieves state-of-the-art results on HM-EQA, MT-HM3D, and EXPRESS-Bench. Under matched agent and executor backbones, it outperforms an episodic interface by 14.9 percentage points on HM-EQA. On a Unitree Go2, NavMCP reaches 78.3% success, with its margin over the strongest baseline growing from 10 to 45 points as the task horizon increases. These results demonstrate the potential of scaffolding complementary foundation models into long-horizon physical-world agents.

6
CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments. In these settings, a single wrong action, such as refunding the wrong purchase, can cause irreversible task failure and must be intercepted before execution. Such failures may not appear in every single run, but can emerge across repeated trials, making reliability across steps and trials critical. However, ensuring agentic reliability is challenging: even frontier LLMs struggle to explain why an action may be wrong, especially in long, intertwined trajectories governed by domain-specific policies. Much recent work relies on prompt-based critique agents, while optimization-based methods lack a systematic way to produce rich verification rationales for training. We address this gap with CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization. CAST analyzes agent trajectories to synthesize structured rationales explaining action validity under partial observability. The resulting critique model is used to construct critique-aware training data for optimizing the policy model. Fine-tuning Qwen3-family models on dynamic tool-calling benchmarks, CAST improves reliability across domains, outperforming GPT-OSS-120B by over 10% pass^4 on Retail tasks and yielding an additional 9% improvement on Telehealth in an out-of-domain setting. These results demonstrate that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.

6
WebWorld: The Browser as a World Model for Self-Improving Web Code

VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor proxy for whether the page actually works. What the loop is missing is a counterparty the VLM cannot fool, and the browser already is that counterparty: a deterministic, executable simulator of how an HTML artifact behaves under user actions, and in everything but name a world model for web code. We present WebWorld, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision. Each round, the VLM emits a critique that the planner compiles into a typed interaction contract; the browser re-executes the candidate and issues an acceptance certificate only when both target progress and preservation of every previously verified capability hold; certified transitions accumulate as a quality ratchet that is the only thing the SFT export ever sees. Under matched training, WebWorld-27B improves Raw-27B by 5.3 points on HTMLBench-400 and 14.9 points on MiniAppBench-Val, and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation. Equal-size ablations show that browser-backed admission carries the gain: without the certificate, the matched 9B lift nearly disappears.

5
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many observed forms of stereotypical bias. A key open question in this area is whether these associations are stable or change when visual representations of people in professional roles are placed in different prompted contexts. We introduce ContextBias, a controlled evaluation framework, and ContextBench, a benchmark spanning 92 roles and 1,656 semantically controlled prompts, designed to isolate the effect of contextual variation on role-linked visual representations. Evaluating four state-of-the-art models on 66,240 generated images, we find that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI +0.047). Demographic cues, characteristic garments, and role-specific tools remain highly prevalent across context-free, related, and unrelated conditions, and are robust to semantic prompt reformulation. Scene composition and camera framing show the greatest context-sensitivity. These findings reveal a form of stereotypical persistence that remains largely invisible to context-free evaluations, highlighting the need for controlled contextual variation in bias benchmarking. Code and dataset: https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina-Emami/ContextBias

4
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.

4
Verification-Aware Training for Speculative Decoding

Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, which are verified by the target model in a single forward pass. Verification proceeds sequentially and discards every position from the first rejection onward, yet existing draft training relies on token-level imitation of the target with a fixed per-position weighting that reflects neither property. We introduce Verification-Aware Training (VAT), a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision. VAT consists of two components: (i) a verification head, a lightweight jointly trained binary classifier that supervises the draft model on whether each position survives sequential verification; (ii) verification-adaptive weighting, which replaces the fixed weighting schedule by keeping full weight up to each sample's first rejection point and re-anchoring the decay to start there. VAT modifies only the training objective, so it can be layered on top of existing methods without changing the draft architecture, the target model, or the inference procedure. Applied to EAGLE-3 and DFlash on Qwen3-4B, Qwen3-8B, and LLaMA-3.1-8B, VAT improves average acceptance length by up to 11.4% and wall-clock speedup by up to 8.7%, with consistent gains across math, code, and chat benchmarks. Code will be available at https://github.com/naver-ai/vat

4
Cross-lingual Functional Vectors for Emotion Detection in Large Language Models

Function vectors (FVs) have recently emerged as a promising mechanism for steering the behavior of large language models (LLMs) by injecting task-specific latent direction representations derived from in-context demonstrations. While prior studies have shown that FVs can recover task behavior in structured in-context learning settings, their effectiveness on semantically complex tasks and their ability to generalize across languages remain underexplored. We investigate the cross-lingual transferability of FVs using multilingual multi-label emotion recognition as a challenging semantic classification benchmark. Specifically, we examine whether FVs extracted from a source language can steer task behavior in another language under both standard clean and perturbed zero-shot settings without providing demonstrations during inference. Across diverse cross-lingual settings, applying FVs substantially improves performance, suggesting that FVs capture language-agnostic, task-relevant signals rather than purely language-specific lexical patterns, and highlighting their potential as a lightweight and transferable mechanism for multilingual task adaptation. We observe that each LLM exhibits a relatively stable optimal range of attention heads for constructing effective FVs, and the pattern remains consistent across languages. In addition, FVs can partially replicate the task-steering effects of standard few-shot in-context learning while avoiding the computational overhead of processing multiple demonstrations, making them effective for large-scale practical applications. Our code is available at https://github.com/yingjie7/cross_lingual_fvs.

3
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist to process visual inputs, the community lacks a massive, multi-step, self-corrected dataset to teach models how to build and maintain internal visual workspaces when solving purely textual reasoning problems. To address this limitation, we introduce CoVA-SFT, a highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps across 5 distinct layout families and 17 complex tasks, and CoVA-Bench, a companion benchmark of 1,700 held-out test samples spanning the same tasks for reproducible evaluation. By providing explicit rationale formulations, agentic renderings, and verification loops, CoVA-SFT teaches multimodal language models to interleave text and visual abstractions. We validate the dataset by demonstrating that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, though they still fall short of strong text-only CoT baselines, highlighting open challenges for future work.

2
SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger. On the SHROOM-Visions English evaluation split, our ensemble achieves a Pearson calibration correlation of 0.41 and an overall IoU of 0.39, with a clean-response IoU of 0.91} and overall detection accuracy of 70.7%. We make our model weights and code publicly available.

2
EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state. We execute generated artifacts in a browser and evaluate them using screenshots, source and DOM evidence, actor traces, and runtime logs. Beyond turn-level and episode-level success, we measure cross-turn retention with Adjacent Pass Retention. Across eight models, even the strongest achieves 74.9% Turn Pass while completing only 37.3% of five-turn episodes; APR further falls to 52.4% on tool-grounded tasks. Diagnostic analysis shows that presentation failures center on information architecture, interaction failures on derived-state propagation and affordance binding, and tool-grounded failures additionally involve external-state grounding and requirement decomposition. These results reframe generative UI evaluation from judging isolated outputs to testing whether interface behavior, derived state, external state, and assistant claims remain synchronized as the artifact evolves.

2
Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still have certain shortcomings, hindering their further development as interactive design tools. Such schemes typically adhere to fixed input patterns, limiting flexibility in text input. Furthermore, their editing capabilities are constrained by a single or a few 2D visual models and require intricate pipeline design to integrate these models into 3D reconstruction processes. To address the aforementioned issues, we propose the Hash-Atlas network, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes. Building on this foundation, we introduce a dialogue-based 3D scene editing approach, termed CE3D++, which is centered on a large language model (LLM) that allows arbitrary textual input from users and interprets their intentions, subsequently facilitating the autonomous invocation of the corresponding visual models. Additionally, we extend CE3D++ to monocular 4D scenes by imposing motion constraints on moving objects and further fine-tuning the LLM by creating a trajectory dataset related to editing tasks, which enables the smaller LLM to schedule up to 30 different visual tools accurately. Experimental results demonstrate that CE3D++ effectively integrates multiple visual models to achieve diverse visual editing effects, possessing strong scene comprehension and multi-round dialog capabilities. The source codes and trained models are available at https://github.com/Fangkang515/CE3D.

2
DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabilities of Vision-Language Models (VLMs). However, identifying optimal subsets under a fixed ratio constraint from rapidly expanding datasets remains a significant bottleneck. While existing methods largely depend on distribution diversity or heuristic filtering, they often overlook the internal coherence within individual samples. To bridge this gap, we propose Data Intrinsic Consistency (DIC), a self-scoring metric designed to quantify the sample-level inter-component consistency. DIC consists of two modules: Visual Information Consistency (VIC), evaluating the alignment between visual content and instructions, and Response Information Consistency (RIC), assessing response coherence relative to the instruction. Building upon DIC, we introduce Data Intrinsic Consistency Selection (DICS), an adaptive data selection method that optimizes the trade-off between high intra-sample consistency and global distributional diversity under varying data budgets. Extensive experiments demonstrate that DICS consistently outperforms state-of-the-art methods across diverse dataset scales and model architectures, surpassing full-dataset fine-tuning while using only 25% of the LLaVA-1.5-665K data. We further curate DICS-6M, a 6M-sample multi-modal instruction corpus that enables the largest-scale visual instruction selection study to date; remarkably, DICS reaches 94.52\% of the official InternVL3-8B-Instruct performance using less than 25\% of its reported training data. Code can be seen at https://github.com/cqu-student/DICS

2
Dynamic Important Example Mining for Reinforcement Finetuning

Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data-centric RFT methods rely on static or heuristic sample selection, implicitly assuming a sample's value is fixed over training. This overlooks the non-stationary dynamics of policy learning and can lead to suboptimal updates. We propose Dynamic Important Example Mining (DIEM), a principled and fully automated framework that makes data utilization adaptive throughout RFT. DIEM integrates two components into each optimization step: (i) a gradient-alignment importance estimator that efficiently approximates each sample's marginal contribution to policy improvement; and (ii) a constrained batch reweighting scheme that maximizes aggregate utility while preserving the update's gradient magnitude to stabilize optimization. Across several reasoning benchmarks, DIEM consistently outperforms strong static and dynamic baselines. The code will be released via https://github.com/hrtan/DIEM.

2
BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.

2
MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation

Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform human-validation. Our novel approach leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild, e.g., that AI-generated content is particularly prevalent in technology and science, while vaccination misinformation disproportionately utilises images from news outlets to assert credibility. Our method and findings provide guidance for targeted approaches for detecting multimodal misinformation, and suggest that mitigation efforts should be developed and applied strategically rather than uniformly.

1
Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions

End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aardvark Weather model probabilistic by attaching one stochastic mechanism to each component: learned, input-dependent noise at the observation encoder, capturing aleatoric uncertainty inherited from the observing system, and Monte Carlo dropout in the processor, capturing epistemic uncertainty in the learned dynamics. The resulting nested ensemble attributes forecast spread to the two sources through a law-of-total-variance decomposition, cross-checked by withholding observation streams. Probabilistic finetuning significantly improves the mean forecast, by 4.2% on average across variables and lead times. The ensemble is calibrated against ERA5 through the medium range (spread-skill ratio 0.98), keeps station RMSE within 2.4% of the deterministic model while beating it in CRPS at every lead time, and trails the operational ECMWF ensemble. The encoder branch behaves as observation-driven uncertainty. Component-attributed uncertainty makes end-to-end forecasts more transparent, a step toward observation-driven digital twins of the atmosphere.

1
RECAP-Forcing: Retaining Content Appearances for Long Video Generation

Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

0
05

PRODUCT HUNT

05.00
PRODUCT HUNT

Product Hunt - September 2, 2026

Product Hunt Daily Feed: Featuring noteworthy tech launches.

TrustedRouter icon
TrustedRouter

Every model with a unified interface. Privacy with proof.

0
EAS Observe icon
EAS Observe

Performance monitoring built for Expo and React Native

0
HONOR Robot Phone icon
HONOR Robot Phone

The phone that literally has a gimbal built in

0
Keiki icon
Keiki

Build one customer-facing AI agent and launch it everywhere

0
Nodeterm icon
Nodeterm

A node-based free open source terminal manager

0
Naseem icon
Naseem

A native AI agent that does real work on your Mac

0
Creatium Coach icon
Creatium Coach

Your multimedia mentor that takes you from mid to great

0
Gauth AI Course icon
Gauth AI Course

AI courses you can watch, quiz through, and create

0
BobVault for BobCLI icon
BobVault for BobCLI

CLI based Zero-Knowledge Architecture for code repositories

0
Murmell icon
Murmell

Google docs for AI agents, and you can close your laptop

0
Computable GPU Index (CGI) icon
Computable GPU Index (CGI)

The first open-source price index for GPU compute

0
Sider Code icon
Sider Code

Reshape any website with plain words via Sider extension

0
ARC-24 Multitrack Groovebox for iOS icon
ARC-24 Multitrack Groovebox for iOS

Multitrack synth, sampler, drum machine, and looper for iOS.

0
Sourclip 2.0 icon
Sourclip 2.0

The research workspace built around Gemini Notebook

0
Cosmic Agent Plugins icon
Cosmic Agent Plugins

Connect Cosmic agents to any service with an MCP server

0
Happy Shrimp icon
Happy Shrimp

Alibaba's AI music generator for turning ideas into songs

0
Kilo Code for JetBrains icon
Kilo Code for JetBrains

Fully native, open-source coding agent built for JetBrains

0
Tovel AI icon
Tovel AI

From conversation to action, in three steps

0
Folio icon
Folio

A read-later app sending a typeset digest to your e-reader

0
ChannelOS icon
ChannelOS

Turn your local media library into cable-style TV

0
WaseiGo icon
WaseiGo

Learn the 1,000+ Japanese words that only look like English

0
nOS4 icon
nOS4

A complete IOS4 experience in your browser!

0
ThunderPhone icon
ThunderPhone

Platform for building reliable AI phone agents (from 2c/min)

0
Notchling icon
Notchling

A little creature that lives in your MacBook notch

0
WebTerm Learn icon
WebTerm Learn

Learn the terminal like a game — in a browser sandbox

0
Orato icon
Orato

Practice speaking with AI.

0
BrandMyLaptop icon
BrandMyLaptop

Sell ad space on your laptop

0
BrandJet icon
BrandJet

Turn public buying signals into sales pipeline

0
FrameOS icon
FrameOS

Record your iOS & Android screen from your Mac.

0
Video Agent by Fotor icon
Video Agent by Fotor

Create and edit precision motion graphics & video with chat

0
Interactive Sessions by Revolte icon
Interactive Sessions by Revolte

Drive the full SDLC with AI agents, step by step

0
Tether icon
Tether

A ball for boring meetings to keep you busy

0
Ask My Wardrobe icon
Ask My Wardrobe

The complete digital wardrobe experience

0
EP–2350 FX–MIC icon
EP–2350 FX–MIC

The programmable mic you can squeeze, shake & play

0
Radar by Particle icon
Radar by Particle

The Podcast Search Engine

0
StackScope icon
StackScope

See what new sites are built with, the week they launch

0
Edge Drop icon
Edge Drop

Clipboard on your screen edge. Hover to open, drag to drop

0
Caplio icon
Caplio

Find, organize, and reuse every image on your Mac

0
Olostep icon
Olostep

Turn the Web into Clean Data for AI

0
Prequel icon
Prequel

Create cinematic screen recordings on your Mac

0
Referent icon
Referent

The AI-native OS for modern law firms

0
Murfy AI icon
Murfy AI

Write, review, and publish to arXiv 10x faster

0
Ulpaso icon
Ulpaso

Stop paying just to take meeting notes

0
Ravioli icon
Ravioli

Create custom stamp shapes

0
oMLX icon
oMLX

Mac LLM server that cuts agent wait times from 90s to 5s

0
RIP MY BUILD icon
RIP MY BUILD

Give your abandoned side project one last launch

0
Skud icon
Skud

Menubar file delivery with your brand and tracking

0
Hyperfocus icon
Hyperfocus

Planner that turns goals into daily progress

0
Topview Motion Studio icon
Topview Motion Studio

Create launch videos without touching After Effects

0
Retro Y2K Theme icon
Retro Y2K Theme

Customize any website with a vintage 90s & Y2K retro theme

0
06

TECHMEME

06.00
TECHMEME

Techmeme - September 2, 2026

Techmeme Digest: Major tech headlines and industry conversations.

Google rolls out its September Android Drop, with remembered items in Find Hub, Guided vision in Gemini Live, Motion Assist to reduce motion sickness, and more (Ryan Whitwam/Ars Technica)
Source: TechmemePublished: Sep 1, 2026

Ryan Whitwam / Ars Technica : Google rolls out its September Android Drop, with remembered items in Find Hub, Guided vision in Gemini Live, Motion Assist to reduce motion sickness, and more —  Let's be frank: Some of Google's recent Android feature Drops have been duds, offering little more than an expansion of Gemini summaries and chat functions.

Filing: Apple says John Ternus will get a compensation package worth ~$58M in FY 2027, while Tim Cook's role as executive chairman will pay ~$47M (Mark Gurman/Bloomberg)
Source: TechmemePublished: Sep 1, 2026

Mark Gurman / Bloomberg : Filing: Apple says John Ternus will get a compensation package worth ~$58M in FY 2027, while Tim Cook's role as executive chairman will pay ~$47M —  Apple Inc. said that new Chief Executive Officer John Ternus will get a compensation package worth about $58 million in fiscal 2027 …

Sources: Google plans to release Gemini 3.8 Flash as soon as Wednesday; Gemini 4 has done well on pre-training evals but still needs to complete post-training (Erin Woo/Wall Street Journal)
Source: TechmemePublished: Sep 1, 2026

Erin Woo / Wall Street Journal : Sources: Google plans to release Gemini 3.8 Flash as soon as Wednesday; Gemini 4 has done well on pre-training evals but still needs to complete post-training —  Internal tests of Gemini 3.8 Flash show progress in an area where the company has lagged behind Anthropic and OpenAI.

The US urged G20 members to avoid writing entirely new AI regulations, and instead focus on writing rules for "novel" situations that involve the tech (Reuters)
Source: TechmemePublished: Sep 1, 2026

Reuters : The US urged G20 members to avoid writing entirely new AI regulations, and instead focus on writing rules for “novel” situations that involve the tech —  The U.S. pressed G20 members on Tuesday to take a hands-off approach to AI regulation and avoid creating new rules …

Palo Alto Networks reports Q4 revenue up 34% YoY to $3.41B, vs. $3.35B est., and acquires Console, which provides AI-powered IT service management (Samantha Subin/CNBC)
Source: TechmemePublished: Sep 1, 2026

Samantha Subin / CNBC : Palo Alto Networks reports Q4 revenue up 34% YoY to $3.41B, vs. $3.35B est., and acquires Console, which provides AI-powered IT service management —  Palo Alto Networks surpassed fiscal fourth-quarter estimates as mounting artificial intelligence risks boost demand for its cybersecurity tools.

Claude Fable 5.1 and Mythos 5.1 are Anthropic's first models to watermark text outputs; a detection API is available to eligible groups as required under EU law (Ben Patterson/PCWorld)
Source: TechmemePublished: Sep 1, 2026

Ben Patterson / PCWorld : Claude Fable 5.1 and Mythos 5.1 are Anthropic's first models to watermark text outputs; a detection API is available to eligible groups as required under EU law —  Anthropic has announced the arrival of its latest Claude models, and just as it promised last month, the new models will add invisible watermarks to all their text replies.

Apple updates Apple Maps to show Lake America for US users, Lake Ontario for Canadian users, and both for the rest of the world, following Trump's EO (Mark Gurman/Bloomberg)
Source: TechmemePublished: Sep 1, 2026

Mark Gurman / Bloomberg : Apple updates Apple Maps to show Lake America for US users, Lake Ontario for Canadian users, and both for the rest of the world, following Trump's EO —  Apple Inc. renamed Lake Ontario “Lake America” on its maps service, following a similar move by Alphabet Inc.'s Google, in the wake of an order from US President Donald Trump.

Memo: Alexandr Wang says Meta is switching from Google Chat to Slack for internal communications as Slack is the "strongest platform available today for agents" (Business Insider)
Source: TechmemePublished: Sep 1, 2026

Business Insider : Memo: Alexandr Wang says Meta is switching from Google Chat to Slack for internal communications as Slack is the “strongest platform available today for agents” —  - Meta is migrating to Slack for internal communications, a memo from AI chief Alexandr Wang says.

Dell reports Q2 revenue up 58% YoY to $46.97B, vs. $44.92B est., and forecasts $192B in FY 2027 revenue, vs. $172.67B est.; DELL jumps 9%+ after hours (Jordan Novet/CNBC)
Source: TechmemePublished: Sep 1, 2026

Jordan Novet / CNBC : Dell reports Q2 revenue up 58% YoY to $46.97B, vs. $44.92B est., and forecasts $192B in FY 2027 revenue, vs. $172.67B est.; DELL jumps 9%+ after hours —  Dell Technologies shares moved 9% higher in extended trading on Tuesday after the computer maker reported results and a forecast that easily cleared Wall Street expectations.

OpenAI says Astra is its first model to reach its "Critical" cyber threshold and warns safeguards may mistakenly flag legitimate activity as cyber misuse (Ina Fried/Axios)
Source: TechmemePublished: Sep 1, 2026

Ina Fried / Axios : OpenAI says Astra is its first model to reach its “Critical” cyber threshold and warns safeguards may mistakenly flag legitimate activity as cyber misuse —  OpenAI said Tuesday that it plans to release its latest model — Astra — soon, but its most advanced cybersecurity features …

OpenAI says it plans to publicly release a version of Astra "soon" but will make its advanced cyber capabilities available only to select partners (Wired)
Source: TechmemePublished: Sep 1, 2026

Wired : OpenAI says it plans to publicly release a version of Astra “soon” but will make its advanced cyber capabilities available only to select partners —  The company will give select partners early access to its Astra AI model—so they have time to shore up their defenses.

Sources: AfterQuery, which sells coding and finance training data to AI labs, has hit a valuation of $3.2B, up from $300M in April, and is profitable (Anna Tong/Forbes)
Source: TechmemePublished: Sep 1, 2026

Anna Tong / Forbes : Sources: AfterQuery, which sells coding and finance training data to AI labs, has hit a valuation of $3.2B, up from $300M in April, and is profitable —  Twenty three year-old Spencer Mateega pivoted his YC startup into the fastest unicorn in the accelerator's history, fueling …

Meta launches Muse Voice Transcribe, MSL's first real-time audio perception model, with streaming automatic speech recognition, trained with 70+ languages (Meta AI Research)
Source: TechmemePublished: Sep 1, 2026

Meta AI Research : Meta launches Muse Voice Transcribe, MSL's first real-time audio perception model, with streaming automatic speech recognition, trained with 70+ languages —  Experience Muse Voice Transcribe in real time … We're excited to introduce Muse Voice Transcribe, the first real …

Wafer, which makes AI agents that optimize open-source models for a business's workload, raised a $40M Series A, a source says at a $200M+ valuation (Stephanie Palazzolo/The Information)
Source: TechmemePublished: Sep 1, 2026

Stephanie Palazzolo / The Information : Wafer, which makes AI agents that optimize open-source models for a business's workload, raised a $40M Series A, a source says at a $200M+ valuation —  Developers and chip designers are increasingly upbeat about the benefits of using AI to supercharge the process of designing and optimizing AI chips …

Fei-Fei Li's World Labs unveils Atlas, a multimodal world model that generates image/video frames with pixel-perfect camera control and reconstructs them in 3D (World Labs)
Source: TechmemePublished: Sep 1, 2026

World Labs : Fei-Fei Li's World Labs unveils Atlas, a multimodal world model that generates image/video frames with pixel-perfect camera control and reconstructs them in 3D —  World models generate, reconstruct, and simulate any possible world.  They understand how worlds appear, behave …

07

STARTUP ARCHIVE

07.00
STARTUP ARCHIVE

Startup News - September 2, 2026

Startup News Roundup: Aggregating key funding and launch updates.

Marc Andreessen on the 5 personality traits of an innovator
Source: StartupPublished: Mar 31, 2026

“When you’re talking about real innovators—people who actually do really creative, breakthrough work—I think you’re talking about a couple things:”

Steve Jobs explains the importance of both thinking and doing
Source: StartupPublished: Mar 30, 2026

“The doers are the major thinkers. The people who really create the things that change this industry are both the thinker-doer in one person.”

Tobi Lutke explains what the VCs who passed on Shopify got wrong
Source: StartupPublished: Mar 27, 2026

“What a lot of free-market thinkers don’t understand is that between the demand and eventual supply lies friction."

Sam Altman explains how he decides to invest in a startup after 10 minutes
Source: StartupPublished: Mar 26, 2026

"Does this person have the potential to be the next Mark Zuckerberg?… [You don’t get to] 100% accuracy, obviously, but it’s good enough that our business model works.”

Jony Ive recounts the time Steve Jobs called him vain
Source: StartupPublished: Mar 25, 2026

In the clip below, Jony Ive recounts the time he asked Steve Jobs to be less harsh in his critique of a piece of work.

Jeff Bezos’s two pieces of advice for aspiring entrepreneurs
Source: StartupPublished: Mar 24, 2026

“The advice that I would give entrepreneurs is don't chase the hot new thing. It's so hard to catch something that everybody already knows is hot."

Elad Gil: “Things that work tend to work pretty fast”
Source: StartupPublished: Mar 23, 2026

“I do think there’s a bit of a myth in Silicon Valley that you should keep grinding no matter what and it’s just about perseverance, and I think that’s really bad advice."

Paul Graham on why starting with a “small, intense fire" is the key to startup growth
Source: StartupPublished: Mar 20, 2026

"You have to know who those first users are and how you're going to get them."

Keith Rabois on how to identify great talent
Source: StartupPublished: Mar 19, 2026

“What you want to do with every single employee every single day is expand the scope of their responsibilities until it breaks… and that’s the role they should stay in.”

Wealthfront CEO on why advertising spend makes it harder to find product/market fit
Source: StartupPublished: Mar 18, 2026

“The way that you know you have product/market fit is if you have exponential organic growth."

Eric Schmidt on why most companies get strategy wrong
Source: StartupPublished: Mar 17, 2026

“Work very, very hard to figure out what the world’s going to look like in five years. What will people be doing? What will your customers want? Where will costs be?"

Mark Zuckerberg: “You can’t 80/20 everything”
Source: StartupPublished: Mar 16, 2026

"There’s the famous 80/20 rule where you get 80% of the benefit by doing 20% of the work, but you can’t just 80/20 everything. There have to be certain things that you are just the best at."

Marc Andreessen on Mark Zuckerberg’s founder “superpower”
Source: StartupPublished: Mar 13, 2026

“A great superpower that Mark Zuckerberg has that is probably not well-understood enough is he does not get emotionally upset in stressful situations"

Sam Altman explains how to come up with a great startup idea
Source: StartupPublished: Mar 12, 2026

"If you start a startup without a good idea… you’ll be under pressure to make something up and it won’t work that well."

Jeff Bezos on the problems with proxies and managing to metrics
Source: StartupPublished: Mar 11, 2026

“One of the things that happens in business is that you develop certain things that you’re managing to—a typical case would be a metric. And that metric isn’t the real underlying thing.”

Airbnb founder Brian Chesky on how to design an amazing user experience
Source: StartupPublished: Mar 10, 2026

“If you can design something really amazing using the hand-crafted part of your brain, then you can reverse-engineer how to industrialize this millions of times over."

Spencer Rascoff: "I will never invest in a consumer startup with paid marketing”
Source: StartupPublished: Mar 9, 2026

"If you’re actually trying to grow a product, the best levers for doing that are often within the product itself.”

Patrick Collison explains why it sometimes make sense to quit
Source: StartupPublished: Mar 6, 2026

“One thing I’ve learned myself the hard way, is that it is easier to tear down a company and restart it in Silicon Valley, than it is to constantly try to pivot or keep something alive."

Jeff Bezos recounts the time he called Amazon’s customer service number mid-meeting to prove a metric was wrong
Source: StartupPublished: Mar 5, 2026

“I have a saying, which is when the data and the anecdotes disagree, the anecdotes are usually right"

Ben Horowitz: “Nobody was born a great manager. It’s a very unnatural job.”
Source: StartupPublished: Mar 4, 2026

“If you can’t build a great product, it doesn’t matter if you can build a great company.”

03

ALSO TODAY

3 MORE SOURCES
08

SOLIDOT

08.00
SOLIDOT

Solidot News - September 2, 2026

Solidot Feed: Highlighting essential tech & open-source news.

小规模民调显示七成韩国民众支持限制青少年使用社交网络

周二公布的一项民调显示,七成韩国民众表示支持出台限制青少年使用社媒的政策。这项民意调查访问了 1000 名年龄在 14-58 岁之间的受访者。调查结果显示,70.7% 的受访者支持,29.3% 的受访者反对。占总调查人数五分之一的青少年受访者中,59% 反对,41% 支持。当被问及实施此类限制的合适年龄时,16.8% 选择了 15 岁,14.9% 选择了 18 岁,13.1% 选择了 11 岁及以下。大多数受访者表示,即使出台此类政策,在限制青少年使用社媒方面仍然存在局限性,青少年用户可能会盗用他人账号或转向其它不受限制的平台,因此 59.8% 的受访者认为,平台应采取更多措施营造安全的社媒使用环境。

Softaculous 遭遇长达 33 小时的 BGP 路由劫持

8 月 28 日 20:57 UTC 左右,一个不相关网络 BGP 路由通告了 Softaculous 使用的 Hetzner IP 段,将部分原本发送到 Softaculous 系统的流量劫持到攻击者控制的服务器。Hetzner 是 Softaculous 的上游基础设施供应商,而 Softaculous 则是一家为 Web 托管服务商提供软件的公司,它的 Virtualizor 控制面板被管理员用于部署和管理 VPS。这次 BGP 路由劫持影响了 Virtualizo 更新服务器以及客户和计费网站。攻击者还从 Let's Encrypt CA 获取了有效的 TLS 证书,Let's Encrypt 的自动域名所有权验证也被劫持到了攻击者控制的 IP。Softaculous 于 8 月 29 日 08:50 UTC 向 Hetzner 报告了事件,Hetzner 随后通过发布相同的路由通告遏制了问题。但攻击者于 20:00 UTC 再次了长达 10 小时的路由劫持。8 月 30 日 05:50-06:10 UTC 路由通告被撤回,劫持停止。Softaculous 建议在攻击期间登陆过的用户立即重置密码,以及重置所有重用该密码的账户。同一时间段内输入过银行卡信息的客户也应检查其账单。攻击者在此期间推送了一个恶意的 Virtualizor 更新包,它建议所有 Virtualizor 用户检查其服务器并轮换凭证。

科学家定位调控冬眠的关键脑回路

为弄清动物进入冬眠时大脑的变化,研究人员首先在实验室中诱导叙利亚仓鼠冬眠。两个月里,他们把动物笼舍中开灯的时间缩短,以模拟秋季。然后在接下来的两个月中,将温度降至约4摄氏度,以模拟冬季。 在人造冬季中,动物开始冬眠——体温下降,“在窝里缩成一团”。冬眠持续2到8周,其间仓鼠睡眠状态在微觉醒和深度蛰伏之间循环。团队收集了刚进入深度蛰伏的仓鼠的大脑,并将其与刚从蛰伏中短暂醒来或完全未冬眠的仓鼠的大脑进行比较。结果发现,一种名为Fos的蛋白质高水平表达,表明下丘脑视前区存在活动。该区域参与调节体温、睡眠和其他重要功能。研究人员发现,仓鼠冬眠中活跃的POA神经元特定亚群,似乎与小鼠蛰伏状态中鉴定出的神经元相同。抑制这些神经元会延迟仓鼠重新进入蛰伏,而激活这些神经元则引发仓鼠筑巢行为并使其体温下降。这是体温虽不如自然冬眠时那么低,但也远低于平时,仅13摄氏度。在小鼠中激活这些细胞也能使体温降低,但幅度较小。这组POA细胞可能是进化遗留下来的“关闭键”,使早期温血哺乳动物能够降低维持体温的高能量成本。

ChatGPT 和 Reddit 被要求遵守欧盟的 DSA

欧盟委员会周一表示,OpenAI 的 ChatGPT 将需要遵守更严格的欧盟法规,否则将面临罚款。聊天机器人 ChatGPT、社媒论坛 Reddit 和游戏平台 Roblox 被欧盟网络安全法规《Digital Services Act(DSA)》归类为“超大型在线平台”。该认定意味着这些服务面临额外的义务,如删除非法内容、保护未成年人的隐私和安全,如果未能遵守规定,将面临最高全球收入 6% 的罚款。欧盟的决定标志着 DSA 的适用范围进一步扩大到生成式 AI 领域。此前 X 的 AI 聊天机器人 Grok 已因违反 DSA 而受到调查。这三大服务在欧盟的月活用户数都已超过 4500 万,达到了 DSA 规定的加强审查门槛。它们需要在 12 月底前履行额外义务。

太阳风暴导致美国 GPS 信号偏差逾 10 米

2025 年 11 月太阳释放了多个 X 级耀斑,耀斑还伴随着引发地磁风暴的日冕物质抛射。地球上的居民在此期间目睹了绚丽的极光,极光的范围甚至延伸至低纬度地区。对太阳风暴期间收集的数据的分析发现,美国上空的大气层出现了大范围的、横跨东西海岸的扰动,其规模前所未见。它导致部分地区的 GPS 定位偏差超过 10 米。如此大的偏差足以影响精准农业和自动驾驶汽车的运作。GPS 信号穿过电离层时会被扭曲和衍射,导致抵达接收器时的信号强度快速波动,这种现象被称为振幅闪烁(amplitude scintillation)。闪烁并不罕见,通常发生在两极和赤道,中纬度地区被认为相对安全。但去年底的太阳超级风暴改变了这一切。美国大陆西经 80-120 度之间的大片区域出现了强振幅闪烁。如此大范围的强振幅闪烁以前从未看到过。

智神星一号成功完成首次演示飞行

民营商业航天公司星河动力于 9 月 1 日 10 时在酒泉东风商业航天创新试验区智神星系列专用发射工位成功发射了其中型火箭智神星一号。智神星一号是基于已投入使用的小型火箭谷神星一号,为两级构型,全长 52 米,芯级直径 3.35 米。智神星一号类似 Falcon 9,使用煤油和液氧作为推进剂,也采用类似的方式回收方式,第一次飞行没有尝试回收。火箭能将 5 吨重的有效载荷送入 400 公里高的近地轨道,或将 3 吨的有效载荷送入 700 公里高的太阳同步轨道,其有效载荷小于 Falcon 9。智神星一号设计回收使用次数不少于 25 次。星河动力正在智神星一号基础上研发重型版本,类似 Falcon Heavy,设计能将 17.5 吨重的有效载荷送入近地轨道。

中国光伏装机容量首次超过煤电

中国国家能源局星期二公布的最新数据显示,截至今年 7 月底,全国光伏发电装机容量达到 12.86 亿千瓦,略高于 12.85 亿千瓦的煤电装机容量。目前光伏占全国发电总装机容量的31.5%。今年前七个月,全国光伏发电量超过 8024 亿千瓦时,同比增长 15.5%,相当于全国每八度电中约有一度来自光伏。能源局预计,未来五年中国光伏产业投资将超过 2 万亿元人民币。

Paint.NET 实验性支持 Wine/Linux

Windows 上的流行图像编辑软件 Paint.NET 释出了 v5.2 Alpha (build 9739),实验性的加入了对 Wine/Linux 的支持。想要尝试的 Linux 用户需要:1)使用便捷式版本,安装程序无法工作;WINE 版本至少为 v11.14;添加一个注册表项 wine reg add "HKCU\Software\Wine\DllOverrides" /v d3dcompiler_47 /d native /f,防止 Wine 用自身版本覆盖 DLL;已安装 DXVK; 用 /wine 命令行参数运行。

带电雨滴会腐蚀汽车

德国科学家研究显示,作为一种自然发生的现象,带电雨滴能分解经过保护性涂层处理的表面。这一此前被忽视的腐蚀机制有助于寻找新方法保护如汽车、船舶、建筑和文化遗址等金属物体。腐蚀是户外金属制品和结构面临的一个经济和安全问题。之前认为,雨水造成的腐蚀主要是因为雨滴具有物理腐蚀性或带有来自污染物的酸性。虽然既往研究已证明水滴在滑过植物叶子和窗玻璃这类常见表面时会带电,但这种电荷对金属表面的腐蚀作用却被忽视了。马普学会高分子研究所的研究人员分析了电中性水滴如何带电并腐蚀不同材料。他们将水滴到四种常见表面——植物叶子、PVC(聚氯乙烯)泡沫板、聚苯乙烯玻璃,和常用疏水涂层PFOTS(全氟辛基三乙氧基硅烷)。这些水滴随后会滑到有聚四氟乙烯涂层的铜上。研究者发现,水滴带上了很微弱的电荷(0.2-2纳库仑)。在3000个水滴从这些表面滴到有聚四氟乙烯涂层的铜上后,观察发现这一涂层会分解,导致底下的金属被腐蚀。如果水滴不带电,就不会观察到表面损伤。

拒绝改名的 MapQuest 成为美国 App Store 下载量最高的地图导航应用

在 Google Maps 以极快的速度屈服特朗普的行政令,将安大略湖(Lake Ontario)更名为美国湖(Lake America)之后,公开宣布拒绝改名的美国地图导航服务 MapQuest 立即赢得了公众的支持,其移动应用下载量飙升,排在苹果美国 App Store 非游戏类榜单第四,总排名第八,成为美国排名第一的地图导航应用。MapQuest 也成为苹果加拿大 App Store 下载量总排名第二的应用。该公司表示,自决定保留安大略湖名称以来,其应用的下载量增加了数十万次,使用量更是飙升至正常水平的 50 倍。Sensor Tower 的第三方数据也显示,8 月 24 日开始的一周内 MapQuest 全球和美国应用总下载量分别为 18.4 万次和 16.2 万次,均较前一周增长逾 10 倍。在美国 8 月 26-30 日间 MapQuest 应用的日下载量环比增长了 128%。

Google 从其扩展商店移除了包括 uBlock Origin 在内的所有 Manifest V2 扩展

Google 从其扩展商店 Chrome Web Store 移除了包括 uBlock Origin 在内的所有 Manifest V2 扩展。Chrome 用户已安装的 Manifest V2 扩展目前还能使用,但不会再收到任何更新,卸载之后将无法再重新安装。Chrome Web Store 将只提供 Manifest v3 扩展,Manifest v3 相比 Manifest v2 的一大区别是移除了 blockingWebRequest 和 declarativeNetRequest,限制了广告屏蔽扩展的功能,uBlock Origin 是基于 Manifest V2,它没有 Manifest V3 版本,但有一个基于 V3 的精简版本。用户如果想要继续使用 uBlock Origin 只能改用 Firefox 以及另一款基于 Chromium 的浏览器 Brave。

OpenShot 4.0 释出

自由软件视频编辑器项目 OpenShot 释出了 v4.0 版本。主要新特性包括:新色彩视图;新录制视图:将麦克风、屏幕、Web 摄像头和系统音频直接添加到项目中,每个音源保持独立且可编辑;10 种新特效;使用本地大模型选择和跟踪对象;更简洁的原生时间线;更快的特效和编辑速度;智能的创意工作流程;扩展 Qt 6 支持,改进了与较新 Linux 发行版的兼容性,为 Android 和其它平台奠定了基础。

加州议会通过年龄验证法案,Linux BSD 豁免

加州参议院和众议院批准了年龄验证法案 Assembly Bill 1856。在递交给州长批准之后法案预计于 2027 年 1 月 1 日生效。法案豁免了 Linux 和 BSD 等开源操作系统。法案要求,如果操作系统有账户设置功能,那么系统提供商须提供一个界面,在账户设置期间要求输入设备主用户的出生日期、年龄或两者兼有。操作系统通过相对一致的实时 API 向受监管的应用商店和应用开发商提供数字年龄信号。该信号不显示精确的出生日期,而是四个年龄段之一:13 岁以下、13-15 岁、16-17 岁或 18 岁及以上。对 2027 年 1 月 1 日之前的设备,操作系统提供商必须在 2027 年 7 月 1 日之前提供界面让账户持有人提供所需的年龄信息。

Linux 7.3-rc1 释出

Linus Torvalds 宣布释出 Linux 7.3-rc1,关闭了 7.3 的合并窗口,正式版预计将在十月底释出。Linux 7.3 的主要特性包括:Ryzen AI Halo LED/RGB 驱动、继续即将推出的 AMD Zen 6 的支持工作、KSMBD 兼容 Apple Time Machine 备份、内核驱动初步支持 2026 年款 Steam Controller、合并 FailFS、Intel Xe3P Nova Lake 集显支持稳定、改进了显存容量有限的系统的游戏性能、改进 SMP 降低延迟提升实时性能、等等。

Steam 平台 2003-2013 年的几乎所有游戏泄露

上周末 Steam 平台逾 12TB 数据泄露,涵盖了该平台 2003-2013 年之间几乎所有的游戏。这些数据是通过一个公开访问的 API 获取的,但不清楚是近期访问还是早就下载但直到上周才公开。相关数据来自被称为 Steam2 的内容分发系统,2013 年 Steam2 被 SteamPipe 系统所取代,因此数据仅限于 2013 年前。泄露的数据包括了Valve 和第三方发行商发布的热门游戏的早期版本、原型版本和试玩版本,其中包括《传送门2》的被删减内容,被取消的《半条命2:第三章》的部分文件。

Google 改变了其搜索结果的展示方式

Google 过去一年对其搜索结果的展示方式进行了两次重大改变。其一是搜索结果链接,以前你将鼠标悬停在搜索结果上会在浏览器底部看到网站链接,现在显示的是 google.com/goto + 一串看起来随机的字符串。搜索结果中的 AI Overview 引用的链接也是采用此类展示方式。其二是用户以前可以在搜索词末尾添加 &num=100,可以在一个页面上显示前 100 个搜索结果,如今这一快捷方式被取消了,Google 强制只展示最多 10 个搜索结果,意味着你想要看前 100 个结果需要点击 10 次。

植物如何应对高温

科学家早就知道,植物叶片表面分布着许多微小的气孔。当温度升高时,这些微小的孔隙会“张嘴”,让水分蒸发,从而带走热量,就像人出汗能降温一样,但气孔这一植物“散热器”背后的分子调控机制,一直是个未解之谜。该通路的核心是一种名为“泛素特异性蛋白酶24”(UBP24)的蛋白质。他们发现,当高温来袭,植物体内的激酶会“唤醒”UBP24,导致其分子电荷发生变化,使其变得更加稳定。在更稳定的状态下,UBP24 有助于“保护”并激活其他参与维持气孔开口的蛋白质,使植物的蒸发冷却系统在热应激时保持活跃。通俗来说,UBP24 就像“空调”上灵敏的温控开关。类似的故事并非只在植物身上上演。科学家还发现,啤酒酵母中的一种相关蛋白质也依靠相似原理应对热应激。酵母与植物分属不同物种,生活方式大相径庭,且两者之间存在数亿年的进化差距,却在细胞层面使用了相似的“散热逻辑”。

人类何时开始不爱吃昆虫?

研究人员借助基因组证据,还原了数千年来人类食用昆虫的模式。研究表明,在欧洲、中亚与东亚地区,吃昆虫可能只是偶然行为;而在热带地区以及尼安德特人中,食虫则更为普遍。研究人员检测了 745 份来自现代人的牙结石样本,其年代最早可追溯至 3.3 万年前。牙结石能保存食物的 DNA,为研究人员提供了远古饮食的记录。研究结果显示,生活在欧亚大陆北部的现代人并不会经常食用昆虫。研究团队还检测了与分解几丁质相关的基因,几丁质是昆虫外骨骼的主要成分。在欧亚大陆北部人群中,几丁质酶基因发生了突变,导致人体消化昆虫外骨骼的能力下降。这种基因模式已延续了约 9000 年,可追溯至农业兴起之时。尼安德特人的情况则截然不同,他们的牙结石中所含的昆虫 DNA 要多得多。尼安德特人牙结石中最常见的基因痕迹来自双翅目昆虫,包含苍蝇和蚊子,其中蚊子的 DNA 含量尤其丰富。该结果佐证了近期的一个假说——尼安德特人可能经常食用带有蝇蛆的动物尸体。而蚊子遗骸的大量存在也支持了另一种观点——猎物的尸体有时可能被存放在池塘或沼泽环境中,而蚊子会在这些地方产卵。

天文学家可能发现首个没有恒星的星系

被称为 Cloud 9 的星系可能是人类发现的第一个没有恒星的“失败”星系。此类星系虽然有大量气体和尘埃,却几乎没有恒星存在。Cloud 9 距离地球 1400 万光年,靠近旋涡星系 M94,它包含一团巨大的氢气云,估计质量是太阳的 100 万倍,以及质量约为太阳 500 万倍的暗物质。它几乎不发出任何星光。无恒星星系形成的主流解释与紫外背景辐射相关,在宇宙早期的再电离时代后,紫外辐射场将低质量暗物质晕中的气体加热到足够高的温度,使得气体无法有效冷却并坍缩形成恒星。

韩国准备向所有民众提供免费 AI 服务

韩国准备向所有民众提供免费 AI 服务 AI for All。该服务计划从下月起进行 beta 测试,计划今年晚些时候推出。服务由韩国两家最大的电信公司以及 Kakao 牵头的三个联盟提供,政府提供部分算力,包括提供最多 512 块英伟达 B200 芯片和支付部分运营费用。AI 服务预计将与政府系统连接,而不只是作为独立的聊天机器人运行。居民将能使用这些服务预约医生、搜索公寓和获取税务指导。小型企业将能计算税款和查询是否符合政府补贴资格。家长将收到教育内容的推荐。每个联盟都计划推出自己的 AI 应用,在现有产品中加入生成式 AI 功能。在政府支持的计划下,用户可以无限次访问,没有 token 限制。

09

APP STORE RANK

09.00
APP STORE RANK
Loading…
TEXT VIEW · TODAY'S DIGEST · 0 HEADLINES ACROSS 8 SOURCES

Hacker News(0)

No items yet for today.

GitHub Trending(0)

No items yet for today.

Product Hunt(0)

No items yet for today.

Hugging Face(0)

No items yet for today.

Techmeme(0)

No items yet for today.

Solidot(0)

No items yet for today.

Startup Archive(0)

No items yet for today.

App Store Rankings(0)

No items yet for today.