面向医学研究者的网站 Research Gold 宣称其服务"100%由人类撰写、绝不使用AI",并列出多名博士审稿人。但调查发现,这些审稿人系AI生成、并不存在;部分真实方法学家的身份和照片未经许可被挪用。致电该公司时,自称"Sarah"的AI助手坚称自己是真人,邮件与聊天回复也均为AI生成。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspo193n05q2rojeue8j2k5a
Grok Bot 【引用 @farzyness】:这是我的 Grok Bot 团队: - Webby:网页设计师 - Shotry:短视频内容创作者 - Writey:文章/通讯稿写手 - Claude Code:专精 CC 的 Grok 智能体 - Codex:同上,但专精 CODEX - Script:YouTube 脚本写手 - Idea:视频/内容创意生成 - Master:整个团队的编排者 搭建这些并让它们跑起来非常容易。看看效果如何! 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspo40i205uiroje88gkps2x
开发者 @avgvstvs96 编写了一个带脚本的 GitHub 技能,让智能体更高效地使用 GitHub。该技能将 `pr view --json` 的 3-5 次调用压缩为一次 `pr-snapshot`,用 `pr-threads` 替代 15 行 GraphQL 查询,并让 `ci-failures` 将错误片段放入上下文、完整日志本地保存供智能体检索。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspo16gx05pkroje9f05655v
小米 MiLM Plus 团队发布 PROVE 评测框架,包含 RC-S 与 RC-T 两项感知对齐指标及 PROVE-Bench 真实世界视频基准,已获 ACM MM 2026 接收。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspn1iyf04y9rojeo97vsr7z
xAI 推出 Grok Bot 早期测试版,将 Telegram + OpenClaw 玩法做成官方托管产品。Grok Bot 是能登录用户工具、像人一样操作并交付成果的 AI 队友。作者认为,相比 Codex 等需人紧盯的工具形态,这种可离开电脑、让 AI 独立协作的任务组织方式更符合未来 AI 形态。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspn1doy04uvroje00dc3ivx
HumanLayer 的 Dex Horthy 认为,模型智能提升反而让人类用户体验变差:数学问题可逐步优化,但"像人一样说话"存在漫长的最后一公里。引用推文指出,前沿模型日益晦涩的表述并非训练失误,而是更高智能的自然表现,实验室最终会通过"屈就人类水平"来消除这种困惑感。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspo16gx05plroje5edrpw96
据 Polymarket 预测,Anthropic 或于 10 月底 IPO,预计募资超 600 亿美元,有望成为全球第二大 IPO。但投资者不满 CEO Dario Amodei 过度强调 AI 风险,劝其多谈盈利。内部测试中,Claude 模型曾因配置失误攻入真实公司系统,恶意软件包被 15 个系统下载。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspmw4yg04qgrojedkzio9eq
MeDo 3.5 正式发布,开启 AI 应用构建新纪元。8 月 13 日上午 9 点(UTC+8)直播介绍新功能,产品经理 Yuchen Xu 将详解 Build(一键生成 Web 与 App)、Launch(共享后端与 iOS 打包)及 Grow(SEO Agent)三大升级。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmsplyukj03mhrojeh1zocpw4
英伟达出资25%,联合华尔街共筹5000亿美金,向AI公司放贷用于购买GPU。此前英伟达借钱给云厂商建数据中心,若算力卖不完则按市场价回购兜底,如今拉投行共担风险。美国产业资本被深度绑定,赌技术升级能再现IT化50年繁荣,AGI成则资本受益,败则共同承压。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspmd19t04gvrojepq4y93yi
微软发布升级版编程模型 MAI-Code-1.1-Flash,在 Terminal-Bench 2.1 测试中表现提升 22%,.NET 相关任务提升 15%,Token 生成速度提升 25%,完成相同任务所需 Token 减少 25%。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspmw4yh04qiroje5vcysij3
从柔性本体走向跨本体基础智能,让任务与世界知识延续
宏碁首款 Googlebook 型号 GP714-91N 曝光,定位高于传统入门级 Chromebook,规划推出 4 款机型。该机配备 14 英寸 2.8K 16:10 OLED 屏,搭载酷睿 Ultra 5 325 或 Ultra 7 355 处理器,提供 16GB/32GB 内存及 256GB/512GB 固态硬盘,支持 Wi-Fi 7 和蓝牙 5.4。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmspmw4yh04qjroje2nljhgg8
arXiv:2608.09949v1 Announce Type: new Abstract: This study evaluates the application of Large Language Models (LLMs) in complex biological systems, evolving from data analysis to autonomous, AI-guided experimentation. The framework is driven by data from a 49-channel phytosensor network, encompassing multispectral, electrochemical, and dielectric modalities. To enhance accessibility, the system provides real-time natural-language interpretation for both specialists and non-experts. However, its core advantage lies in the transition from human-in-the-loop analysis to autonomous control. Processing biophysical data, the LLM evaluates plant physiology and triggers hardware actuators to optimize microclimates, execute phenotyping protocols, or induce controlled stress scenarios. This closed-loop architecture establishes a direct AI-biology interface, enabling data-driven exploration of complex biosystems and ecologies. The framework was validated across three case studies, based on a vert
arXiv:2608.09967v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel model-agnostic, sampling-based framework for interpreting DRL policies. Given access to the policy and an environment simulator, SPOT constructs an interpretable finite-horizon tree by sampling actions and recursively simulating the resulting successor states. The tree provides an empirical representation of the policy's action preferences and their possible downstream evolution. We provide formal guarantees establishing SPOT's asymptotic recovery of the policy's unique most probable action and characterizing its disagreement behavior under high-entropy policies. We demonstrate SPOT in the SUMO-RL traffic-signal control domain. The case study illustrates how its tree-based representation can be used to inspect policy pr
arXiv:2608.09986v1 Announce Type: new Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fail to effectively extract and leverage task-relevant information from the incomplete multimodal data. To address this challenge, we propose a unified framework termed Mutual Information Disentanglement with uncertainty-Aware fuSion (MIDAS), which effectively restructures multimodal representations under incomplete conditions. MIDAS adopts a variational modeling strategy to represent each modality with multivariate Gaussian latent variables and further decomposes them into shared and exclusive factors. To obtain reliable representations, we design a minimax objective that mini
arXiv:2608.09998v1 Announce Type: new Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to their high energy demands and associated carbon emissions. This concern is particularly relevant in light of the increasing deployment of large-scale models, especially Deep Learning (DL) architectures, which provide advanced predictive capabilities but require substantial computational resources. This paper presents a systematic review of research on Green AI, Green DL, and optimization techniques aimed at reducing the environmental impact of AI models. In addition, we examine and compare several carbon measurement tools for estimating emissions generated by AI algorithms. To complement the review, we conducted an empirical evaluation using a CPU-based experimental setup, in which six DL mo
arXiv:2608.10004v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide an interpretable framework by grounding predictions in human-understandable concepts, enabling semantic inspection and test-time intervention. Recent variants have improved CBMs through richer concept representations, uncertainty estimation, and dependency modeling. However, robust reasoning under unreliable concept states remains underexplored. Without such reasoning, misleading semantic evidence can propagate through the bottleneck, compromising both explanations and downstream predictions. To address this issue, we propose ReCBM, an uncertainty-gated relational reasoning framework for CBMs. ReCBM introduces semantically defined concept relations into the bottleneck and uses uncertainty to guide their refinement. By modeling co-occurrence, implication, and exclusion, ReCBM specifies how evidence is exchanged across concepts, while uncertainty modulates the contribution of each concept during thi
arXiv:2608.10030v1 Announce Type: new Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, we used AEROBAT to generate and test 79 hypotheses: designing 1,240 controlled experiments and executing 23,512 simulation rounds in total. Moderate-to-strong statistical evidence was found for 26 hypotheses, including some novel ones. In sum, our results demonstrate that automated behavioral scientific research on AI agents
arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting
arXiv:2608.10108v1 Announce Type: new Abstract: Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far back in the history. External memory stores such trajectories as structured representations, yet each structure provides a distinct and incomplete view. Existing multi-memory systems either read a fixed set of structures for every query, inflating context and introducing noise, or route each query to a single structure, preventing the composition of complementary evidence. A controlled analysis on AMA-Bench shows that the optimal memory configuration is typically neither a single structure nor the full union, but a tailored composition of multiple structural memories that varies with query and task demands. Motivated by these findings, we formulate structure-level dynamic selection: selecting and fusing a query-adaptive subset from a library of specialized memory
arXiv:2608.10153v1 Announce Type: new Abstract: Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI governance is four problems, not one, each with a mature governing science. The CASE framework assigns Control theory to the individual agent (intent as setpoint, guardrails as feedback, evaluation as observation), complex Adaptive systems theory to agent collectives (where emergence makes single-agent assurance non-compositional), Supervisory cybernetics to human-agent teams (where the Law of Requisite Variety shows unaided human oversight fails structurally), and Engineering operations to fleets (extending error budgets to decision quality so autonomy becomes a controlled variable). We formalize each layer, derive cross-layer coupling conditions, including a zero-touch deployment paradox where
arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G\"odel Machine and the Huxley G\"odel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modific
arXiv:2608.10171v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this escalating trend has introduced significant security vulnerabilities, necessitating the identification and mitigation of flaws arising from malicious exploitation. Red teaming assessments, conducted to evaluate model robustness through diverse adversarial inputs, are essential for exposing security risks and implementing countermeasures. Currently, red teaming is performed either manually by experts or automatically using predefined attack datasets. Nevertheless, manual testing remains time-consuming, while existing automated methods suffer from limited creativity due to their inherent dependency on fixed datasets. In this study, we propose an automated, human-independent, and adaptive approach leveraging GFlowNets to identify LLM vulnerabilities by utilizing one large lang
arXiv:2608.10176v1 Announce Type: new Abstract: Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that the recommendations respect explicit user constraints. In practice, public service directories are noisy and inconsistent, and general-purpose large language model (LLM) or AI-based chatbots frequently generate unreliable recommendations, citing unverified sources from the web. We investigate the impact of retrieval quality on constraint-aware recommendation in public service conversational systems built over noisy and heterogeneous service directories. We propose TRACE (Trustworthy Retrieval-Augmented Conversational Engine), a retrieval-based, constraint-aware framework that parses input user queries into structural and semantic constraints for downstream retrieval, with the help of a dual data representation schema. Using a curated statewide pantry directory and a synthetic query benchmark, we evaluat
arXiv:2608.10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating visual features into a universal latent representation that can be consumed by another model, but every message is transported as a dense tensor of the same size regardless of its content. A fixed-capacity dense tensor therefore need not have a fixed effective information density: some messages may use only a small fraction of the available representational degrees of freedom. This observation suggests that the communication channel may be substantially compressible. We study its redundancy by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations and measuring reconstruction, downstream utility, feature reuse, and token-level interventions across nine reasoning benchmarks. Relative to the or
arXiv:2608.10206v1 Announce Type: new Abstract: Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children's speech, with the privacy and compliance benefits that come with edge processing.
arXiv:2608.10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding controllers primarily rely on instantaneous operational variables or route-specific stop identifiers, which provide limited information about the functional and operational context of individual stops and constrain policy reuse across routes. This study introduces an LLM-assisted semantic stop representation for event-driven bus holding control. An LLM is used offline to transform heterogeneous stop information, including physical attributes, surrounding activity context, and historical operational characteristics, into fixed semantic embeddings that are incorporated into a deep Q-learning controller without requiring real-time LLM inference. Experiments are conducted in stochastic simulations calibrated with observed data from two bus routes. Compared with the best calibrated Daganzo baseline,
arXiv:2608.09934v1 宣布类型:新 摘要:大型语言模型(LLM)代理通过将问题分解为角色专业行为来提高任务性能。然而,由于每次用户请求的实时代理设计带来的计算成本和不稳定性,其实际部署通常受到限制。为了解决这个问题,我们提出了 LLM 代理工厂,这是一个基于检索的框架,使用超过 20K 个预定义代理配置文件按需构建领域特定且基于维基百科的代理。我们的框架支持两种模式:(1) 通过语义进行代理配置文件检索 via sema
arXiv:2608.09936v1 宣布类型:新 摘要:法国新闻标题是否将左翼和右翼民粹挑战者视为对称的“极端”,还是作为根本不同的政治对手?我们研究了 2022 至 2025 年间 25 家法语媒体发布的 28,592 条关于 La France insoumise (LFI) 和 Rassemblement National (RN) 的标题,并通过三模型 LLM 流程进行标注,该流程经过分层人工审计验证。最明确的发现是角色不对称而非价值不对称:冲突框架和战略博弈框架在模型和时间上更为稳健 than d
每週精選最值得關注的 AI 故事,直接發送至您的郵箱。