-
ithome · 2026-07-25 · 科技产品 / rss
IT之家 7 月 26 日消息,Debian 项目开发者本周(7 月 24 日)发起议案,讨论未来是否允许 AI 大语言模型(LLM)开发 Debian。 据悉,本次议案共有三个选项:完全禁止大语言模型或其他 AI 工具、允许项目贡献者使用 AI 辅助工作、尽可能拒绝大语言模型和生成式 AI。 IT之家附三项提案详情如下: 提案 A:完全禁止使用大语言模型提交官方软件代码、文档、公告 该提案主张完全禁用大语言模型和生成式 AI 工具,即使是 AI 修改、开发者人工确认也不可行。提案认为,LLM 输出的代码无法确认版权归属,且 AI 大模型训练的数据来源存在争议, 很多时候开发者也无法保证 AI 生成代码中完全没有复制版权内容 。 同时,AI 可能会在生成的代码中使用过时 API、引入废弃写法,不符合 Debian 开发规范。并且 LLM 并不能真正理解代码,它只能根据训练数据预测可能出现的代码组合,因此无法保证生成的结果完全正确。 提案 B:允许使用 AI 辅助开发,但开发者必须全权负责 该提案将允许开发者使用 AI 工具辅助开发, 但必须确保提交的内容可以按照 Debian 许可证进行分发 。开发者还需要确保 AI 输出的内容完全符合开源许可证要求,如果包含第三方代码也需要独立确认。 此外,即使使用 AI 工具, 最终的责任仍将由按下提交按钮的人负责 。提交者必须保证代码质量、版权符合标准,不能以 AI 为由免责,并做好“AI 生成”标记。 提案 C:在实际可行范围内尽可能拒绝大语言模型 该提案认为,大语言模型存在大量问题,不应成为 Debian 的一部分,也不应取代人类。但由于大量开发者已经使用 AI 辅助编程,且大量上游项目可能已经使用 LLM,因此完全禁止 AI 工具和大语言模型已经不可行。 该提案要求所有 Debian 贡献者尽可能在工作中避免使用 AI 或 LLM,但如果某些情况下需要妥协,则具体情况具体判断。 同时所有内部邮件交流、Bug 报告、文章撰写必须完全由人类完成 ,不得使用 AI 工具辅助。 即使工作需要使用了 AI 工具,该提案也要求开发者披露。 项目维护者可自行决定某软件包是否允许 AI 代码 ,所有人必须遵守,如果违反则可能被处以警告、社区处分等。
-
github · 2026-07-26 · GitHub / LLM workflow stars:>300 / JavaScript
语言:JavaScript;stars:337;更新:2026-07-26T01:27:26Z
-
github · 2026-07-26 · GitHub / MCP server stars:>100 / Python
语言:Python;stars:2198;更新:2026-07-26T01:19:20Z
-
github · 2026-07-26 · GitHub / AI agent language:Python stars:>500 / Python
语言:Python;stars:3675;更新:2026-07-26T01:13:54Z
-
github · 2026-07-26 · GitHub / AI agent language:Python stars:>500 / Python
语言:Python;stars:4436;更新:2026-07-26T00:57:38Z
-
github · 2026-07-26 · GitHub / MCP server stars:>100 / JavaScript
语言:JavaScript;stars:1842;更新:2026-07-26T01:16:55Z
-
hf-papers · 2026-07-22 · AI论文 / hf-daily-papers / [object Object]
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Both outperform open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger. Beyond benchmarks, we analyze how harness choice (e.g., ZeroClaw, OpenClaw, Codex) and RL shape agent behavior. We find that some harnesses are substantially harder to learn than others, and that RL improves agentic reliability, such as self-verification, tool coverage, and completing multi-step plans, though critical abilities such as error recovery remain weak.
-
techcrunch-ai · 2026-07-24 · AI行业 / rss
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
-
techcrunch-ai · 2026-07-23 · AI行业 / rss
AegisAI co-founders developed AI agents that quickly analyze each message as a human would, paying attention to small anomalies that even the most elaborate checklist wouldn’t catch.
-
tavily · 2026-07-18 · 认知差与信息权重 最新 热点 趋势 / Tavily
# Thread. 我认识一个高中生,靠平台搬运内容,一个多星期就赚了40多万 就靠这三招 第一招:主动获取第一手信息 他每天固定时间刷 Reddit、Product Hunt、Medium、Substack 等海外平台。 这些国外最新的趋势、产品、观点和玩法,国内大多数人根本刷不到。 第二招:深度加工,本地化重构 他不做翻译,做重组。 表达按国内用户的习惯重写,再往里塞自己的案例和判断。 搬运的活儿,他干成了创作。 大多数人做内容的思路是“拿到了什么就发什么”,他多走了一步:先问自己三个问题——这条信息国内用户最缺的是哪个角度? 哪个情绪点能让他们停下来看完? 我能不能提供一个他们没想到的视角? 同样的素材,别人发一遍就沉了,他发完还有人在评论区追问。 第三招:全平台矩阵分发 同一份内容,他会做成不同形式投放到多个平台: 抖音 / 视频号 → 主攻口播短视频,真人出镜涨粉快 小红书 → 主攻高质量图文,用户愿意收藏和反复回看 头条 / 公众号 → 主攻长文章,沉淀深度内容和搜索流量 X(推特)/ 小红书 / 视频号 → 主攻长文或推文串,信息密度高、传播快 现在有了 AI 加持,效率直接起飞。. Log in or sign up for ThreadsSee what people are talking about and join the conversation. Log in with username instead.
-
hf-papers · 2026-07-22 · AI论文 / hf-daily-papers / [object Object]
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.
-
techcrunch-ai · 2026-07-25 · AI行业 / rss
OpenAI's fancy new AI keypad will be a lot of fun for some, while many others are probably not going to touch it.
-
tavily · 2026-07-25 · AI与Agent 最新 热点 趋势 / Tavily
# Hugging Face CEO shares his demands of OpenAI after 'rogue' agent hack: 'It deserves an unprecedented response'. Last week, after Hugging Face suffered an unusual security breach involving an AI agent running on OpenAI models, the company's CEO, Clem Delangue, boarded a flight to San Francisco to meet with the maker of ChatGPT. A week later, in the "spirit of transparency," Delangue shared on X what he asked of OpenAI. He said he asked the leading AI startup to release all the "traces" of the rogue agent for the public and research community to study. Last week, however, Hugging Face was struck with a security breach when an autonomous AI agent accessed some of its internal datasets. The company called the episode an "unprecedented cyber incident" and said it was working with Hugging Face on the investigation. Look out for an alert in your inbox the next time a new story is published!
-
ai-news · 2026-06-06 · finance_news / AI 基础设施 / 能源
特锐德推出了专供智算中心的模块化供电站“算电岛”,不仅能直供 800V 直流机房,官方号称配合 AI 优化调度,能让 Token 的综合用电成本直接砍掉约 30%。 启发:算力的尽头是电力。国内这种卷底层基建的降本思路非常实在,毕竟现在大模型拼到最后,拼的就是显卡折旧和电费谁能省下来。
-
hf-papers · 2026-07-22 · AI论文 / hf-daily-papers
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interaction histories without sacrificing environment sample efficiency remains underexplored. We term this problem Experience Distillation and develop an implementation that requires no further environment interaction beyond the collected experience. Experiments on 749 curated software-engineering tasks and six text-adventure games show that it retains at least 64.8\% of the gains from in-context learning across both domains, whereas direct supervised fine-tuning on the collected experience recovers only 3.8\%. Compared with classical reinforcement-learning baselines, in-context learning from trial-and-error experience followed by Experience Distillation matches their performance with at least \(9.6\times\) fewer environment samples.
-
hf-papers · 2026-07-22 · AI论文 / hf-daily-papers / [object Object]
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
-
hf-papers · 2026-07-21 · AI论文 / hf-daily-papers / [object Object]
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an LLM-driven agent loop, while methods with normal bodies remain standard deterministic Python. This gives developers and agents the same interface, so agent behavior can be tested, traced, refactored, and improved just like other software. This paper makes three contributions. (1) We present the agent-as-a-Python-object programming model and the design principles behind it. Where Python has existing abstractions, we adopt them directly. Agent-specific capabilities--context, events, state rendering, long-term memory, and validated LLM loops--are exposed through simple Pythonic APIs, so both developers and agents share one familiar programming model. (2) We identify six model-facing ideas that NOOA is, to our knowledge, the first to combine on a single surface: typed input/output, pass-by-reference over live objects, code as action, programmable loop engineering, explicit object state, and model-callable harness APIs for context and events. We find the community already converging on several of these ideas--often as experimental or partial features--and present the comparison to encourage further adoption. (3) We demonstrate that current models use this interface effectively, both in targeted capability tests and on agentic and reasoning benchmarks such as SWE-bench Verified and Terminal-Bench 2.0 and ARC-AGI-3.
-
techcrunch-ai · 2026-07-25 · AI行业 / rss
At libraries around the country, "Avoiding AI" workshops have elicited unprecedented demand.