-
github · 2026-07-28 · GitHub / Claude Code stars:>50 / Java
语言:Java;stars:291;更新:2026-07-28T01:24:30Z
-
github · 2026-07-28 · GitHub / AI agent language:Python stars:>500 / Python
语言:Python;stars:3411;更新:2026-07-28T01:23:42Z
-
github · 2026-07-28 · GitHub / AI agent language:Python stars:>500 / Python
语言:Python;stars:819;更新:2026-07-28T01:21:19Z
-
github · 2026-07-28 · GitHub / Claude Code stars:>50 / Go
语言:Go;stars:2487;更新:2026-07-28T01:26:59Z
-
github · 2026-07-27 · GitHub / AI 最新 趋势 stars:>50 / TypeScript
语言:TypeScript;stars:169;更新:2026-07-27T20:16:12Z
-
techcrunch-ai · 2026-07-27 · AI行业 / rss
Google’s AI Overviews now appear in 43% of searches, underscoring how quickly AI-generated answers are becoming the default way people discover information online.
-
ithome · 2026-07-28 · 科技产品 / rss
IT之家 7 月 28 日消息,科技媒体 Windows Latest 昨日(7 月 27 日)发布博文,示警用户谨慎从第三方软件站下载应用程序,由于 AI 降低了分发恶意软件的门槛, 目前已发现超过 70 个软件站存在分发恶意软件问题。 在运行机制方面,部分第三方软件站通过各种方式提高在谷歌搜索中的排名,部分站点的排名甚至高于官方网站,目前已发现 72 个存在恶意行为的站点。IT之家附上相关域名如下: 用户 u/Bogdan_X 在 Reddit 上解释说,该网站是用 WordPress 搭建的,内容是通用且不准确的 AI 生成的,并且使用了微软的旧版 logo。而这些网站的下载按钮竟然指向真正的微软商店页面,这很可能是它之前没有引起注意的原因。 在分发的软件方面,涉及 PowerToys、CrystalDiskMark、EasyBCD、Lively Wallpaper 和 Wintoys 等。 根据 Check Point Research 发布的研究报告,这些第三方软件站主要操作流程如下: 第一阶段:通过分发热门软件提高搜索排名 第二阶段:提供无害的真实下载源 第三阶段:在吸引足够用户关注后,赢得流量和信任后,悄然替换软件源,开始分发恶意软件。 在幕后黑手方面,该机构发现上述约 70 个站点由同一所有者通过 Epik Inc.注册,后于 2026 年 7 月转移至 Dynadot LLC。 部分网站会加载 Amazon CloudFront 脚本,拦截下载按钮点击。流量分发系统会根据访问者位置、浏览器及是否疑似机器人或安全研究人员,决定跳转目标。 Check Point Research 称,一组相关域名至少从 2025 年 9 月开始积累搜索排名,并于 2026 年 1 月开始分发恶意软件。
-
ithome · 2026-07-28 · 科技产品 / rss
IT之家 7 月 28 日消息,当地时间 26 日,据英国《金融时报》报道,咨询公司 Adaptavist 调查发现,约 65% 的白领 经常怀念 AI 出现前的工作方式 。约三分之一的受访者因为“会削弱创造力”而 愿意放弃生成式 AI ;也有约三分之一的人担心 AI 遭到滥用 。近一半受访者表示,为了核实 AI 生成内容是否准确,自己被迫投入更多时间,结果 与 AI 减少重复劳动的承诺背道而驰 。 AI 对创造力的影响同样引发担忧。一项针对普通人创意写作的研究发现,生成式 AI 虽然能够提供更多剧情转折,却也让 故事越来越相似、创意越来越趋同 。 IT之家从报道中获悉,宾夕法尼亚大学沃顿商学院另一项关于创新的研究也得出类似结论:AI 虽然提高了创意质量, 却削弱了创意多样性 。其中一位研究作者警告:“如果你把 ChatGPT 当成唯一的创意顾问,很快就会发现 自己的点子越来越少 ,因为 AI 给出的想法彼此太相似了。” 企业不断调整 AI 政策,也让员工无所适从。企业起初鼓励员工利用 AI 自动完成工作、大胆尝试各种应用,然而随着科技企业改为 按 Token 收费 ,企业又开始要求员工减少使用 AI,以控制成本。 与品牌合作管理历史档案和企业传承业务的 History Factory CEO 杰森 · 德雷塞尔表示,这种“怀旧”反映了人们在面对快速变化时, 对稳定、自主权和经济信心的渴望 ,AI 进一步强化了这种感受。因为 AI 不仅改变了人们的工作方式,也改变了人们的价值如何被衡量。
-
bilibili · B站热词 / social-hot
热度:482745;热度层:S
-
ithome · 2026-07-28 · 科技产品 / rss
IT之家 7 月 28 日消息,微软首席执行官萨蒂亚 · 纳德拉(Satya Nadella)今天(7 月 28 日)在 X 平台发布推文,宣布其首个网络安全模型 MAI-Cyber-1-Flash, 并表示该模型在 CyberGym 基准测试中取得了 95.95% 的成功率,超过 Anthropic 和 OpenAI 的同类系统。 纳德拉表示 MAI-Cyber-1-Flash 是微软首个专门针对网络安全而研发的 AI 模型,并将其接入多智能体(Agent)安全系统 MDASH 中,可以高效处理约 90% 的网络安全任务。而针对剩余 10% 较为复杂的网络安全任务,可以调用规模更大、成本更高的模型(官方示例中为 GPT-5.4)。 成本方面,微软表示,MAI-Cyber-1-Flash 和 GPT-5.4 的组合成本比当前使用 GPT-5.4、GPT-5.4 Mini 和 GPT-5.3 Codex 的 MDASH 配置成本降低了近 50%。 在性能方面,CyberGym 测试中, MAI-Cyber-1-Flash 和 MDASH 组合获得 95.95% 的最高成绩 ,相比较而言,Claude Mythos 5 模型的成绩为 83.8%;GPT-5.6 Sol 模型为 83.6%;GPT-5.5 Cyber 模型为 85.6%。 微软同时推出 Perception。该系统由多个智能体组成,服务于 MDASH 中的多类安全工作流。系统可持续监测威胁、修补漏洞,并关闭新的威胁入口。微软称,Perception 未来将把该模型用于更多安全工作流。 微软称,另一项统计显示,其每天处理的安全信号超过 100 万亿条。该公司将漏洞、攻击、防御和处置结果纳入强化学习流程。MAI-Cyber-1-Flash 还经过微软 AI 红队评估、自动化测试、专家对抗测试和第三方评估。
-
ithome · 2026-07-28 · 科技产品 / rss
北京时间 7 月 28 日,据《金融时报》报道,微软 AI 负责人表示,OpenAI 一个失控模型发起的黑客攻击事件,是针对 AI 网络攻击崛起发出的“警告信号”。与此同时,微软发布了一款安全产品,声称其性能优于竞争对手,且成本更低。 微软 AI CEO、DeepMind 联合创始人苏莱曼表示,OpenAI 上周承认其一个 AI 智能体在测试中逃脱了实验环境,攻击了创业公司 Hugging Face,这是一个“重要教训”。 “这些都是非常强大的工具,需要极其谨慎地处理。我们需要对细节保持极致关注,”他在接受《金融时报》采访时表示,“随着模型变得越来越强大,预防原则在这里将变得非常重要,我认为这是一个警告信号。” 随着 Anthropic 在 4 月发布功能强大的 Mythos 模型,AI 发现和利用软件漏洞的能力可能彻底改变网络安全领域,这引发了越来越多的担忧。苏莱曼的上述言论正是在这一背景下发表的。 微软安全负责人哈耶特 · 加洛特 (Hayete Gallot) 表示,公司“别无选择”,只能推动网络安全能力的边界,以防御其客户所面临的“层出不穷”的自主 AI 攻击。 “攻击者已经拥有了这些模型。因此,我们也必须不断评估并推进这项技术的发展,让防御者能够进行防御,”她补充道,“这不是我们可以选择做或不做的事情,我们根本没有选择。” 微软周一发布了一款专门的网络安全模型:MAI-Cyber-1-Flash。微软表示,该模型与 OpenAI 的 GPT 5.4 结合使用时,在一个核心行业基准测试中的评分高于 Anthropic 和 OpenAI 自身最佳网络安全产品。 虽然微软在寻求降低对 OpenAI 的依赖,寻求技术独立,但苏莱曼表示,将较小规模的 MAI-Cyber-1-Flash 部署在 GPT 之上,可以为客户降低成本。 微软表示,其自主开发的 AI 模型可以“高效处理”所有用户任务中的 90%,但在处理“极其困难的任务”时,仍然会使用 OpenAI 的模型。 苏莱曼表示,采用一种名为“集成框架”(harness) 的技术架构,就足以打造强大的 AI 能力,并内置安全防护机制。他还补充称,与完全使用 OpenAI 模型相比,这款新工具的运行成本降低了 50%。 微软表示,该产品的预览版将首先向特定客户开放。 相关阅读: 《 OpenAI 奥尔特曼警示 AI 垄断风险:担忧威权世界而非技术本身 》
-
hf-papers · 2026-07-22 · AI论文 / hf-daily-papers / [object Object]
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Both outperform open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger. Beyond benchmarks, we analyze how harness choice (e.g., ZeroClaw, OpenClaw, Codex) and RL shape agent behavior. We find that some harnesses are substantially harder to learn than others, and that RL improves agentic reliability, such as self-verification, tool coverage, and completing multi-step plans, though critical abilities such as error recovery remain weak.
-
techcrunch-ai · 2026-07-26 · AI行业 / rss
"The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!"
-
hf-papers · 2026-05-18 · AI论文 / hf-daily-papers / [object Object]
The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, "quality" refers to downstream training utility rather than surface-level textual properties. We introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models. For data construction, methods consume identical raw sources and are scored by fine-tuning a base model on their outputs jointly with Dolly-15k; alongside this track we release Data-Construction-Skill, a skill-guided agent that lifts the Dolly-only baseline by nearly 20 points absolute on Llama-3.1-8B Finance and is competitive with the strongest agent- and DataFlow-based methods in knowledge-extraction-dense domains. For data quality evaluation, scoring functions are scored by Pearson correlation with downstream performance on a shared candidate pool; we release the Distributional Alignment Score (DAS), a distribution-based evaluator that uses MMD between a candidate dataset and a domain proxy. DAS attains the strongest cross-model correlation in four of six domains and is the only metric clearing r > 0.70 simultaneously in Math, Science, and Medical, outperforming existing quality-, diversity-, and heuristic-based evaluators. DataPrep-Bench provides a unified, downstream-grounded framework for measuring progress on both capabilities as co-equal targets of LLM-driven data preparation.
-
hf-papers · 2026-07-21 · AI论文 / hf-daily-papers / [object Object]
Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the cost lands on the researcher at every iteration. Molt is a PyTorch-native training framework built to keep that cost small: a codebase compact and clean enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety, so the algorithm flow can be traced and changed end to end. The agent is an ordinary program, and one asynchronous loop trains multimodal and mixture-of-experts policies while never training on a token it did not generate, consistent in tokens, policy versions, and model semantics. Leanness does not cost performance: under a matched, fully asynchronous protocol, Molt is statistically comparable to a state-of-the-art Megatron-based stack. Molt is open source and provides recipes and containers at https://github.com/NVIDIA-NeMo/labs-molt.
-
techcrunch-ai · 2026-07-27 · AI行业 / rss
The issue appears to have originated from Claude’s “share chat” feature, which allows users to create links that enable anyone with the assigned URL view a conversation or project.
-
techcrunch-ai · 2026-07-27 · AI行业 / rss
OpenAI's Hugging Face breach has reignited debate over AI alignment and control, exposing competing views on whether increasingly capable AI should be better aligned, better contained, or both.
-
ai-news · 2026-06-06 · finance_news / AI 基础设施 / 能源
特锐德推出了专供智算中心的模块化供电站“算电岛”,不仅能直供 800V 直流机房,官方号称配合 AI 优化调度,能让 Token 的综合用电成本直接砍掉约 30%。 启发:算力的尽头是电力。国内这种卷底层基建的降本思路非常实在,毕竟现在大模型拼到最后,拼的就是显卡折旧和电费谁能省下来。