AI Builders Digest — 2026-08-05
🧠 模型发布
Shieldstral — Mistral 发布 3B open-weights 多模态安全分类器,把 moderation 表述为可配置的自然语言问答:推理时直接写入 policy,无需重新训练,同时覆盖 text、image、prompt-response 与 refusal detection,并输出连续的 yes/no safety score。官方称其在多项 text safety、refusal detection、policy adaptability 与 multimodal benchmark 上可匹敌或超过规模大至 7 倍的 guard model。可用性: Apache 2.0 开源权重,可在单张 16GB NVIDIA GPU 上运行。评价: 用小模型和可编程 policy 把安全 moderation 从固定 taxonomy 变成了部署时可调的基础设施。
LFM2.5-2.6B — Liquid AI 发布面向 agentic workload 的 2.6B 参数模型,预训练约 34T tokens,提供 128K context extension,支持 planning、tool calling 与 multi-step workflows;在 ToolSandbox 得分 77.83、Claw-Eval(EN)62.85、BrowseComp+(OpenClaw)26.89,且在多个 instruction-following benchmark 领先对比模型。官方测试中,M5 Max CPU 可达 220 tokens/s,手机上约 30 tokens/s,显存占用低于 2.5GB。可用性: Hugging Face 提供 base 与 post-trained open-weight 版本,原生支持 llama.cpp、MLX、vLLM、SGLang 与 ONNX。评价: 它把能实际调用工具的 agent 下沉到手机和 edge device,重点不是参数规模而是速度、隐私与零云端 token 成本。
Alpamayo 2 Super — NVIDIA 发布 34B open reasoning vision-language-action model,由 32B Cosmos 3 Super Reasoner 与 2B Action Expert 组成,支持最多七路摄像头的 360°感知,并统一输出轨迹、Chain-of-Causation reasoning trace、meta-action、2D grounding VQA 与 reasoning auto-label。官方技术资料给出 LingoQA 79.2,在 37 个模型中排名第一,超过 Qwen2.5-VL 72B 17.0 分、Gemini 2.5 Pro 15.1 分和 GPT-4o 23.2 分。可用性: Hugging Face 权重与 GitHub inference notebooks,OpenMDW-1.1 许可,支持 fine-tuning、衍生模型和商业再分发。评价: 这是把 AV 的 planning、reasoning、evaluation 与数据标注串成 cloud-to-car 工作流的开放基础模型。
🔥 新产品发现
- Nightcrawler — 运行在 smartphone 上的本地 AI pentesting agent,展示了移动端 agent 与安全工具链的结合。[HN 115👍]
- Soup — 让开发者在 4GB laptop GPU 上 fine-tune 8B model,降低个人设备进行 post-training 的门槛。[HN 116👍]
- Armature — 面向 MCP 的 agent session product analytics 与 evals,帮助观察工具调用、会话质量和 agent 行为。[HN 41👍]
- CostPerPrompt — 实时 AI API pricing 与真实 workload cost calculator,用于比较不同模型和工作流的实际成本。[HN 20👍]
- Implement-spec — harness-agnostic 的 spec-to-verified-PR agent skill,把规格、实现和验证连接起来。[HN 12👍]
- video-use — 让 coding agent 编辑视频,把软件 agent 的操作范围从代码扩展到 video production。[GitHub Trending]
- loopx — 面向 long-running AI agent teams 的状态 kernel,提供 durable goals、quota-aware auto-wake、executable todos 与 evidence logs。[GitHub Trending]
- rtk — 单一 Rust binary 的 CLI proxy,目标是在常见开发命令中减少 60–90% 的 LLM token 消耗。[GitHub Trending]
⚙️ 工具更新
Claude Code 2.1.222
- 修复 worktree-isolated session 与 subagent 可能对 main checkout 执行 destructive git command 的问题,隔离现在覆盖所有 session 类型的 file edits 与 Bash。
- 修复 background agent task 中 PreToolUse auto-allow hook 绕过 tool restriction 的问题。
- 修复 Team 和 Enterprise 的
/usage-credits在先前请求被 dismiss 后错误阻止重新申请的问题。 - 修复 HTTPS proxy 后 startup connectivity check 卡住并失败的问题,改用与 API request 相同的 proxy-aware transport,并提供清晰 timeout 信息。
- 修复已完成 response 被报告为 “Connection closed mid-response”、MCP usage 归因过度,以及新建 PR 未正确关联 session 等问题。
查看 Claude Code v2.1.222 changelog
📰 博客更新
Google AI Blog
- The latest AI news we announced in July 2026:Google 汇总了 2026 年 7 月的 AI 更新,涵盖面向大规模 agent 构建的 Gemini 模型、Gemini Robotics ER 2,以及音乐和视频生成工具。
Anthropic Engineering
- claude code auto mode
- harness design long running apps
- infrastructure noise
- building c compiler
- AI resistant technical evaluations
- demystifying evals for ai agents