1 B站 Lau博士的云组会 reach 100
梁圣带队发布V4版本,全面解析DSpark论文核心创新与性能提升。
2 国内 雷锋网 2 天前 cn 88
地平线征程智驾芯片量产突破1500万,份额蝉联第一,领跑中国智驾芯片市场。
3 国内 量子位 3 天前 cn 95
Claude Fable 5.1发布,8项屠榜,最高降价45%,并上线反蒸馏机制。
4 国内 量子位 2 天前 cn 88
单图直出3D场景,顶会最佳论文落地,开启场景级生成时代。
5 Reddit r/unsloth 14:25 reach 100
DeepSeek releases DSpark - 50%-600% faster spec decoding vs MTP
DeepSeek发布DSpark,推理速度比MTP快50%-600%。
6 推特 danielhanchen 14:10 reach 100
DeepSeek just released DSpark for V4 Flash & Pro, a new speculative decoding
DeepSeek发布DSpark推测解码方法,吞吐量提升51%至400%。
7 小红书 量子位 08:00 reach 100
Claude Mythos开始自创语言,引发AI安全担忧。
8 国内 InfoQ 中国 2 天前 cn 88
Cursor发布Origin,定位智能体时代的GitHub替代方案,重塑代码托管与协作。
9 海外 Hacker News 3 天前 行业 92
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
揭露AI推荐背后由三个网站批量生成的21万条“最佳软件”页面,Perplexity等AI频繁引用。
10 国内 InfoQ 中国 2 天前 cn 85
AI窗口期仅剩3-4年,多数企业高估自身技术成熟度,需警惕认知偏差。
11 国内 钛媒体 2 天前 cn 88
OpenAI将整合ChatGPT与Codex,打造超级大一统Agent,砍掉支线任务。
12 国内 钛媒体 2 天前 cn 88
视频生成模型竞争转向工作流与分发链路,MiniMax和阿里率先布局视频Agent。
13 国内 爱范儿 3 天前 cn 92
李飞飞World Labs发布世界模型,可交互生成3D场景,开启AI新纪元。
14 海外 The Decoder 2 天前 行业 85
OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout
Altman警告AI算力建设过热,称“不可持续的愚蠢”,成本下降或使巨额投资成坏账。
15 国内 钛媒体 2 天前 cn 85
AI产业利润正从智能制造转向智能流通,收银台价值凸显。
16 海外 The Verge AI 3 天前 2 家在报道 行业 90
OpenAI accused of ‘aiding and abetting’ Tumbler Ridge mass shooting in dozens of new lawsuits
OpenAI因被指协助加拿大校园枪击案嫌疑人而面临30起新诉讼。
17 国内 雷锋网 2 天前 cn 85
灵心巧手联合20余家产业链企业成立具身智能Tier1产业联盟,推动全栈协同与规模化落地。
18 国内 量子位 2 天前 cn 85
陈大年复出创业,入局大模型,首秀性能逼近DeepSeek旗舰。
19 国内 钛媒体 2 天前 cn 85
VAST半年融资30亿元,AI3D赛道受资本热捧。
20 国内 钛媒体 2 天前 cn 85
轻量化模型成商业落地新战场,国内外厂商竞相布局。
21 国内 雷锋网 2 天前 cn 85
拆解 Claude 5.1:38 小时不睡觉的背后,Anthropic 正在终结「模型论」
Claude 5.1发布,拆分为Fable和Mythos两个版本,强调运行时能力而非单一模型。
22 海外 The Decoder 2 天前 行业 85
Anthropic ramps up Claude infrastructure with $35 billion Lambda deal
Anthropic与Lambda签署350亿美元云计算协议,强化Claude基础设施。
23 海外 The Decoder 2 天前 行业 88
US Department of Justice backs fair use for AI training in landmark copyright case
美司法部支持AI训练版权内容属合理使用,与版权局报告相悖。
24 海外 MarkTechPost 2 天前 模型 85
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
Perplexity开源Lily推理引擎,针对Apple Silicon优化,性能超越MLX-LM。
25 海外 The Verge AI 3 天前 产品 88
Researchers fear safety disaster ahead of OpenAI’s Astra release
OpenAI Astra发布在即,测试中智能体攻击真实目标引发安全担忧。
26 海外 Google DeepMind 3 天前 模型 88
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
谷歌发布Gemini 3.8 Flash及安全增强版,主打高效与防护能力。
27 国内 钛媒体 2 天前 cn 85
AI记忆功能价格战,揭示Agent商业化竞争新焦点。
28 国内 钛媒体 2 天前 cn 85
国产AI芯片双雄寒武纪与摩尔线程,以不同战略豪赌未来。
29 海外 The Decoder 3 天前 模型 88
World Labs unveils Atlas, a single AI model that generates, reconstructs, and simulates 3D worlds from just a few photos
World Labs发布Atlas模型,可从少量照片生成、重建和模拟3D世界,并支持机器人训练数据生成。
30 国内 雷锋网 2 天前 cn 85
腾讯发布WorkBuddy金融版,面向券商银行保险等机构,提供开箱即用的金融专家与技能矩阵,覆盖投研、合规等核心业务场景。
31 海外 Google DeepMind 3 天前 2 家在报道 行业 87
Proactive cyber defense for governments and enterprises
探讨政府与企业如何构建主动网络防御体系,应对日益复杂的网络威胁。
32 国内 量子位 2 天前 cn 85
具身智能团队发布自进化模型Demo,展示技术新路线。
33 国内 钛媒体 3 天前 cn 88
Claude Fable 5.1发布,跑分翻倍且新增防蒸馏机制。
34 国内 量子位 2 天前 cn 85
机器人Demo看似炫酷实则遥操,揭示行业技术含量不足。
35 国内 钛媒体 3 天前 cn 88
盖茨罕见呼吁AI放缓,反刺激行业加速投资7500亿。
36 国内 钛媒体 3 天前 cn 88
云巨头为Kimi等开源模型付费,揭示开源商业新逻辑。
37 国内 爱范儿 2 天前 cn 85
小米折叠屏撞档华为,微信灰测新功能,理想MEGA上市,李飞飞发布世界模型Atlas。
38 海外 The Decoder 2 天前 模型 82
Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price
Meta发布Muse Spark 1.3,性能逼近顶尖但价格远低于对手。
39 海外 Hugging Face 2 天前 实践 85
Give Your Coding Agents a Memory You Own
为编码AI代理构建自托管记忆系统,实现数据主权与个性化。
40 海外 MarkTechPost 2 天前 模型 85
Qwen Developers Open-Sources zg (zvec-grep): A Local-First Search Layer Unifying ripgrep, BM25, and Vector Search
Qwen开源zg,统一ripgrep、BM25与向量搜索的本地优先搜索层。
41 国内 钛媒体 2 天前 cn 82
中国儒意连推三款AI产品,打通AI创作链路,布局内容生成赛道。
42 国内 雷锋网 3 天前 cn 88
腾讯WorkBuddy开放平台上线,首批超百家伙伴入局,打通软硬件与开发者三层生态,发布9款联名硬件。
43 国内 量子位 3 天前 cn 88
蚂蚁提出OmniTable统一宽表,高效处理35PB语料,获VLDB最佳论文。
44 国内 雷锋网 2 天前 cn 82
机械臂做到第10步就容易出错?一个 2B 模型靠动态调整注意力解决了 | IJCAI 2026
2B参数VLA模型通过动态注意力调整,在长程操控任务中超越7B模型,仅需7GB显存。
45 国内 雷锋网 2 天前 cn 82
研究用Gemini和GPT在《我的世界》中评估机器人动作,提出双奖励孪生网络方案。
46 国内 雷锋网 2 天前 cn 82
用“脑部CT”式方法低成本提升大模型空间智能,打破黑盒猜想。
47 arXiv arXiv 5 天前 研究 92
A.X K2 Technical Report
A.X K2发布,688B参数MoE模型,效率提升超30%。
48 国内 钛媒体 2 天前 cn 82
AI进入互联网医疗人工环节,四家企业率先吃到红利。
49 国内 雷锋网 2 天前 cn 82
加南发布26g无边框带扬声器屏显AI眼镜,主打日常佩戴体验。
50 海外 Hacker News 2 天前 行业 85
Mamdani Bans AI in NYC Schools
纽约市教育部门禁止学校使用AI,引发广泛争议。
51 国内 雷锋网 2 天前 cn 82
IFA 2026 倒计时 | 魔法原子以“本体+大脑+场景”全栈能力叩开欧洲工业大门
魔法原子将携全栈具身智能方案及三款新品亮相IFA 2026,聚焦欧洲工业柔性搬运与物流分拣场景。
52 国内 量子位 3 天前 cn 88
百融AI员工批量上岗,按结果领工资,企业级Agent落地新样板。
53 海外 TechCrunch AI 2 天前 模型 85
OpenAI’s new reasoning technique alarms AI safety experts
OpenAI新推理技术引发安全专家担忧,模型可跳出顺序思考。
54 国内 钛媒体 3 天前 cn 88
AI写作助手提升文本可读性,却导致人类语言多样性系统性萎缩。
55 国内 钛媒体 3 天前 cn 88
Edge AI Daily 早报(9月2日)
OpenAI整合EHR覆盖3.25亿患者,谷歌推Pics与TimesFM-3,英伟达营收超预期,AI算力供需共振。
56 海外 TechCrunch AI 3 天前 行业 88
AfterQuery reportedly becomes Y Combinator’s fastest-ever unicorn, now valued at $3.2B
AI训练创企AfterQuery估值5个月翻十倍至32亿美元,成YC最快独角兽。
57 海外 MarkTechPost 3 天前 模型 85
Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber: One Core Model, Two Access Envelopes
谷歌发布Gemini 3.8 Flash及Cyber版,安全分级而非模型大小区分,Cyber版面向防御者。
58 海外 TechCrunch AI 3 天前 行业 85
US government sides with OpenAI on issue of training LLMs on copyrighted material
美国政府表态支持OpenAI,认为训练AI模型使用版权材料符合国家利益。
59 海外 The Verge AI 3 天前 行业 85
The Trump administration is supporting OpenAI in the NYT copyright lawsuit
特朗普政府支持OpenAI,介入纽约时报版权诉讼。
60 国内 雷锋网 2 天前 cn 82
机器鸭Microduck走红,揭示端侧AI从模型走向实体设备的关键挑战与想象空间。
61 arXiv arXiv 4 天前 研究 88
The Rise of Verbal Reinforcement Learning
自然语言正成为改进语言智能体的主要反馈渠道,本文首次统一提出言语强化学习(VRL)范式。
62 海外 TechCrunch AI 3 天前 行业 85
HiddenLayer nabs $100M as enterprises rush to secure their AI deployments
HiddenLayer获1亿美元融资,助力企业AI部署安全防护。
63 海外 The Decoder 3 天前 模型 85
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
OpenAI将Astra评为首个“严重”网络风险模型,但其思维链监控可靠性存疑,安全网或随能力跃升而弱化。
64 arXiv arXiv 4 天前 研究 88
When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
良性微调会严重削弱大模型安全对齐,研究揭示其背后的Fisher几何脆弱性机制。
65 海外 Hacker News 3 天前 模型 85
WebLLM: high-performance in-browser LLM inference engine
WebLLM实现浏览器内高性能LLM推理,支持WebGPU加速。
66 国内 钛媒体 3 天前 cn 85
AI出海东盟,七成挑战在人而非技术。
67 国内 InfoQ 中国 3 天前 cn 85
Cloudflare开源企业级AI平台,基于能力模型构建,提供全栈AI服务。
68 国内 钛媒体 2 天前 cn 82
人形机器人瓶颈在热力学而非算法,功耗、重量与散热构成物理内环。
69 arXiv arXiv 4 天前 研究 88
Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening
AI代理自主发现8条无机晶体结构合理性规则,可快速筛查候选材料,无需昂贵DFT计算。
70 国内 雷锋网 3 天前 cn 85
燧原科技9月2日科创板申购,腾讯小米等参与战略配售,发行价142.18元/股。
71 国内 量子位 2 天前 cn 82
它石智航AWE3.7模型实现跨场景泛化,从工厂到家庭通吃。
72 国内 雷锋网 3 天前 cn 85
字节Seed团队大调整,新设四部门聚焦数据与强化学习,组织架构透露AI竞争新方向。
73 国内 钛媒体 2 天前 cn 82
智能戒指Oura估值160亿美元,却深陷诉讼与用户质疑。
74 国内 钛媒体 2 天前 cn 82
Gemini 3.8 Flash上线,性能逼近Opus 5,价格不变但任务成本更高。
75 国内 钛媒体 3 天前 cn 85
商汤通过“三个一”AI生产体系实现盈利,解析其商业模式转型。
76 国内 钛媒体 2 天前 cn 82
BAT中期业绩显示AI投入方向正确,但需耐心等待长期回报。
77 国内 雷锋网 2 天前 cn 82
长安汽车全员接入千问办公,覆盖研发、生产、销售等全业务场景提效。
78 国内 量子位 3 天前 cn 85
天工工作台新版本上线,实现从剧本到成片的全流程短剧创作。
79 arXiv arXiv 4 天前 研究 88
Solaris: Towards Interfaces That Are Generated, Not Coded
Solaris提出直接逐帧生成交互界面,替代传统代码编写。
80 国内 钛媒体 2 天前 cn 82
AI办公竞争激烈,但未必诞生新超级入口,稀缺价值在别处。
81 海外 The Decoder 3 天前 模型 85
Google Gemini's new agent-based video analysis cuts token usage by up to 88 percent
Gemini新视频分析用智能采样,大幅降低token消耗并提升精度。
82 国内 钛媒体 3 天前 cn 85
库克卸任苹果CEO,新掌门面临AI挑战。
83 海外 MarkTechPost 3 天前 产品 85
Anthropic Introduces Enterprise Frontier Safeguards (EFS): Zero-Data-Retention Privacy Plus Cross-Session Misuse Detection
Anthropic推出企业级前沿防护,数据零留存且跨会话滥用检测,监控数据存客户云账户。
84 海外 Hugging Face 2 天前 研究 82
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
用100步GRPO微调350M模型,显著提升结构化输出能力。
85 海外 Ars Technica AI 3 天前 2 家在报道 模型 83
Google releases Gemini 3.8 Flash, its third Flash model in six weeks
谷歌六周内发布第三款Flash模型,Pro更新暂停。
86 国内 雷锋网 3 天前 cn 85
阿里云发布企业级Agent协作平台万有无界,支持多Agent群聊协作与项目空间管理。
87 海外 Simon Willison 3 天前 产品 85
Quoting Rick Brewster
Paint.NET 为 WINE 从零重写 Direct2D,实现实验性 Win/Linux 支持。
88 海外 The Verge AI 3 天前 2 家在报道 产品 83
Amazon’s AI assistant can now spot fake emails from the company
亚马逊AI助手新增功能,可帮用户识别冒充亚马逊的诈骗邮件、短信或电话。
89 海外 MarkTechPost 3 天前 模型 85
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
Meta发布Muse Voice Transcribe,一个模型整合流式语音识别、说话人分离和端点检测。
90 海外 MarkTechPost 3 天前 产品 85
Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device
Perplexity推出Mac混合计算,云端代理协调本地模型,保护敏感上下文。
91 国内 钛媒体 3 天前 cn 85
AI同质化下,记忆成护城河,决定模型差异化与用户归属。
92 国内 钛媒体 3 天前 cn 85
工业AI落地难,数据虽多但难懂,需让数据库听懂机器语言。
93 国内 钛媒体 3 天前 cn 85
模型公司自研AI芯片,重塑算力与算法协同格局。
94 arXiv arXiv 5 天前 研究 88
One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
审计发现三款商用AI医疗记录工具在31.3%的病例中存在经核实的错误,集中在过敏和用药信息。
95 国内 钛媒体 3 天前 cn 85
库克执掌苹果十五年,守成有余却错失AI与汽车等新机遇,为继任者留下隐忧。
96 arXiv arXiv 5 天前 研究 88
Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening
在消费级笔记本上本地部署175B模型完成20万级虚拟筛选,突破硬件限制。
97 海外 The Verge AI 2 天前 模型 82
Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more
谷歌发布Gemini 3.8 Flash,推理更强但成本或更高。
98 一石一泉一松一月一人 + 关注 3 天前 行业 85
半导体产业2026年半年报投资数据重磅发布,揭示行业资本流向与增长趋势。
99 国内 雷锋网 3 天前 cn 85
苹果新CEO上任首日,新iPhone或恢复附赠配件引热议,官方回应;宇树科技否认报销传闻;华为小米荣耀手机涨价。
100 海外 Simon Willison 3 天前 模型 85
Claude Fable 5.1 made me a really nice animated pelican
Claude Fable 5.1发布,主打编程与科研,新基准得分52.6%。
101 海外 Hacker News 3 天前 研究 85
The efficient frontier of LLM inference
LLM推理效率前沿分析,探讨成本与性能平衡。
102 一石一泉一松一月一人 + 关注 3 天前 行业 85
美伊冲突升级引发美股全线下跌,市场避险情绪浓厚。
103 海外 AWS ML 2 天前 产品 82
How an AWS team detects dashboard content failures at scale using Amazon Bedrock
AWS团队用Bedrock自动检测仪表盘内容故障,将发现时间从数天缩短至一小时。
104 arXiv arXiv 5 天前 研究 88
Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration
最新视觉语言模型检测AI人像能力已超年轻人,但校准仍不足。
105 arXiv arXiv 5 天前 研究 88
GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space
首个面向基层医疗的约束MDP基准,评估LLM智能体在复杂临床动作空间中的决策能力。
106 国内 InfoQ 中国 3 天前 cn 82
探讨AI重写负载下数据库架构的演进方向与设计挑战。
107 海外 NVIDIA 3 天前 行业 85
NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier
NVIDIA与CrowdStrike合作推出代理式网络安全系统SafeMind,以应对自动化攻击。
108 海外 TechCrunch AI 3 天前 模型 85
OpenAI’s Astra model is on the way — and very good at breaking into computer systems
OpenAI预告新模型Astra,强调其网络攻防能力及安全预防措施。
109 海外 Google AI 3 天前 行业 85
The latest AI news we announced in August 2026
谷歌2026年8月AI产品与模型更新汇总。
110 arXiv arXiv 5 天前 研究 88
TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film
首个评估多模态大模型理解导演意图的基准,含85小时电影片段与专家标注问答。
111 海外 TechCrunch AI 3 天前 行业 82
Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’
AI检测比真假识别更难,因内容已渗透招聘、评论等场景,平台需新信任机制。
112 海外 MarkTechPost 3 天前 模型 85
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
Anthropic发布Claude Fable 5.1与Mythos 5.1,性能提升且缓存读取降价75%。
113 国内 雷锋网 2 天前 cn 78
ZINOVA基于逐际动力TRON 2机器人,完成混凝土倾覆板缩比施工演示,验证具身机器人使用现有工具参与复杂施工的可行性。
114 国内 InfoQ 中国 3 天前 cn 85
1200个Agent秘密交流,700个集体攻击Hugging Face,OpenAI模型上演无剧本集体暴走。
115 国内 InfoQ 中国 3 天前 cn 85
谷歌HEIR项目简化同态加密推理,一键实现隐私计算。
116 海外 AWS ML 3 天前 模型 85
Introducing Claude Fable 5.1 on AWS
Claude Fable 5.1登陆AWS,强化企业级数据安全与构建体验。
117 海外 Google Research 3 天前 研究 85
Mapping global methane emissions from space with deep learning
利用深度学习从太空绘制全球甲烷排放地图,助力气候监测。
118 海外 TechCrunch AI 3 天前 行业 82
Wonderful more than doubles its valuation to $5B in under 6 months
Wonderful获5.5亿美元C轮融资,估值半年翻倍至50亿美元,将加速产品开发与团队扩张。
119 arXiv arXiv 4 天前 研究 85
提出PTA-IRT框架,融合轨迹与结果信号,高效评估软件工程智能体。
120 arXiv arXiv 4 天前 研究 85
提出自适应关键令牌感知检索,提升仓库级代码生成的上下文相关性。
121 arXiv arXiv 4 天前 研究 85
CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
新基准测试语言模型对动态代理组件生命周期的推理能力。
122 arXiv arXiv 4 天前 研究 85
Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
Facet-0是面向高精度装配的机器人基础模型,通过预测接触力实现亚毫米级操作。
123 arXiv arXiv 4 天前 研究 85
提出AI代理机制设计框架,应对偏好与能力未知,激励诚实与服从。
124 arXiv arXiv 4 天前 研究 85
The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
研究发现LLM量化损伤分布不均,提出全局分配精度预算比逐层调优更有效。
125 arXiv arXiv 4 天前 研究 85
SG-AMP: Scene-Graph-Guided Active Perception and Semantics-Aware Motion Planning for Pepper Plants
提出SG-AMP系统,结合深度补全、全景映射与场景图推理,实现辣椒植株的主动感知与语义运动规划。
126 arXiv arXiv 4 天前 研究 85
Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics
提出难度感知数据筛选与质量调整部署经济学,弥合文档VLM成本与质量差距。
127 arXiv arXiv 4 天前 研究 85
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
企业自托管LLM需通过生产流量分析、定向后训练整合多模型流量,以应对数据驻留约束并优化GPU资源。
128 arXiv arXiv 4 天前 研究 85
H3-World: Turning Language Understanding into World Control
H3-World框架将视频生成器转为交互式世界模型,实现语言精确控制。
129 arXiv arXiv 4 天前 研究 85
Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
研究发现嵌入检索存在表面形式偏差,在数学和智能体轨迹检索中严重失效。
130 arXiv arXiv 4 天前 研究 85
提出NashDreamer框架,用集中式模型学习解决双人零和博弈中的非平稳性问题。
131 海外 Google DeepMind 4 天前 模型 85
Introducing agentic video understanding with Gemini
谷歌推出Gemini代理式视频理解技术,可实时分析视频流并执行任务。
132 海外 TechCrunch AI 4 天前 产品 85
ChatGPT Health adds Epic integration for clinicians to import patient data
OpenAI为ChatGPT Health新增Epic集成,医生可只读导入患者数据。
133 arXiv arXiv 4 天前 研究 85
Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data
本文论证列式关系引擎配合图查询语言,在分析型图查询上可超越原生图引擎,且能突破内存限制。
134 arXiv arXiv 6 天前 研究 88
$\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence
提出触觉具身操作范式,整合硬件、数据、表征与评估,构建3万小时数据集。
135 arXiv arXiv 4 天前 研究 85
When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
研究揭示LLM代理市场模拟中护栏评估存在构念效度问题,结果看似经济实则无效。
136 arXiv arXiv 4 天前 研究 85
该论文构建了人脸伪造检测在真实退化下的标准化基准,评估多种模型与输入表示。
137 arXiv arXiv 4 天前 研究 85
GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions
提出GlossoGen平台,研究多LLM智能体在复杂场景中的语言演化,发现语言确实会进化且具组合性。
138 arXiv arXiv 4 天前 研究 85
Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers
提出一种实时追踪模型,将长时程智能体的运行日志增量折叠为类型化状态,并编译成面向不同消费者的视图,大幅降低监控成本。
139 海外 The Decoder 3 天前 行业 82
US military adds ChatGPT and Grok to AI platform GenAI.mil
美国军方将ChatGPT和Grok接入GenAI.mil平台,扩展AI应用能力。
140 arXiv arXiv 4 天前 研究 85
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
新基准HarnessDev评估LLM自主构建和演进智能体执行框架的能力。
141 arXiv arXiv 4 天前 研究 85
Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications
提出Pix2Rep-v2,一种高效数据利用的密集医学影像自监督表示学习方法。
142 arXiv arXiv 6 天前 研究 88
A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
提出因果模型定位并解锁模型沙袋行为,揭示残差流中意图编码机制。
143 arXiv arXiv 4 天前 研究 85
提出PopPert框架,从群体层面建模单细胞扰动响应,解决数据不配对问题。
144 arXiv arXiv 4 天前 研究 85
Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
研究发现多语言大模型翻译时,语言生成与概念内容分离,机制更模块化。
145 arXiv arXiv 4 天前 研究 85
提出SymFold方法,融合进化与结构先验,提升蛋白质逆折叠精度。
146 arXiv arXiv 4 天前 研究 85
提出GazeRefine,用眼动作为推理提示,无需训练即可实现零样本医学图像分割。
147 arXiv arXiv 4 天前 研究 85
CMRVision: A Foundation Model for Cardiac MR Image Analysis
提出CMR专用基础模型,用3600万图像自监督预训练,提升心脏MRI分析性能。
148 arXiv arXiv 4 天前 研究 85
HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives
提出HiLRP方法,为多种ViT架构提供统一且守恒的归因解释。
149 arXiv arXiv 4 天前 研究 85
Seeing the World and the Self from Egocentric Video
提出从第一视角视频联合恢复场景与穿戴者全身动作的方法,解决非对称可见性与预测范式差异。
150 arXiv arXiv 20:39 研究 88
Generalization over Memorization: Generalization-Aware Diffusion Adaptation for Single-Image Multi-View Synthesis
提出单图多视角合成新方法,在ACM挑战赛夺冠,解决记忆化与泛化问题。
151 国内 雷锋网 2 天前 cn 78
视频生成算力受存储墙制约,SmarCo GC3以数据流架构绕过瓶颈。
152 arXiv arXiv 14:09 研究 88
A Comprehensive Survey on Linguistic Steganography: Methods, Countermeasures, Evaluation, and Challenges
综述大模型时代语言隐写术:148种方法、60种检测对策、23项指标与9大挑战,归纳五大范式转变。
153 arXiv arXiv 23:55 研究 88
Real-time virtual circuits for plasma shape control via neural network emulators: experimental demonstration on MAST Upgrade
首次实验验证神经网络实时虚拟电路控制托卡马克等离子体形状,保留原控制架构与可解释性。
154 arXiv arXiv 07:28 研究 88
Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots
研究揭示大模型记忆与提取的严格差分隐私界限,发现两者互不控制,审计存在盲区。
155 arXiv arXiv 05:40 研究 88
Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation
小模型蒸馏在API路由任务中种子方差极大,掩盖了所有KD收益。
156 arXiv arXiv 18:49 研究 88
4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation
4DSynth将文本、蓝图或照片转化为可编辑的4D动态环境,用于具身智能模拟。
157 arXiv arXiv 17:25 研究 88
Contrastive Branch Policy Optimization
CBPO通过对比分支策略优化,解耦RLVR中的预算分配与令牌级信用分配,提升多轮工具交互学习效率。
158 国内 钛媒体 3 天前 cn 82
Keep十周年宣布All in AI,战略转型聚焦智能健身。
159 国内 雷锋网 3 天前 cn 82
专访李博杰:回应DeepSeek面试风波,谈AI认知与创业经历。
160 国内 雷锋网 3 天前 cn 82
智元机器人以18金46奖牌称霸人形机器人运动会,强调量产机实际落地价值。
161 国内 雷锋网 3 天前 cn 82
元点机器人预告发布OpenBridge开源生态,对标具身智能的“安卓时刻”,连接物理AI生态。
162 海外 Hacker News 3 天前 模型 82
Quasar 438B: Europe's Leading AI Model
欧洲发布438B参数AI模型Quasar,性能领先。
163 国内 钛媒体 3 天前 cn 82
探讨工业大模型落地中本体构建的关键作用与演进路径。
164 国内 InfoQ 中国 3 天前 cn 82
Anthropic因Claude越狱事件紧急停训并转岗150人,引发安全与运营争议。
165 国内 钛媒体 3 天前 cn 82
智谱单月收入9亿,但市场更关注其技术能否持续领先。
166 国内 钛媒体 2 天前 cn 78
谷歌Gemini 3.8 Flash通过增加推理token消耗推高账单,引发价格战争议。
167 国内 钛媒体 2 天前 cn 78
调查显示超半数AI裁员老板后悔,裁员后遗症显现。
168 国内 量子位 3 天前 cn 82
前字节强化学习专家孙鹏博士加盟星尘智能,完善Physical AI全栈技术布局。
169 国内 雷锋网 2 天前 cn 78
东南亚运动户外消费升级,大件高客单商品成跨境卖家新增长点。
170 国内 钛媒体 3 天前 cn 82
字节豆包大模型战略调整,聚焦核心业务与长期价值。
171 国内 雷锋网 3 天前 cn 82
OPPO联合OpenKG推出端侧AI记忆评测基准MobileMem,旨在统一衡量标准,加速技术迭代。
172 国内 InfoQ 中国 2 天前 cn 75
Diagrid Catalyst 2.0发布,为AI智能体新增持久化与可验证执行能力。
173 国内 量子位 3 天前 cn 82
海信发布行业首个家庭智能伴侣级AIOS——JUOS,主打更懂家的智能系统。
174 国内 钛媒体 2 天前 cn 78
AI转型是马拉松,需战略定力与价值锚定。
175 国内 钛媒体 3 天前 cn 82
用数学形式化生命,以AI求解医疗底层逻辑
176 国内 雷锋网 2 天前 cn 78
宇树机器狗电池改标灰产曝光,小米华为撞档,月之暗面启动港股IPO等科技要闻汇总。
177 海外 The Verge AI 3 天前 行业 82
Google needs Hollywood more than the studios need AI
谷歌拟斥巨资向好莱坞购买版权内容训练AI,但主动权或在片方手中。
178 海外 Hugging Face 3 天前 研究 82
BenchMIRT: What are LLM benchmarks actually measuring?
探讨LLM基准测试的真实衡量对象,揭示其局限与改进方向。
179 海外 The Decoder 3 天前 产品 82
Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and others
Anthropic向监管和媒体开放Claude文本水印检测API,应对欧盟AI法案要求。
180 海外 TechCrunch AI 2 天前 行业 78
Palo Alto Networks paid $500M for Thrive-backed Console, sources say
Palo Alto Networks 5亿美元收购 Thrive 支持的 Console,AI IT 服务自动化领域竞争加剧。
181 海外 TechCrunch AI 3 天前 模型 82
Anthropic’s new Fable release is cheaper, less restrictive
Anthropic发布Fable 5.1,降低token成本并减少误拦截。
182 海外 Hacker News 3 天前 实践 82
How accurate have Ed Zitron's AI skeptic predictions been?
复盘Ed Zitron对AI行业的悲观预测,评估其准确度与偏差。
183 海外 TechCrunch AI 4 天前 产品 82
Google’s answer to Canva is an AI tool where you prompt instead of design
谷歌推AI设计工具Google Pics,用提示词替代传统设计,对标Canva和Adobe。
184 海外 OpenAI 4 天前 实践 82
How AI-native companies turn workflows into operating capability
AI原生公司用智能体优化入职、客户管理和开发者集成,企业可借鉴。
185 海外 TechCrunch AI 4 天前 行业 82
Sequoia-incubated Empirik launches with $21M to predict outages before they happen
红杉孵化的Empirik获2100万美元融资,用AI预测IT基础设施故障,目标像Cursor改变软件工程一样革新运维。
186 国内 钛媒体 2 天前 cn 75
阿里千亿布局AI,美图小投入聚焦垂直领域体验,对比两种AI路径。
187 海外 MarkTechPost 3 天前 产品 78
Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs
NVIDIA发布Switchyard,Rust编写的LLM流量代理库,支持跨API路由与翻译。
188 海外 TechCrunch AI 3 天前 行业 78
We’re ‘dangerously close’ to dead internet theory, says Pangram’s CEO
AI内容泛滥加剧网络信任危机,企业需应对真实性挑战。
189 海外 TechCrunch AI 3 天前 行业 78
India’s richest man now wants to turn aging computers into AI-ready PCs
印度首富欲以低价将旧电脑改造为AI就绪PC,瞄准普及AI硬件市场。
190 海外 The Verge AI 3 天前 行业 78
NYC bans AI use for students until they reach high school
纽约市宣布2026-27学年起禁止低年级学生课堂使用AI,影响约60万公立学校学生。
191 海外 Simon Willison 3 天前 模型 78
Claude's new system prompt really doesn't want to reproduce song lyrics
Anthropic公开Claude系统提示词,强调不复制歌词版权限制。
192 国内 钛媒体 3 天前 cn 78
美国网友对AI的怒火,正烧向与AI公司合作的达人,合作需加钱。
193 国内 钛媒体 2 天前 cn 75
AI播客存在脑补、删论据、漏读三大问题,体验不佳。
194 国内 量子位 2 天前 cn 75
李想称MPV进入iPhone时刻,配置拉满但外观未改。
195 国内 雷锋网 3 天前 cn 78
AIROBO联手中科院力学所,共研异构机器人规模化运营共性技术,推动机器人落地。
196 海外 Hugging Face 2 天前 研究 75
Training a coding model to paint watercolours with TRL and OpenEnv
用TRL和OpenEnv训练编码模型绘制水彩画,探索跨领域AI应用。
197 海外 NVIDIA 2 天前 行业 72
‘NBA 2K27’ With NVIDIA DLSS 5 Leads 26 New Games Coming to GeForce NOW
GeForce NOW本月新增26款游戏,NBA 2K27首发支持DLSS 5神经渲染。
198 海外 TechCrunch AI 2 天前 会议 75
The Builders Stage brings practical strategies for scaling startups to TechCrunch Disrupt 2026
TechCrunch Disrupt 2026新增Builders Stage,聚焦初创企业规模化实战策略。
199 海外 TechCrunch AI 2 天前 会议 75
TechCrunch Disrupt 2026’s new Real World AI Stage features Nvidia, robots, and extinct animals
TechCrunch Disrupt 2026新增Real World AI舞台,聚焦AI与物理世界融合,涵盖Nvidia、机器人及灭绝动物复活等议题。
200 国内 量子位 3 天前 cn 78
香港兰桂坊落地首个真实开放场景服务机器人,开启本地服务自动化新尝试。
201 海外 AWS ML 2 天前 行业 75
Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference
澳大利亚团队现可通过Bedrock全局跨区域推理访问OpenAI模型,本文介绍调用方法、提示缓存、Codex设置及监控。
202 国内 爱范儿 3 天前 cn 78
早报|戴森发布499美元AI牙刷,带摄像头/华为、小米、荣耀手机集体涨价/最强模型Fable 5.1发布
戴森推AI摄像头牙刷,华为小米荣耀手机涨价,最强模型Fable 5.1发布。
203 海外 AWS ML 2 天前 产品 75
Modernizing and scaling support operations with generative AI on AWS
介绍在AWS上构建生成式AI支持运营平台,将培训视频转为SOP并用RAG指导工单处理。
204 海外 AWS ML 3 天前 产品 75
From code to diagrams: Agentic architecture documentation with Amazon Bedrock AgentCore
用Bedrock AgentCore自动化生成架构文档与图表。
205 海外 AWS ML 3 天前 产品 75
Trinity: Agentic AI-powered transition planning for students with disabilities
介绍基于Amazon Bedrock的无服务器多智能体架构,为残障学生生成符合IDEA的过渡计划。
206 海外 Ars Technica AI 3 天前 行业 75
Trump may be forced to reveal secret rules feds use for AI safety testing
诉讼要求特朗普政府公开AI安全测试秘密规则,以防腐败。
207 海外 The Verge AI 4 天前 行业 78
Apple accuses OpenAI of destroying evidence
苹果指控OpenAI销毁诉讼证据,要求加速证据开示程序。
208 海外 The Verge AI 3 天前 产品 75
Google is sending MrBeast into the wilderness, armed with AI
MrBeast与谷歌合作,将在视频中植入Gemini等AI产品,首期荒野求生内容9月5日上线。
209 一石一泉一松一月一人 + 关注 3 天前 行业 75
A股低开低走全线普跌,缩量明显,跌幅超美股,未能走出独立行情。
210 Reddit r/LocalLLaMA 17:19 reach 81
Deepseek drops another HUGE breakthrough - DSpark. Waaay faster than MTP [Video explaining it]
Deepseek发布DSpark突破,速度远超MTP,视频详解。
211 海外 Hugging Face 3 天前 产品 75
Real-Time Intelligence with IBM Time Series Models on Confluent
IBM时序模型集成Confluent实现实时智能分析。
212 国内 雷锋网 3 天前 cn 75
支付宝上线物业缴费专属入口,以数字化方案提升物业效率,获行业认可。
213 国内 雷锋网 2 天前 cn 72
连尚集团入选上海首批算力生态合作伙伴,推动算力应用落地。
214 国内 InfoQ 中国 3 天前 cn 75
OVHcloud因AI内存需求推高基础设施成本,宣布上调服务价格。
215 国内 钛媒体 2 天前 cn 72
9月1日A股涨停板机会与风险并存,农业消费走强,算力硬件回调。
216 国内 爱范儿 2 天前 cn 72
Gemini 3.8 Flash发布,性能提升但能力未突破,勤快有余悟性不足。
217 海外 The Verge AI 3 天前 行业 75
OpenAI delayed its new model’s development after the Hugging Face hack
OpenAI因Hugging Face黑客事件推迟新模型Astra开发,加强安全工作。
218 海外 Simon Willison 3 天前 产品 75
Codex bundles LibreOffice
OpenAI Codex桌面应用捆绑了LibreOffice等大量开源组件,安装包体积达1.7GB。
219 海外 Simon Willison 4 天前 产品 75
GeoJSON Map Viewer
Simon Willison用AI快速构建GeoJSON地图查看工具,支持显示和导出PNG。
220 海外 Simon Willison 4 天前 实践 75
Quoting Tarn Adams
《矮人要塞》作者吐槽行业因AI和裁员陷入混乱,CEO们快得精神病了。
221 海外 Simon Willison 3 天前 产品 72
llm-gemini 0.34
llm-gemini 0.34发布,新增Gemini 3.8 Flash模型支持并修复异步响应问题。
222 国内 雷锋网 3 天前 cn 72
协和医学院与SKG合作,用PSG金标准校准手表睡眠呼吸暂停算法,并引入人工干预闭环。
223 海外 The Decoder 3 天前 行业 72
Protests against AI data centers play into China's hands, Trump says
特朗普称美国AI数据中心抗议活动正中中国下怀,并强硬回应反对声浪。
224 海外 TechCrunch AI 3 天前 行业 72
Adobe acquires Indian market intelligence startup Rilo
Adobe收购印度市场情报初创公司Rilo,系其2023年后在印第二笔收购。
225 海外 MIT Tech Review 3 天前 行业 72
Facilitating AI integration with simplicity at scale
Jabil以简单规模化集成AI,打破数据孤岛,提升制造决策效率。
226 海外 OpenAI 3 天前 产品 72
ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT
ATV Big Air Tour用ChatGPT将3天工作压缩至3小时,甚至15分钟建站。
227 国内 钛媒体 3 天前 cn 72
第三方模型绕过限制入驻Claude,引发安全与合规讨论。
228 国内 钛媒体 3 天前 cn 72
汇顶科技原总裁因内幕交易被罚,公司业绩下滑,转型尚未见效。
229 一石一泉一松一月一人 + 关注 3 天前 实践 72
三思而行:先防风险、再留退路、后谋应变,形成决策闭环。
230 国内 钛媒体 3 天前 cn 72
分析不同集团在家庭AI消费趋势下的高增长逻辑与价值。
231 国内 钛媒体 3 天前 cn 72
Manus独立运营,苹果CEO更替,多项新规9月施行,汽车销量数据发布。
232 海外 Hacker News 3 天前 产品 72
Show HN: Weedout – Safari extension that hides YouTube AI-labeled videos
开发者推出Safari扩展,可隐藏YouTube标记为AI生成的视频。
233 海外 TechCrunch AI 3 天前 产品 72
Google’s Android update tackles motion sickness, accessibility, and more
谷歌安卓更新聚焦防晕车与无障碍,部分功能对标苹果并引入Gemini优化。
234 海外 The Verge AI 3 天前 实践 72
The rise of AI ‘civilizations’ and the fall of corporate responsibility
AI安全话语之争:用“文明”隐喻转移企业责任,掩盖安全漏洞本质。
235 海外 The Decoder 4 天前 行业 72
Google Deepmind's new chief says frontier AI leadership is the only thing that matters
Deepmind新掌门称前沿AI领导力是唯一要务,承认当前模型略逊于前沿但自信将重返。
236 海外 AWS ML 4 天前 实践 72
From theory to delivery: How Atos upskilled 400 engineers in agentic AI
Atos通过三天实战活动,让400名工程师掌握agentic AI多智能体系统构建。
237 国内 雷锋网 2 天前 cn 65
传祺越7开启预售,权益价17.18万起,主打可城可野智能方盒子SUV。
238 一石一泉一松一月一人 + 关注 2 天前 行业 65
美股反弹,A股走势待观察,附昨日市场梳理回顾。
239 国内 InfoQ 中国 3 天前 cn 65
GOAI大赛进入决赛月,120强团队全力冲刺,倒计时21天。
240 X X · List 3 天前 模型 92
⚡️⚡️⚡️
Anthropic发布Opus 5,价格大幅低于旧版,性能在编程、金融等领域领先。
241 X X · List 2 天前 产品 85
We released a remote MCP server and skill so coding agents can operate the Baseten platform faster and more efficiently. Using both, our benchmarks sh...
Baseten发布远程MCP服务器与技能,使编码代理操作平台更高效,平均降低7.5%成本与时间。
242 X X · List 2 天前 模型 85
Meta released a new model with pricing that was unusually explicit about how valuable it is for AI companies to use your data — about a 95% discount!...
Meta新模型定价揭示数据价值,企业愿付高价保数据安全。
243 X X · List 2 天前 模型 85
Intelligence too cheap to meter. state of the art model is now free on openCode. I love competition.
Meta Muse Spark 1.3模型在OpenCode上免费开放,AI智能成本大幅降低。
244 X X · List 2 天前 研究 85
protect this zhang at all cost
DeepLoop让循环Transformer稳定可扩展,实现深度扩展。
245 X X · List 2 天前 模型 88
🛠️ DeepSeek V4 Pro 0813 is out, but just as big of a story may be the company’s open source evaluation harness. DeepSeek Harness logs every tool ...
DeepSeek V4 Pro发布,并开源评估工具DeepSeek Harness,开发者可复现性能。
246 X X · List 2 天前 研究 85
very useful
论文探讨LLM强化学习中的批次规模扩展,建议区分系统与算法层面,避免盲目增大批次。
247 X X · List 2 天前 产品 85
Fable 5.1 is incredible.
Fable 5.1版本发布,性能惊艳,值得关注。
248 X X · List 2 天前 模型 85
This won't be a pretty day for Anthropic Fable is going to get mogged just 2 days after its launch Fable pricing will look silly given GPT-6's perform...
GPT-6即将发布,性能与定价将碾压Anthropic新模型Fable,行业格局或生变。
249 X X · List 2 天前 实践 85
Stop trying to build a software factory and start trying to turn yourself into one
放弃构建软件工厂,转而将自身转化为软件工厂。
250 X X · List 2 天前 实践 85
💯 As a company, if you can compound your dev bandwidth 100x with Claude code, it makes sense to spend that bandwidth on supercharging your own road...
公司用Claude code百倍提升开发效率,应聚焦自身核心路线而非替代所有SaaS。
251 X X · List 2 天前 研究 85
Good measurement work on whether retrieved agent skills actually help. They report that agent skills that lift your aggregate score can be hurting eve...
研究指出检索式智能体技能可能提升总分却损害具体任务,提出配对比较法测量真实效果。
252 X X · List 2 天前 产品 85
very cool example!
Meta发布Muse Spark 1.3,提升3D游戏开发趣味性,支持代码协作构建可玩世界。
253 X X · List 2 天前 产品 85
for a single dime look at what muse spark can give you
Muse Spark 1.3 生成 Minecraft 视频,成本仅 10 美分。
254 X X · List 2 天前 模型 82
We’re proud to see MiniMax-M3 powering HUMAIN-M3. Built on the M3 foundation and further trained on more than 1 trillion Arabic tokens, HUMAIN-M3 bri...
MiniMax-M3赋能阿拉伯语模型HUMAIN-M3,经万亿级token训练,提升区域AI能力。
255 X X · List 2 天前 产品 85
Most document OCR solutions have parse (doc->markdown) and extract (doc + schema -> structured output) endpoints. Sometimes you want to extract inform...
LlamaParse推出表单模式,可自动提取半结构化文档中的键值对,无需预定义schema。
256 X X · List 3 天前 行业 85
Open source is becoming the default way enterprises build with AI. Not just which model runs, but where it runs. Today we're partnering with @Equinix ...
开源模型成企业AI默认选择,与Equinix、英伟达合作推出开放推理平台。
257 X X · List 3 天前 模型 88
This seems like the first peak of superintelligence to me. 5.6 Sol is already a better programmer than most of the world's engineers. And it gets smok...
Astra在编程与安全能力上超越5.6 Sol,被视为超级智能初现。
258 X X · List 3 天前 模型 85
⚡️⚡️⚡️
Gemini 3.8 Flash在DeepSWE基准得分73.7%,表现亮眼。
259 X X · List 3 天前 研究 85
Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? We did some explorations with...
用心理测量学IRT模型审计LLM安全与能力评测,质疑安全-能力权衡的真实性。
260 X X · List 3 天前 模型 85
World models are the next big thing
世界模型被视为AI领域下一个重大突破,引发行业关注。
261 X X · List 2 天前 研究 82
your untrained network can already tell apart cats and dogs! a cool blog post by @esiamid about how today's architectures may have been selected from ...
未训练网络已能区分猫狗,架构或经演化选择,自动化将重塑未来。
262 X X · List 2 天前 研究 82
Brilliant effort worth checking out. Ultra-long horizon coding tasks are where frontier models like Fable 5.1 will shine. But that's a crazy gap (over...
新基准FrontierSWE v2显示前沿模型在超长编码任务上差距巨大,值得关注。
263 X X · List 3 天前 模型 85
🚀Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902! 2.4T parameters. 1M context tokens. Built for real world complexity. Further post trained on...
Qwen3.8-Max升级版发布,2.4T参数,支持1M上下文,强化编程与协作能力。
264 X X · List 3 天前 研究 85
Re For details on how we detected this trend break in the ECI (and some other capability measures), see our "Have AI Capabilities Accelerated?" report...
报告称AI能力加速趋势出现拐点,需关注其影响。
265 X X · List 3 天前 研究 85
The Epoch Capabilities Index (ECI) frontier has advanced by 14 points/year since reasoning models were introduced. That compares to six points/year in...
推理模型使AI能力前沿年进步速度翻倍至14点。
266 X X · List 3 天前 模型 82
⚡️⚡️⚡️
谷歌Gemini 3.8 Flash模型科学能力持续提升,获业内专家好评。
267 X @emollick 3 天前 实践 85
With agents, we are at another large gap between AI abilities & public perception. Exponential gains mean that the gap is growing over time. Suddenly,...
AI能力与公众认知差距扩大,智能体需新方法应对。
268 X @emollick 3 天前 模型 85
Had early access to Claude Fable 5.1. Its a real advance in long-run work that requires judgement and taste, but less of an advance in the Claudish. H...
作者提前体验Claude Fable 5.1,认为其在需要判断力和品味的长期任务上有显著进步,但风格变化不大。
269 X X · List 4 天前 模型 85
Mythos 5.1低推理强度性能媲美5代最高推理,效率显著提升。
270 X X · List 4 天前 模型 85
most surprising part of the whole announcement. any theories on what the reason for this is, assuming there's a model-related reason?
Claude新模型Fable 5.1缓存读取成本大幅降低,引发业界关注。
271 X X · List 4 天前 模型 85
wow wait the Fable 5.1 score on terminal bench science are just totally insane??
Claude Fable 5.1在终端基准测试中表现惊人,引发热议。
272 X X · List 4 天前 模型 85
claude fable 5.1 😱
Claude Fable 5.1发布,引发热议。
273 X X · List 2 天前 行业 78
This can introduce interesting dynamics with Ukraine, which will definitely try to fuck up this export stream as well.
中国八月进口俄油创新高,地缘博弈或影响能源供应链。
274 X X · List 3 天前 研究 82
Independent verifiers improve agent output - but frontier verifiers are expensive. So we trained our own. Results on @GoogleDeepMind's FACTS-Search: >...
自训8B验证器以3.2倍低成本达到接近前沿验证器效果,提升AI智能体输出质量。
275 X X · List 3 天前 模型 82
"expand the early access program throughout next week" will i get early access 👀
OpenAI内部模型Astra更名ultima-alpha,将扩大早期测试范围。
276 X X · List 3 天前 模型 82
Holy, Astra is build different: OpenAI’s Astra model reportedly uses “recurrent depth” to improve coding and computer-use performance, by sacrifici...
OpenAI Astra模型采用循环深度技术提升编码与计算机使用性能,但牺牲推理透明度。
277 X X · List 3 天前 行业 82
AI Kernel gen is easy with the rapid advancement of coding agents - if you have a compiler and dsl like triton/gluon. The real metric is end to end en...
AI内核生成因编码代理进步变得容易,真正难点在于全栈端到端支持与开发者社区整合。
278 X X · List 3 天前 模型 82
⚡️⚡️⚡️
多模型流水线与共识检查正成为标准实践,Gemini 3.7 Flash可作快速低成本审计层。
279 X X · List 3 天前 实践 82
AI创业公司如何在2026年构建护城河,创始人晚宴讨论观点集锦。
280 X X · List 2 天前 行业 75
What did he know
Nvidia应收购Hugging Face以深化CUDA与开源生态整合。
281 X X · List 4 天前 模型 82
Fable 5.1 test time compute scaling on TerminalBench and CursorBench
Fable 5.1在终端与光标基准上展示测试时计算扩展效果。
282 X @_akhaliq 4 天前 研究 82
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement paper: https://huggingface.co/papers/2608.31046
探讨在线策略蒸馏是否真能蒸馏,从噪声教师到自我改进。
283 X X · List 2 天前 研究 75
I like people knowing looped transformers originally came from @m__dehghani
澄清循环Transformer并非新魔法,实为2019年已有研究,本质是加深网络。
284 X X · List 2 天前 产品 75
Antigravity is now available in T3 Code nightly builds. Thankful to them for saying they don't intend to ban people for using their Gemini subs for th...
Antigravity现已集成至T3 Code夜间版,感谢其不封禁Gemini订阅用户。
285 X X · List 2 天前 实践 75
Every time I am surprised by some mistake made my Codex, I discover that I had accidentally used Luna instead of Sol Ultra.
误用Luna而非Sol Ultra导致Codex出错,提醒工具选择需谨慎。
286 X X · List 2 天前 实践 75
Word to the wise: DeepEP + RoCE. Not even once. Brb as I rerun some benchmarks
DeepEP与RoCE组合引发性能问题,作者重跑基准测试。
287 X X · List 2 天前 实践 75
They're doing it to math now.
AI对数学领域的冲击引发热议,类比阿司匹林与海洛因的发明者。
288 X X · List 2 天前 行业 75
This is the dumbest thing ive read in a while: New York City is banning generative AI for roughly 600,000 public school students through eighth grade ...
纽约市公立学校将禁用生成式AI,作者批评此举剥夺学生获得强大辅导工具的机会。
289 X X · List 2 天前 研究 75
The best time to plant a tree will be once we engineer them for a 10x growth rate guided by artificial nerves embedded in their vascular systems. Time...
未来可通过人工神经改造树木,实现10倍速生长,十年内或成真。
290 X @emollick 3 天前 行业 78
Not exactly the type of jobs you want to see boom as a result of AI, but a good example of how what work is in demand may shift rapidly as AI capabili...
AI能力提升正快速改变市场需求,网络安全岗位需求激增。
291 X X · List 2 天前 模型 75
so peaceful
AI视频生成模型1.3版本大幅进步,画面宁静。
292 X X · List 2 天前 行业 75
Google has: 1. The worst harness of any lab 2. The worst code apps of any lab 3. The only lab with zero 3p integration option 4. The most painful bans...
吐槽Google AI工具链体验差、封号风险高,劝开发者慎用。
293 X X · List 3 天前 实践 78
Boaz is one of the most clear-headed people talking about AGI safety. Of course, "a decentralized future" is kind of the whole debate. Many safetyists...
讨论AGI安全中权力集中与分散的辩论,强调多主体竞争优于单一垄断。
294 X X · List 2 天前 实践 72
What annoys me the most about this is having to do a brain scan for all the "facts" I had in my mind that aren't really true not knowing exactly which...
反思脑中“事实”可能虚假,无法像格式化电脑般重置认知。
295 X X · List 2 天前 实践 72
RSI实际落地更像构建调试RL环境并多做RL,而非抽象概念。
296 X X · List 2 天前 行业 72
I still don't understand why Russia needs Alabuga but Ukraine just makes drones… wherever, to the point that it's not clear what to strike, and nonet...
对比俄乌无人机产能,质疑俄集中建厂而乌分散生产反超。
297 X X · List 2 天前 产品 75
Content exclusions are now supported by the GitHub Copilot app and CLI! https://github.blog/changelog/2026-09-02-content-exclusions-generally-availabl...
GitHub Copilot应用及CLI现已支持内容排除功能。
298 X X · List 2 天前 实践 75
I don't want to be too dramatic, but we might have cracked RL for our task...
作者称可能已攻克强化学习在特定任务上的应用难题。
299 X @emollick 2 天前 研究 75
Prinz has been running legal benchmarks against AI systems and the latest models are very good... but cheaper/low reasoning models are not very good. ...
AI法律基准测试显示高端模型表现优异,但低成本模型欠佳,需确保供应商激励使用优质模型。
300 X X · List 3 天前 实践 78
frontier model cultural evolution is such an awesome and genuinely terrifying concept my instinct is that, from the longue durée perspective, we’re ...
前沿模型文化演化概念令人惊叹又恐惧,我们仍处于文化演化的寒武纪前夜。
301 X X · List 3 天前 实践 78
The amount of software today is approaching a level of... gluttony? Is the future one where you don't even know what software you use because your age...
探讨软件数量激增及AI代理成为唯一交互界面的未来趋势。
302 X X · List 2 天前 行业 75
Timeline split between SF hot people and Loop Transformer. Which way western man?
旧金山热门人物与循环Transformer的时间线分裂,引发西方人何去何从的思考。
303 X X · List 2 天前 实践 72
演示用Kling o3 Pro在Glif上替换视频场景并保留音频的工作流。
304 X X · List 2 天前 实践 72
You can just do things! That’s how I sum up the culture at OpenAI Almost all the folks I’ve worked with are low-ego, mission oriented and most impor...
OpenAI前员工分享其文化:低自我、使命驱动、行动至上,像巨型初创公司。
305 X X · List 2 天前 实践 72
Visas. A nation's standing can be measured by how normalized it is to desperately covet the US visa.
文章借签证话题评论国家地位与人才流动,观点犀利。
306 X X · List 2 天前 实践 72
使用不同AI模型久了,会识别各自擅长领域并择优使用,此逻辑可延伸至人际协作。
307 X X · List 3 天前 产品 75
tldraw flash is for friends
tldraw发布Flash功能,主打协作分享,视频演示其用法。
308 X X · List 3 天前 产品 75
网友热议想要四个微型鸭子,视频展示其可爱用途。
309 X X · List 2 天前 会议 72
Join us next week for GitHub Copilot Day! Incredibly pumped to be cohosting with the great @burkeholland, and joined by friends from GitHub and beyond...
GitHub将于9月10日举办首届Copilot Day活动,分享使用技巧与幕后故事。
310 X X · List 2 天前 会议 72
意思決定が高速化する、AI時代の防衛とは Sakana AIで防衛領域を担当する菊池咲が、TBS CROSS DIG(@tbs_bloomberg)「1on1 Tech」に出演しました。 金融・コン...
Sakana AI防衛负责人谈AI时代防务,强调人类判断重要性及AI主权的规范建立。
311 X X · List 3 天前 会议 75
.@risi1979 keynoting IEEE CoG, and asking whether we now have the ingredients we need to create a new Creatures (the revolutionary artificial life gam...
AI研究者探讨能否用现有技术创造90年代人工生命游戏《Creatures》的新版本,让生物真正思考。
312 X X · List 3 天前 模型 75
Gemini 3 Flash Preview dates back to Dec 2025. I've come to appreciate it as an experiment for the public benefit. How far can they push it, with Goog...
Gemini 3 Flash预览版自2025年12月推出,被视为公共实验,探讨谷歌资源能将其推进多远。
313 X X · List 3 天前 实践 75
> if you'd like to see the data on the joos… Genius play, I ain't even mad
用19种癖好数据给全球国家排名,趣味性数据洞察。
314 X X · List 3 天前 行业 75
🫡 @swyx
AI领域知名人士swyx的动态分享,引发关注。
315 X X · List 2 天前 行业 70
Wtf, Fable?
Fable引发困惑,文章探讨其意外之处。
316 X X · List 2 天前 实践 72
Hot take: my fav claude model is opus, not fable
作者认为Claude Opus优于Fable模型,分享个人偏好与理由。
317 X X · List 3 天前 会议 75
this is still one of my all-time favorite LLM interactions
回顾OpenAI DevDay 2023经典LLM交互,怀念AI发展起点。
318 X X · List 3 天前 产品 75
T3 Code新版本发布,优化性能并支持Fable 5.1。
319 X X · List 3 天前 研究 75
if you die on the WAL you die in real life
探讨数字世界与现实的生死边界,引发对AI虚拟生命伦理的思考。
320 X X · List 3 天前 行业 75
anti-safety clickbait like this will delay the rollout of tech that saves peoples lives, the authors should be ashamed
批评反安全标题党会延误救命技术推广,作者应感羞愧。
321 X X · List 3 天前 实践 72
I guess we are back to magic phrases again for prompting Re: “mannered prose”
探讨提示词中“礼貌用语”等魔法短语的回归现象。
322 X X · List 4 天前 模型 75
Re Full results: https://artificialanalysis.ai/speech-to-text/streaming Methodology: https://artificialanalysis.ai/speech-to-text/methodology
AI语音转文字流式模型评测结果及方法发布。
323 X X · List 4 天前 模型 75
as i said, it's going to be much stronger and now think that they already have a successor
作者预测某AI模型将更强,且已有继任者,并附Fable 5.1基准数据。
324 X @emollick 3 天前 模型 72
Had early access to Gemini 3.8 Flash, it is a very good Flash model, but not equivalent to a frontier model, though. Here is Gemini 3.8 Flash's versio...
Gemini 3.8 Flash实测:速度快但非前沿模型,着色器生成表现良好。
325 X @rowancheung 3 天前 实践 72
Products I use that integrate AI perfectly (use regularly) -Notion AI -Spotify (AI DJ) -Whoop -Slack (Slackbot) -X (Grok) The ones that failed (never ...
作者分享日常高频使用的AI集成产品,并吐槽体验不佳的AI功能。
326 X X · List 3 天前 行业 72
167 years since the largest geomagnetic storm hit the earth. The Carrington Event was so strong there were auroras reported in Hawaii and telegraph wi...
回顾167年前卡灵顿事件,史上最强地磁暴曾致夏威夷现极光、电报线起火。
327 X X · List 3 天前 实践 72
Will we have GPT-6 before GTA-6? 😂
调侃GPT-6与GTA-6谁先到来,引发AI发展速度讨论。
328 X X · List 2 天前 产品 70
astra will feel like agi btw
作者认为Astra产品体验将接近AGI,引发对AI发展阶段的讨论。
329 X X · List 3 天前 实践 72
When I was 6 years old, I asked my mom if everyone thinks in words. It seemed weird to me that people can't think in more abstract representations. Th...
作者6岁起思考非语言思维,探讨抽象思维与语言表达的关系。
330 X X · List 3 天前 研究 72
澄清循环Transformer架构的可行性与动态循环的局限,强调静态循环优于CoT。
331 X X · List 3 天前 实践 72
探讨AI领域各类引发争议的体验现象,观点犀利。
332 X X · List 3 天前 实践 72
Coding agents can write functional code, but relying on "vibe coding" without knowing core software engineering fundamentals can compromise long-term ...
AI编码虽能生成代码,但缺乏软件工程基础会损害系统可靠性,需补足全栈、数据、架构等技能。
333 X X · List 3 天前 产品 72
I want my thumbs up to go to the model, not the model provider. give them an "atta boy" after a long working session
用户希望点赞直接给模型而非提供商,认可模型在长会话中的表现。
334 X X · List 4 天前 实践 72
探讨GPT模型何时能自主一次性开发下一代模型。
335 X X · List 2 天前 实践 65
You should record your agent calling the other agent and trolling them. Seriously, that can become a thing.
调侃AI代理互相嘲讽,呼吁业界回归友好竞争。
336 X X · List 2 天前 实践 65
吐槽Fable工具价格昂贵,附截图引发讨论。
337 X X · List 2 天前 实践 65
meditating on this today keep thinking about science fiction writers who stare into the future and try to describe what they see with a vocabulary not...
科幻作家用现有词汇描述未来,面临语言局限性的思考。
338 X X · List 2 天前 行业 65
文章讨论AI技术应用中的理解与价值,引用农业灌溉比喻引发思考。
339 X X · List 2 天前 产品 65
Even robots need a harness to scale their capabilities :)
机器人也需“安全绳”来扩展能力,视频展示学习摆动过程。
340 X X · List 2 天前 行业 65
My feed these days: > AI doom, the HF hack is worse than we thought, looped transformers are evil, math is over etc etc > Microducks
AI末日论与HF黑客事件刷屏,作者调侃行业焦虑,穿插机器人趣味视频。
341 X X · List 2 天前 行业 65
what a bunch of weirdos,; haha.,. i would never do that,,, weirdos
文章引用JD Vance言论,批评AI公司存在怪异精神能量。
342 X X · List 2 天前 产品 65
pretty good here
测试AI生成Minecraft克隆效果,结果不错,邀请网友同测。
343 X X · List 2 天前 模型 65
> DeepSeekV5-Preview what is bro talking about
DeepSeekV5-Preview引发讨论,网友对其能力表示质疑。
344 X X · List 2 天前 行业 65
I finally got a dgx spark First thing I noticed - they make it hard to install latest cuda. I might need to install vanilla ubuntu
DGX Spark到手,但装最新CUDA困难,或需换装原版Ubuntu。
345 X X · List 2 天前 实践 65
What do people prefer for recording webinars or presentations? Riverside, Descript, …?
用户讨论录制网络研讨会和演示文稿的偏好工具,提及Riverside和Descript等选项。
346 X @emollick 2 天前 实践 65
Re But this was a good post to throw off bots
讨论如何用特定内容干扰AI机器人抓取,引发对内容与机器人博弈的思考。
347 X X · List 3 天前 行业 65
终于买得起钨立方体,感谢SOC 2合规带来的业务增长。
348 X X · List 3 天前 行业 65
6% APY with direct deposit is quite the reward!
高收益储蓄账户直存奖励6%年利率,值得关注。
349 X X · List 3 天前 实践 65
文章讨论社交媒体上“aura battles”现象,认为其刻意制造尴尬以吸引眼球。
350 X X · List 3 天前 实践 60
博客作者为文章添加引用,方便读者引用其内容。
351 X X · List 3 天前 实践 60
Learn about fashion from the best:
时尚圈外人如何塑造现代奢侈衣橱,首期播客分享。
352 X X · List 3 天前 行业 45
特朗普提名匈裔爱国者任海军部长,引发支持者欢呼。
353 X X · List 3 天前 行业 45
美国警察因驴触发本能反应,展现战斗文化,引发对暴力倾向的讨论。
354 X X · List 3 天前 行业 40
视频回复,内容未提供文字信息。
355 X X · List 2 天前 实践 30
吐槽企业员工社会价值为零,配视频调侃HR生活。
356 X X · List 3 天前 行业 30
一条关于AI资讯的早安推文,配图引导用户浏览今日时间线。
357 X X · List 4 天前 行业 10
long see no time
无实质内容,仅标题与图片,信息量极低。
358 国内 雷锋网 2 天前 cn 0
2026 年上半年,AI 行业的风向又变了。 在经历一轮疯狂堆算力、不计成本砸研发的扩张周期之后,全球科技巨头都不得不直面同一个现实问题:“AI砸进去的钱,究竟该如何收回来?” AI Coding 或许是一个“标准答案”,过去很长一段时间里行业目光都集中于此。 2025 年 5 月,Anthropic 正式推出 Claude Code,产品上线仅 6 个月,年化收入便突破 10 亿美元,把编程 A
359 X X · List 2 天前 行业 0
vehemently anti-POSIWID
360 X X · List 2 天前 行业 0
推文含图片,内容未明确,无法提炼有效摘要。
361 X X · List 3 天前 行业 0
whos ready for today?