<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LLM Comparison on Chengyu Wang</title><link>https://chengyu.eu/tags/llm-comparison/</link><description>Recent content in LLM Comparison on Chengyu Wang</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Chengyu</copyright><lastBuildDate>Sun, 08 Feb 2026 14:00:22 +0000</lastBuildDate><atom:link href="https://chengyu.eu/tags/llm-comparison/index.xml" rel="self" type="application/rss+xml"/><item><title>AI Benchmarks: Opus 4.6 vs. GPT-5.2 vs. Gemini 3 Pro, Head to Head</title><link>https://chengyu.eu/posts/ai-benchmark-opus-gpt-gemini/</link><pubDate>Sun, 08 Feb 2026 14:00:22 +0000</pubDate><guid>https://chengyu.eu/posts/ai-benchmark-opus-gpt-gemini/</guid><description>A benchmark chart making the rounds shows large models shifting from chatbots to agents — with Opus 4.6 dominating agentic computer use, GPT-5.2 still king of pure reasoning, and Gemini 3 Pro owning vision and multilingual work.</description></item><item><title>Head-to-Head: How Mainstream AI Models Judge a Real Traffic Accident</title><link>https://chengyu.eu/posts/ai-models-traffic-accident-liability/</link><pubDate>Thu, 15 Jan 2026 18:51:39 +0000</pubDate><guid>https://chengyu.eu/posts/ai-models-traffic-accident-liability/</guid><description>After my own Tesla was hit by a car pulling out from under a bridge, I ran the accident photo past ChatGPT, Gemini, Grok, Copilot, Claude, and a few domestic models to see which ones actually got liability right.</description></item></channel></rss>