{"provider_url":"https://hatena.blog","height":"190","provider_name":"Hatena Blog","title":"LLM\u3067MTP\u3092\u8a66\u3059","version":"1.0","html":"<iframe src=\"https://hatenablog-parts.com/embed?url=https%3A%2F%2Fdevops-blog.virtualtech.jp%2Fentry%2F20260706%2F1783299600\" title=\"LLM\u3067MTP\u3092\u8a66\u3059 - \u3068\u3053\u3068\u3093DevOps | \u65e5\u672c\u4eee\u60f3\u5316\u6280\u8853\u306eDevOps\u6280\u8853\u60c5\u5831\u30e1\u30c7\u30a3\u30a2\" class=\"embed-card embed-blogcard\" scrolling=\"no\" frameborder=\"0\" style=\"display: block; width: 100%; height: 190px; max-width: 500px; margin: 10px 0px;\"></iframe>","published":"2026-07-06 10:00:00","image_url":null,"author_name":"ytooyama","categories":["LLM","\u30ed\u30fc\u30ab\u30ebLLM","AI"],"description":"\u5c11\u3057\u51fa\u9045\u308c\u305f\u611f\u306f\u3042\u308a\u307e\u3059\u304c\uff64LLM\u306eMTP\u3092\u8a66\u3059\u3053\u3068\u306b\u3057\u307e\u3057\u305f\u3002 \u6b21\u306eOllama\u516c\u5f0f\u30a2\u30ab\u30a6\u30f3\u30c8\u306e\u6295\u7a3f\u3092\u898b\u305f\u306e\u304c\uff64LLM\u306eMTP\u3092\u8a66\u305d\u3046\u3068\u601d\u3063\u305f\u304d\u3063\u304b\u3051\u3067\u3059\u3002 Gemma 4 is now nearly 90% faster on Apple Silicon with Ollama using MLX!The speedup comes from improved multi-token prediction (MTP), now on by default for Gemma 4, with more models to come.Ollama automatically tunes how\u2026","author_url":"https://blog.hatena.ne.jp/ytooyama/","type":"rich","width":"100%","blog_title":"\u3068\u3053\u3068\u3093DevOps | \u65e5\u672c\u4eee\u60f3\u5316\u6280\u8853\u306eDevOps\u6280\u8853\u60c5\u5831\u30e1\u30c7\u30a3\u30a2","blog_url":"https://devops-blog.virtualtech.jp/","url":"https://devops-blog.virtualtech.jp/entry/20260706/1783299600"}