GPT-5.6 Luna Gets an 80% Price Cut: Output Falls to $1.20 per 1M Tokens
OpenAI cuts GPT-5.6 Terra by 20% and Luna by 80%. Luna now costs $0.20 input and $1.20 output per 1M tokens, while Sol pricing and subscription plans stay unchanged.
24 articles
Compare large language models, inference approaches and deployment requirements through model analyses and practical guides.
Start with your task, context requirements and hardware budget, then follow the relevant model analysis and deployment guide. Compare quantization, inference cost and practical limitations alongside benchmark results.
OpenAI cuts GPT-5.6 Terra by 20% and Luna by 80%. Luna now costs $0.20 input and $1.20 output per 1M tokens, while Sol pricing and subscription plans stay unchanged.
Kimi K3 launches with 2.8 trillion parameters, native multimodality, a 1M-token context window, and API pricing well below GPT-5.6 Sol and Claude Fable 5. Here is what matters.
Qianwen, Doubao, Iflytek in China. OpenAI, Google globally. But what are the actual rules of this war? What does winning mean? Will voice input methods still exist in five years? Everyone might be fighting the wrong battle.
DeepSeek-V4 released with 1M context as standard! Agent capabilities surpass open-source models, reasoning performance rivals GPT-4o and Claude Opus.
DeepSeek-V3.2 is the latest open-source LLM rivaling GPT-5 and Gemini-3.0-Pro, featuring DSA sparse attention and gold-medal results in IMO/IOI 2025.
vLLM now supports Xiaohongshu's dots.ocr for free multilingual OCR. Follow our tutorial to deploy this powerful model in two steps for parsing text, tables, and formulas.
Qwen3-Next series: a hybrid architecture with Gated DeltaNet × Gated Attention. 80B total parameters with ~3B active per step, optimized for long context, high concurrency, and low latency. Instruct and Thinking target production chat and deep reasoning respectively.
GLM-4.5 is Zhipu AI's new agent-native model, featuring efficient MoE architecture, strong reasoning, coding, and open-source ecosystem.
Kimi K2 by Moonshot AI: A trillion-parameter open-source model with MuonClip optimizer, advanced data rephrasing, and efficient sparse MoE architecture. Discover the core innovations and engineering behind this SOTA AI model.
Introducing Alibaba Cloud's Qwen3 LLM series (0.6B-235B), featuring MoE and Dense architectures for optimal performance and efficiency. Supports 119 languages with advanced coding and math capabilities.
OpenAI releases next-generation inference models o3 and o4-mini, demonstrating exceptional performance in coding, mathematics, science, and visual tasks. Both models show outstanding results in benchmark tests with major breakthroughs in safety and cost-effectiveness.
Meta Llama 4 faces ranking manipulation allegations while dealing with executive departures, putting its AI strategy to the test
Explore four pillars of modern AI: Agent, RAG, Function Call, and MCP. Learn how these technologies work together to enhance AI capabilities through practical examples and analogies.
DeepSeek-V3-0324 is the latest model upgrade from DeepSeek, achieving significant improvements in reasoning capabilities, frontend development, and content creation.
Mistral Small 3.1 vs Gemma 3: Compare 24B vs 27B parameters, performance benchmarks, hardware requirements, and real-world applications. Discover which lightweight LLM offers better efficiency and multimodal capabilities.
Alibaba Cloud unveils QwQ-32B, a groundbreaking 32B-parameter model challenging DeepSeek R1 (671B) through pure reinforcement learning, showcasing comparable performance in reasoning tasks.
Microsoft OmniParser V2.0 is a next-gen AI visual parsing tool that converts GUI to structured data, with faster speed, higher accuracy, and seamless LLM integration.
Learn to implement LLM-Reasoner framework for enhanced logical reasoning like DeepSeek R1. Step-by-step guide for building AI systems with advanced thinking capabilities.
DeepSeek R1, a breakthrough AI model using reinforcement learning, achieves human-level reasoning and matches OpenAI's o1-1217 performance
ByteDance open sources Eino, a Golang-based LLM application development framework providing stable and scalable development. With clear component definitions and process orchestration, it helps developers quickly build high-quality LLM applications with reliability and extensibility.
Learn how to deploy LangChain Open Canvas - a powerful open-source alternative to OpenAI Canvas for AI writing and coding. Step-by-step guide for installation and private deployment of your own AI platform.
CogAgent-9B: A 9B-parameter GUI agent by Zhipu AI and Tsinghua University that excels in interface understanding and automation, outperforming other models in MM-Vet and more benchmarks
In-depth comparison and analysis of popular AI model deployment tools including SGLang, Ollama, VLLM, and LLaMA.cpp, helping developers and users choose the most suitable AI model deployment tool
"Deep dive into DeepSeek-V3 model. Its architecture combines MLA and DeepSeekMoE with innovative load balancing. Trained on 14.8T tokens, powered by HAI-LLM framework and FP8 technology. Enhanced by innovations like MTP, performance surpasses open-source and approaches closed-source models. Cost-effective with low training and API costs, a key reference in AI advancing language models."