📊 Full opportunity report: What Is Inkling And How Does It Revolutionize Artificial Intelligence? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has launched Inkling, a large-scale open multimodal AI model on Hugging Face, capable of processing text, images, and audio. Its significant size and hardware requirements limit direct use but open new possibilities for research and domain-specific applications.

Thinking Machines has released Inkling on Hugging Face, a highly large-scale multimodal AI model designed for processing text, images, and audio. The 975-billion-parameter model aims to expand multimodal reasoning capabilities but requires substantial hardware resources, limiting its accessibility for most users. This release marks a significant step in open AI model availability and multimodal AI research.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters, of which 41 billion are active during processing, according to Hugging Face. It was trained on approximately 45 trillion tokens across multiple data types, including text, images, and audio. The architecture employs 256 experts, with a routing mechanism that activates six experts per input, and combines global and sliding-window attention patterns. This approach is discussed in detail in Welcome Inkling By Thinking Machines. The model supports a one-million-token context window, enabling extensive reasoning across multimodal data.

Deployment requires high-end hardware: Hugging Face reports that the BF16 checkpoint demands about 2 TB of VRAM, while the NVFP4 version needs roughly 600 GB. This highlights the importance of understanding AI hardware requirements, as detailed in the original analysis. Consequently, full operation on consumer hardware is impractical, with hosted inference services or specialized hardware being the primary options. The release includes support for popular inference frameworks such as Transformers, SGLang, and llama.cpp, and offers both BF16 and lower-precision NVFP4 checkpoints for inference.

While the model’s architecture and deployment options are detailed, crucial information remains undisclosed. No independent benchmark results, safety evaluations, or licensing terms are available yet, raising questions about its performance, safety, and usage restrictions. The model is positioned for domain-specific fine-tuning, especially in scientific, media, and enterprise sectors that handle multimodal data.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face, emphasizing its open access and multimodal reasoning capabilities.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Open Multimodal Architecture

The release of Inkling signifies a major advancement in multimodal AI, offering the potential to unify language, image, and audio reasoning within a single model. Its open availability could accelerate research and development in fields requiring complex data analysis, such as scientific research, media analysis, and enterprise applications. However, its enormous hardware requirements limit immediate practical use, making it primarily accessible through hosted services or specialized hardware setups. The model’s scale and architecture also push the boundaries of current AI capabilities, setting a new benchmark for future multimodal models.

PNY VCNRTXA6000-SB NVIDIA RTX A6000 Graphics Card 48GB GDDR6

PNY VCNRTXA6000-SB NVIDIA RTX A6000 Graphics Card 48GB GDDR6

NVIDIA Virtual PC (vPC)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal AI Development

Prior to Inkling, multimodal AI models have generally been smaller and more specialized, often requiring separate systems for text, images, or audio. The trend toward larger, more integrated models has accelerated over recent years, driven by advancements in hardware and training techniques. Notable examples include multimodal models from OpenAI and Google, but none have matched Inkling’s scale of 975 billion parameters. The open release by Thinking Machines on Hugging Face marks a shift toward more accessible yet resource-intensive models, reflecting ongoing industry efforts to unify multimodal reasoning capabilities.

Previous large models, such as GPT-4 and PaLM 2, have demonstrated multimodal capabilities but with less transparency about architecture and training data. Inkling’s release emphasizes openness and multimodal versatility, although details about licensing and safety remain sparse. The model’s training on 45 trillion tokens across diverse data types highlights the trend toward comprehensive, multi-input AI systems.

“This model is huge.”

— Hugging Face

MX3 M.2 AI Accelerator

MX3 M.2 AI Accelerator

High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

Key details remain unconfirmed, including independent benchmark results, safety evaluations, and licensing terms. The model’s actual performance on practical workloads, especially with video inputs, has not been evaluated or disclosed. Additionally, the cost and accessibility of hosting or fine-tuning Inkling are unclear, given the enormous hardware requirements.

AI Vision & Voice Interaction Robot for Arduino Scratch Python Programming 17DOF Humanoid Robot Large AI Model STEM Project Education Voice Command Walking Dancing Self-Stand Up, Tonybot Advanced kit

AI Vision & Voice Interaction Robot for Arduino Scratch Python Programming 17DOF Humanoid Robot Large AI Model STEM Project Education Voice Command Walking Dancing Self-Stand Up, Tonybot Advanced kit

【Humanoid Robot with ESP32】 Powered by ESP32 and 17 intelligent servos, Tonybot smart humanoid robot delivers smooth, dynamic…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing, Benchmarking, and Licensing Clarity

Developers and research teams are expected to begin testing Inkling through supported inference engines, which will clarify its real-world performance, latency, and accuracy. Independent evaluations, safety assessments, and benchmark results are anticipated to be published in coming months. Clarification on licensing, usage restrictions, and potential for domain-specific fine-tuning will also shape how the model is adopted and integrated into workflows.

NIMO AI NAS, Agentic Computer Mini PC and AI Server, AMD Ryzen 7 PRO 8845HS

NIMO AI NAS, Agentic Computer Mini PC and AI Server, AMD Ryzen 7 PRO 8845HS

Next-Gen AI & LLM Local Deployment: Powered by the 8845HS processor and RTX 5060 GPU, this NAS provides…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a multimodal AI model from Thinking Machines with 975 billion parameters, capable of processing text, images, and audio within a unified architecture.

Can Inkling process video?

While the architecture supports image inputs with a temporal dimension, native video processing has not yet been evaluated or confirmed by Hugging Face or Thinking Machines.

Can I run Inkling on my personal computer?

Probably not, as the required hardware—up to 2 TB of VRAM for BF16—exceeds typical consumer systems. Hosted inference services or specialized hardware are likely necessary for deployment.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing details, restrictions, or whether training code and data are publicly available.

What are the practical applications of Inkling?

Potential uses include scientific research, media analysis, and enterprise workflows that require reasoning across multiple data types, especially with domain-specific fine-tuning.

Source: ThorstenMeyerAI.com

You May Also Like

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Explore how Mistral’s focus on sovereignty, open weights, and enterprise control shifts the AI race. Is it a smart move or a sign of losing ground?

Drones in 2025: What Are They Being Used for Beyond Photography?

The transformative uses of drones in 2025 extend far beyond photography, revolutionizing industries and saving lives; discover how they’re shaping our future.

The Tech Taking Us to Mars: How SpaceX and Others Drive Space Innovation

Keen advancements in space tech are revolutionizing Mars travel, but the full story of these innovations and their impact is just beginning.

How to claim a WhatsApp username

Learn the steps to reserve your WhatsApp username, understand restrictions, and what it means for your privacy and security as Meta rolls out the feature globally.