Gemma 4 26B A4B lands on Hugging Face: MoE, Apache 2.0, and three things to check before you load it
A 26B-parameter Mixture-of-Experts model from Google DeepMind's Gemma 4 family hits the hub under Apache 2.0. The specs are strong, but the upload's provenance and the MTP-drafter question matter more than the parameter count.
What happened
On 20 September 2026, user 7tianan published a checkpoint called gemma-4-26B-A4B-it-assistant to the Hugging Face model hub. The model card identifies it as a Mixture-of-Experts variant from Google DeepMind's Gemma 4 family, licensed Apache 2.0, with the hub task label set to "any-to-any."
Context
Gemma 4 ships in five sizes β E2B, E4B, 12B, 26B A4B, and 31B β mixing dense and MoE architectures. The 26B A4B is the MoE variant, positioned for consumer-GPU and workstation deployment rather than data-centre inference. The family adds a 256K-token context window, support for 140+ languages, and a hybrid attention mechanism that interleaves local sliding-window attention with full global attention. This upload sits at 0 likes and 0 downloads at the time of collection, consistent with a very fresh posting.
How it works
The 26B A4B is a Mixture-of-Experts model: roughly 26 billion total parameters with about 4 billion active per token, which is what makes single-GPU deployment realistic. It accepts text and image input and generates text output; audio input is limited to the E2B, E4B, and 12B variants. Attention layers alternate between a local sliding window and full global attention, with the final layer always global. Global layers use unified Keys and Values and Proportional RoPE (p-RoPE) for positional encoding.
The most operationally relevant feature is Multi-Token Prediction (MTP): a smaller draft model predicts several tokens ahead, and the target model verifies them in parallel. The card states this yields up to 3Γ decoding speedup with identical output quality to standard generation. The checkpoint uses the transformers library, safetensors format, and references arXiv paper 2607.02770.
Our read
Three things in this upload deserve scrutiny before anyone wires it into a pipeline. First, the provenance. The hub shows a non-Google account, 7tianan, as the uploader, yet the embedded model card is Google DeepMind's official Gemma 4 documentation, linking to ai.google.dev and github.com/google-gemma. The brief does not confirm whether this is a verbatim mirror of Google's official weights, a fine-tune, or a repackaged subset. The 0-download count means there is no community verification yet.
Second, the card carries a prominent note: "This model card is for the MTP drafters" β while the checkpoint name and parameter hint describe the full 26B A4B assistant. Whether the MTP drafter weights are inline or a separate download is not stated. A studio that loads this expecting the 3Γ decode speedup may find the drafter missing.
Third, the hub task label "any-to-any" contradicts the model card, which specifies text output only. If a ComfyUI graph is built on the assumption that this node can emit image or video tokens, it will not work. The label is a hub-side tag, not a capability guarantee.
What this changes
For a studio running a local ComfyUI pipeline, this is a text-generation model, not a media generator. The practical role is conditioning-text: scene descriptions, shot lists, VFX direction copy, prompt expansion for downstream video or image nodes. The 256K context window handles long multi-shot scripts in a single pass.
If the MTP drafter is inline, batch-generating that conditioning text at up to 3Γ decode speed with identical output is a genuine workflow gain. The catch: it requires speculative-decoding support in the inference stack, not just transformers.
Audio input is not supported on the 26B A4B variant, so audio-reactive workflows are out. No source in the brief states a minimum VRAM requirement; the ~4B active parameter count suggests 24β32 GB, but that is inference, not specification.
License
Apache 2.0. Commercial use is permitted without royalty or revenue ceiling. A studio can ship client work built on this checkpoint without additional licensing fees. No source in the brief documents training-data provenance or any usage restriction beyond the licence itself.
Key takeaways
- The 26B A4B is an MoE variant (~26B total, ~4B active) from Google DeepMind's Gemma 4 family, licensed Apache 2.0, and targeted at consumer-GPU deployment.
- The hub label "any-to-any" contradicts the model card's text-only output specification; do not wire this into a pipeline expecting image or video tokens.
- The provenance of the 7tianan upload is unconfirmed: the card is Google's official documentation, but the uploader is a third-party account with zero downloads.
- The MTP speculative-decoding path (up to 3Γ faster, identical output) is the most practical performance lever, but it requires drafter weights and speculative-decoding-capable inference infrastructure.
- Training-data provenance is not documented in any source; the model card does not name datasets.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 1
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260920T223933Z