Nonobench v1.2: 43 LLMs, a Hard Ceiling No Open Model Clears
The new 20×20 Hard nonogram tier is a pass/fail boundary, not a percentage-point gap. The top of the board is closed-weight, and the reasoning traces are public.
What happened
Nonobench v1.2, posted to r/LocalLLaMA on September 27, tests 43 LLMs on nonogram logic puzzles and introduces a new 20×20 Hard mode. No open-weight model solves that Hard mode; DeepSeek V4 Pro, the top open-weight entry, ties for fourth.
Context
The January edition of Nonobench covered 23 models. v1.2 roughly doubles the field and adds the Hard tier. Nonograms are a clean proxy for constraint-satisfaction reasoning: each row and column carries a count constraint, and the solver must reconcile conflicting partial information without a generative shortcut. That makes the benchmark a useful stress test for chain-of-thought discipline in models that do not have a dedicated planning head. The explicit per-run reasoning-effort parameter is new in v1.2; the January edition did not separate it.
How it works
Nonobench feeds each model a nonogram grid with row and column count clues and asks it to fill the grid. In v1.2 the 20×20 Hard mode is the new ceiling: at that size the constraint space is large enough that a model cannot shortcut by pattern-matching smaller grids. Two design choices matter for anyone who wants to reproduce or inspect results. First, reasoning effort is an explicit per-run parameter, so a single model can be scored at different compute budgets. Second, every prompt and output is public, meaning the full reasoning trace for each of the 43 models is available for inspection. The source post does not state the scoring metric (accuracy, partial credit, or time-normalised) or the hardware required to reproduce a run locally. The full leaderboard beyond DeepSeek V4 Pro's 4th-place tie is truncated in the captured text.
Our read
The obvious reading is "open-weight models still trail closed ones on hard reasoning." That is true but incomplete. The more useful signal for a small studio is structural: the top of the board is closed-weight, and the gap at the Hard ceiling is not marginal. It is binary. No open model clears it at all. That is not a five-percentage-point deficit; it is a pass/fail boundary.
For a video-production pipeline that routes complex multi-step prompts through an LLM (ComfyUI workflow planning, multi-constraint asset QA, scheduling logic), the transfer is indirect. Nonograms are not video generation. But the underlying skill, reconciling many simultaneous constraints without a shortcut, is the same skill an orchestration layer needs when it plans a five-node render chain with conflicting resolution and aspect-ratio requirements. If the model cannot close a 20×20 grid, the question for Monday's workflow is whether it can hold a six-constraint prompt without dropping one.
The public prompts and outputs change the calculus in one practical way: you can pull the exact reasoning traces for the models that clear lower tiers and see where the open-weight models start losing coherence. That is more useful than any aggregate score.
What the post does not say, and the truncated leaderboard means we cannot verify, is whether the 1st-through-3rd positions are all closed-weight. The "no open model" claim at Hard implies it, but the 4th-place tie means the gap between "top" and "open" may be one or two positions, not a gulf.
What this changes
Nothing in the ComfyUI node graph, the model-loading path, or the render pipeline shifts because of this post. The relevant action is narrower: if your orchestration layer uses an open-weight model for multi-step reasoning (prompt routing, workflow planning, agent-based QA), treat the Hard-ceiling result as a caution, not a verdict. Before committing compute to a larger open-weight model for those tasks, run your own constraint-satisfaction probe. The explicit reasoning-effort parameter in v1.2 is a template: test whether increasing it on your own tasks changes output quality before you buy the tokens. And because every prompt and output is public, pull the traces for the models you are considering and read where they break. That costs an afternoon, not a GPU month.
License
The sources do not state a licence for the Nonobench v1.2 benchmark code, prompts, or outputs, nor for the DeepSeek V4 Pro weights specifically. The post describes DeepSeek V4 Pro as "open-weight" but does not name a licence. Check the model card and the benchmark repository before building anything commercial on either.
Key takeaways
- No open-weight model solves the 20×20 Hard nonogram; the top of the board appears to be closed-weight.
- DeepSeek V4 Pro ties for 4th among 43 models, the best open-weight result in the captured text.
- Reasoning effort is now an explicit per-run parameter, enabling compute-budget A/B testing on the same model.
- All 43 models' prompts and outputs are public, making trace inspection the cheapest way to evaluate reasoning quality.
- The benchmark measures constraint-satisfaction reasoning, not generation; its transfer to video-pipeline orchestration is indirect but relevant to multi-step planning tasks.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 1
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260927T150513Z