A 10k-Token Benchmark That Never Landed
A truncated r/LocalLLaMA post tests an unidentifiable model called IQ3_S for throughput. No TPS number, no hardware, no licence. What the gaps tell you is more useful than the one number that might have been.
What happened
On 4 October 2026 a user posted on r/LocalLLaMA describing a roughly 10,000-token generation test of a model they call "IQ3_S," with llama.cpp suggested as the inference backend in the post title. The thread is truncated before any tokens-per-second figure is reported.
Context
Local LLM inference through llama.cpp is a well-established community workflow, and informal throughput tests on that subreddit are routine. What makes this particular post slightly unusual is the embedded question the model was asked to answer: it dismisses Kimi K2 and GPT OSS 120b as "outdated" and frames sycophancy as a selection axis for open-weights LLMs used in creative-assistant and agentic-coding roles. The author also refers to the model under test as an unnamed "miracle engine everyone is talking about," a phrase that signals a specific release the community is mid-hype-cycle on. The model's own thinking trace, visible in the post, references a system date of 22 June 2026 and a tool-calling function called get_datetime, placing it among the newer agentic-capable checkpoints rather than a bare text generator.
How it works
What the post confirms: the author generated approximately 10,000 tokens in a single pass, with tokens-per-second as the stated measurement goal. The model accepted a system prompt that included a current date and exposed a get_datetime tool, meaning the harness was configured for tool-calling rather than plain autoregressive completion. The "IQ3_S" label is not expanded anywhere in the source; the vendor, parameter count, architecture, and quantisation level are all unknown. No GPU model, VRAM figure, CPU specification, sampling parameters, or context-length setting is stated. The TPS measurement was the purpose of the run; the result is absent from the post as published, and the phrase "10k tokens late" at the tail end suggests the text was cut off mid-sentence rather than concluding.
Our read
The honest read is that this is a work-in-progress benchmark that never shipped its number. The author set up the test, ran the tokens, and the thread died before the result landed. What the post does confirm, even incomplete, is a shift in how the community evaluates local models: the question is no longer "can it generate coherent text" but "can it sustain a 10k-token agentic task with tool-calling at a pace that makes it usable on the hardware I already own." The embedded sycophancy question is the more interesting signal. Users are filtering local checkpoints not just by perplexity or leaderboard scores but by behavioural temperament under open-ended creative prompting — whether the model pushes back on a weak prompt or sycophantically agrees. That is a criterion no vendor benchmark surfaces, and it only appears in a thread where someone has spent an afternoon reading raw output. The reticence around the model's name — "miracle engine everyone is talking about" instead of a direct callout — is itself a small data point: the release is either proprietary, under a non-disclosure constraint, or so recent that naming it felt premature. We cannot confirm which from this source. What we can say is that the community's informal evaluation loop is maturing into something closer to a QA pipeline, and that loop is currently running on Reddit threads and llama.cpp rather than on a dashboard.
What this changes
Nothing in the studio's pipeline changes on Monday. No ComfyUI node, no video-generation model, no hardware purchase is affected by a truncated community post that reports no TPS, names no hardware, and states no licence. If the studio runs local LLMs for script drafting or agentic QA of generated assets, the single transferable note is the sycophancy axis: when comparing candidate models for a copywriting or script-critique step, "does this model push back on a weak prompt, or does it agree with everything?" is now a legitimate question to ask alongside throughput and licence. That is a general observation the community is converging on, not a recommendation testable from this one post.
License
No licence is stated anywhere in the source material. The identity and weight licence of the model called "IQ3_S" are unknown; the reader should check the original model card or release repository before building anything commercial on it. An informal community post is not a licence grant, and the absence of one here is expected but worth flagging explicitly.
Key takeaways
- A user on r/LocalLLaMA ran a ~10,000-token generation of a model called IQ3_S to measure tokens-per-second; the post is truncated and no TPS figure is reported.
- The model's thinking trace shows tool-calling capability (get_datetime) and a system date of 22 June 2026, indicating an agentic-capable checkpoint with a custom system prompt.
- The vendor, parameter count, architecture, and quantisation level of "IQ3_S" are not identified in the source; no hardware specification is given.
- An embedded question in the post frames sycophancy as a selection criterion for local LLMs in creative and agentic-coding workflows, a criterion absent from standard vendor benchmarks.
- No licence is stated, no performance number is reported, and no hardware requirement is documented; nothing in this post warrants a change to a production pipeline.
Sources
- Need maybe say "Use llama.cpp" — tier 3
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 1
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20261004T125822Z