Harness
your curiosity.
Research, case studies, builds, and playbooks on the technology shaping how we live.
On the bench
Llama 3.1
ComfyUI
Orpheus TTS
Whisper
Ollama
Flux.1
Mistral
Stable Diffusion
LM Studio
OpenCode
Qwen 2.5
DeepSeek
Claude Code
Llama 3.1
ComfyUI
Orpheus TTS
Whisper
Ollama
Flux.1
Mistral
Stable Diffusion
LM Studio
OpenCode
Qwen 2.5
DeepSeek
Claude Code
The question we keep asking
How are the tools we use daily changing us?
The question every piece traces back to.
01
Two Unsloth UD quant names, one GGUF file, and not a single IQ1 or IQ2 tensor inside
Note · 5 min
An Unsloth UD quant name describes the recipe, not the GGUF tensors. Two Nemotron builds are byte-identical under…
Read
02
Nemotron 3.5 Lightning’s 4-bit GGUF is 12 percent 16-bit, and the weight is a head most people never switch on
Note · 6 min
The Nemotron 3.5 Lightning GGUF from NVIDIA is 22.46 GB, and 2.67 GB of it is BF16. That…
Read
03
Your 4-bit vision model sees in 16-bit, and the best anyone offers is 8
Note · 5 min
An mmproj GGUF holds a whole vision encoder. Four publishers ship it unquantised at BF16 or F32, and…
Read
04
Three publishers, one schedule: the quantisation policy we called a choice is a line of llama.cpp
Note · 4 min
A llama.cpp quantisation rule decides which layers get more bits. One line predicts all 20 promoted blocks in…
Read
05
The speculative decoding head ships two different ways, and neither one is in the model you downloaded
Note · 4 min
An MTP draft model GGUF ships either as an extra block inside the model or as a separate…
Read
06
Experts really do go missing from MoE models, and the builds that remove them say so in the metadata
Note · 4 min
REAP pruned MoE builds remove a quarter of the experts and declare it. We read four GGUF tensor…
Read
Editor's pick
Stridenalysis
How AI companies hide tracking code, from Claude Code to your local tools
№ 033 · 5 min Read the piece
Stridenalysis
Brand voice for AI agents: the layer that beats a better model
№ 028 · 6 min Read the piece
Stridenalysis
Set up local LLM tool calling in Hermes and OpenCode on a Mac
№ 026 · 5 min Read the piece- № 058 Two Unsloth UD quant names, one GGUF file, and not a single IQ1 or IQ2 tensor inside Note
- № 057 Nemotron 3.5 Lightning’s 4-bit GGUF is 12 percent 16-bit, and the weight is a head most people never switch on Note
- № 056 Your 4-bit vision model sees in 16-bit, and the best anyone offers is 8 Note
- № 055 Three publishers, one schedule: the quantisation policy we called a choice is a line of llama.cpp Note