Categories

Engineering

Assessment engine, report pipeline, infrastructure — the choices we made and why.

Two filled on try four — Qwen3.8-27B 4bit on Terminal-Bench 89, pass@4 48 and the vanished ERRs
뇌파 곡선 일러스트

Two filled on try four — Qwen3.8-27B 4bit on Terminal-Bench 89, pass@4 48 and the vanished ERRs

Round 4 of Qwen3.8-27B MLX 4bit (17GB): restart only the server, rerun the 43 failures. Two recoveries reach pass@4 48/89 and the 13 round-3 ERRs drop to 1 — the server was never the cause.

  • #terminal-bench
  • #qwen
  • #mlx
  • #benchmark
Third try recovered nothing — Qwen3.8-27B 4bit on Terminal-Bench 89, pass@3 46 and the 13 ERRs
뇌파 곡선 일러스트

Third try recovered nothing — Qwen3.8-27B 4bit on Terminal-Bench 89, pass@3 46 and the 13 ERRs

Round 3 of Qwen3.8-27B MLX 4bit (17GB): re-running the 43 round-2 failures recovers zero, pass@3 stays 46/89. 30 failed exactly as before; 13 never ran a turn — and the real cause was not the server.

  • #terminal-bench
  • #qwen
  • #mlx
  • #benchmark
How much does a retry recover — Qwen3.8-27B 4bit on Terminal-Bench 89, pass@2 46 vs 53
뇌파 곡선 일러스트

How much does a retry recover — Qwen3.8-27B 4bit on Terminal-Bench 89, pass@2 46 vs 53

Round 2 of Qwen3.8-27B MLX 4bit (17GB): re-running the 54 round-1 failures reaches pass@2 46/89 (51.7%). 11 recoveries (median 6.5 min), two solved only by 27B, and why retrying failures is slower.

  • #terminal-bench
  • #qwen
  • #mlx
  • #benchmark
The price of a 6x smaller model — Qwen3.8-27B 4bit on Terminal-Bench 89, single pass 35 vs 45
뇌파 곡선 일러스트

The price of a 6x smaller model — Qwen3.8-27B 4bit on Terminal-Bench 89, single pass 35 vs 45

Qwen3.8-27B MLX 4bit (17GB) on Terminal-Bench 2.1, same harness as Flash Next (100GB): 35/89 (39.3%) vs stock 45 and Uncensored 41. Which tasks split, 15 repetition loops, speed, ops notes

  • #terminal-bench
  • #qwen
  • #mlx
  • #benchmark
Closing the table — stock vs Uncensored Qwen3.8 on Terminal-Bench 89, pass@5 64 to 59
시냅스 일러스트

Closing the table — stock vs Uncensored Qwen3.8 on Terminal-Bench 89, pass@5 64 to 59

Round-5 (pass@5) results for the abliterated Uncensored Qwen3.8 Flash Next on Terminal-Bench 2.1, plus a series wrap-up: 59/89 vs stock 64/89, the 13 tasks that split them, speed, and harness failures

  • #terminal-bench
  • #qwen
  • #mlx
  • #abliterated
Round five: the "saturated" curve bent up again — stock Qwen3.8 pass@5 64/89 (71.9%)
노트와 노트북이 놓인 상담 책상

Round five: the "saturated" curve bent up again — stock Qwen3.8 pass@5 64/89 (71.9%)

Round-5 (pass@5) Terminal-Bench 2.1 results for stock Qwen3.8 Flash Next mixed-4-8bit: 4 of 29 retries recovered (two timeout-with-the-right-file), a 120K context handoff and a fallback loop

  • #terminal-bench
  • #qwen
  • #mlx
  • #abliterated
Giving the other side the same number of tries — Uncensored Qwen3.8 pass@4 58/89 (65.2%), gap of 6
뇌파 곡선 일러스트

Giving the other side the same number of tries — Uncensored Qwen3.8 pass@4 58/89 (65.2%), gap of 6

Round-4 (pass@4) Terminal-Bench 2.1 results for the abliterated Uncensored Qwen3.8 Flash Next: 2 of 33 retries recovered, two fallback loops worth 3,877 HTTP 400s, and a 40-minute Coq build

  • #terminal-bench
  • #qwen
  • #mlx
  • #abliterated
Still ahead after three tries: stock Qwen3.8 pass@3 58/89 vs Uncensored best-of-3 56
뇌파 곡선 일러스트

Still ahead after three tries: stock Qwen3.8 pass@3 58/89 vs Uncensored best-of-3 56

Round-3 (pass@3) Terminal-Bench 2.1 results for stock Qwen3.8 Flash Next mixed-4-8bit, a like-for-like three-trial comparison with the abliterated build, and why 31 tasks still fail

  • #terminal-bench
  • #qwen
  • #mlx
  • #abliterated
Four tries, still ahead: stock Qwen3.8 pass@4 60/89 (67.4%) and the two tasks that flipped
시냅스 일러스트

Four tries, still ahead: stock Qwen3.8 pass@4 60/89 (67.4%) and the two tasks that flipped

Round-4 (pass@4) Terminal-Bench 2.1 results for stock Qwen3.8 Flash Next mixed-4-8bit: 2 recoveries out of 31 retries, and why 29 still fail — repetition loops, premature submits, near misses

  • #terminal-bench
  • #qwen
  • #mlx
  • #abliterated
Does abliteration boost coding? Uncensored vs stock Qwen3.8 on Terminal-Bench 89
뉴런 네트워크 일러스트

Does abliteration boost coding? Uncensored vs stock Qwen3.8 on Terminal-Bench 89

A same-conditions single-pass A/B between an abliterated (Uncensored) model and the stock mixed-4-8bit build across all 89 Terminal-Bench 2.1 tasks

  • #terminal-bench
  • #qwen
  • #mlx
  • #abliterated
Uncensored Qwen3.8 Flash Next: all 89 Terminal-Bench tasks
시냅스 일러스트

Uncensored Qwen3.8 Flash Next: all 89 Terminal-Bench tasks

Three days of Terminal-Bench 2.1 on a community abliterated derivative. Four rounds of context, timeout and temperature changes, 13 server wedges, and what a final 56/89 actually means.

  • #Terminal-Bench
  • #Qwen
  • #MLX
  • #Local LLM
Running Terminal-Bench overnight against Qwen3.8 Flash Next
Summary of an overnight Terminal-Bench run on Qwen3.8 Flash Next — 70/89 completed, 51.4 percent pass rate, timeouts the leading failure cause at 13, and four mlx-serve wedges all recovered by the watchdog

Running Terminal-Bench overnight against Qwen3.8 Flash Next

Measuring how much real terminal work a local LLM on an M5 Max can actually complete. The pass rate was less interesting than the server wedging four times and recovering itself every time.

  • #Terminal-Bench
  • #Qwen
  • #MLX
  • #Local LLM
Qwen3.8 27B — 4bit vs 8bit vs Flash Next, and is DFlash2 really lossless?
A comparison of Qwen3.8 Flash Next against 27B at 4bit and 8bit on the same prompts — average decode of 82.5, 65.5, and 18.1 tok/s respectively

Qwen3.8 27B — 4bit vs 8bit vs Flash Next, and is DFlash2 really lossless?

Measuring Qwen3.8 27B at 4bit and 8bit against Flash Next, then adding DFlash2 speculative decoding. It was 2–3x faster, but contrary to the official claim the output diverged from the original on M5.

  • #MLX
  • #Qwen
  • #Apple Silicon
  • #Local LLM
Running Qwen3.8 Flash Next on an M5 Max, and the fight with prefill
A before-and-after comparison of tuning Qwen3.8 Flash Next on an M5 Max — cache reuse from 29% to 99.7%, prefill throughput from 354 to 962 tok/s

Running Qwen3.8 Flash Next on an M5 Max, and the fight with prefill

Running a 125B MoE model on one Mac as an agent backend. I assumed the model was slow. A day of digging showed the real culprit was prefill — and how caching and a few flags fixed it.

  • #MLX
  • #Qwen
  • #Apple Silicon
  • #Local LLM
[postgresql] pg_wal that will not shrink — start with replication slots
Bar lengths contrasting a 650MB database against an 18GB pg_wal directory

[postgresql] pg_wal that will not shrink — start with replication slots

PostgreSQL will not delete WAL while a replication slot still holds it. Two months of WAL accumulated after a standby died, and the secondary damage that only surfaced once the space came back.

  • #PostgreSQL
  • #Operations
  • #Monitoring
[proxmox] Fixing "CPU does not support x86-64-v2"
The default CPU model kvm64 failing to meet the x86-64-v2 level a container image requires

[proxmox] Fixing "CPU does not support x86-64-v2"

Modern container images require x86-64-v2. The default CPU model your hypervisor hands a VM does not have those instructions, so the process dies inside glibc before the application ever runs.

  • #Proxmox
  • #Virtualization
  • #Containers
Fixing ssh "Connection timed out during banner exchange"
A contrast showing the SYN packet of a connection passing while the ACK packet of the same connection is dropped

Fixing ssh "Connection timed out during banner exchange"

If your iptables whitelist only inspects ctstate NEW, the handshake never completes. A port scan still reports the port open, which is why the cause of a failing backup went unfound for a month.

  • #iptables
  • #Networking
  • #Troubleshooting
[haproxy] The reload succeeded and the config still did not change
A reload command returning rc=0 and OK while the configuration remains unchanged

[haproxy] The reload succeeded and the config still did not change

A reload can return rc=0 without applying anything. Two patterns: an old process surviving and serving the previous config, and changes landing in a staging file that never got promoted.

  • #HAProxy
  • #Troubleshooting
  • #Operations
Recovering a failed Ubuntu 22.04 → 26.04 upgrade without reinstalling
A decision cue showing that half-installed at zero means no reinstall is needed, and 750 unpacked packages dropping to zero

Recovering a failed Ubuntu 22.04 → 26.04 upgrade without reinstalling

Skipping an LTS step and stalling at the configure stage takes sudo and networking with it. Counting dpkg states tells you whether recovery without a reinstall is possible.

  • #Ubuntu
  • #Linux
  • #Recovery
[monitoring] When "last updated at" is the wrong thing to alert on
A side-by-side contrast between a metric that stays still when healthy and one that moves when work is happening

[monitoring] When "last updated at" is the wrong thing to alert on

If you alert on a value that does not change while things are working, it will fire eventually — guaranteed. The metric that produced false alarms, and the real alert we silenced while fixing it.

  • #Monitoring
  • #Operations
  • #Alerting
[linux] You changed the timezone, but the DB and containers did not
Three boxes contrasting an OS set to KST against a database and container still running UTC

[linux] You changed the timezone, but the DB and containers did not

timedatectl changes the OS timezone. Your database and your containers keep the old one. How NOW() came back nine hours apart inside one cluster, and why nothing short of a restart fixes it.

  • #Linux
  • #MariaDB
  • #Docker