git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness /tmp/0.2.2 && cp -r /tmp/0.2.2/changelog/0.2.2/2026-08-11-humanizer- ~/.claude/skills/0.2.22026-08-11-humanizer-skill.md
# New library skill: humanizer - **Date:** 2026-08-11 - **Type:** feature - **Scope:** `skills` - **PR:** [#256](https://github.com/Prism-Shadow/penguin-harness/pull/256) [中文版](2026-08-11-humanizer-skill.zh.md) `humanizer` joins the built-in library under the Office Productivity group as a manual-install skill — `preinstall: false`, so it stays out of default_agent's preinstalled set and installs from the Skill Library on demand ([#256](https://github.com/Prism-Shadow/penguin-harness/pull/256)). ## What it does Rewrites or edits prose in any language so it reads like edited human writing in the register of books, newspapers and encyclopedias rather than default AI output. The SKILL.md working surface is deliberately small: seven drafting principles (vary every pattern, density from anchored facts in whole grammar, a quota on quotable lines, a real writer with real material, structure serving content, native idiom and typography, verify-and-calibrate) under one meta-rule — aim for the natural distribution of edited prose, not a perfect scorecard. The detailed instrument ships as reference files read during the diagnostic census rather than during drafting: a 25-tell catalog in three layers (sentence patterns, discourse shapes, and the humanizing pass's own fingerprint), per-language surface forms for Chinese, English, Japanese, French, German and Spanish, and the measured case study behind every rule. ## How it was derived Five empirical rounds, documented with counts in the skill's `reference/case-study.md`: a controlled baseline-versus-rewrite exercise; reader field tests that yielded the discourse-layer tells; a ten-piece, five-language blind cross-review by independent editor agents (two pieces double-reviewed for agreement, all pieces revised against accepted flags, the four worst re-scored blind with the average verdict dropping from ~71 to 30) that yielded the scrub layer; and two further user passes that caught successive overcorrections (telegraph compression, then clipped conditionals and connective monotony). One recurring result shaped the design: drafting against a long checklist produces compliance-shaped text, so the rules live in the diagnosis step and the drafting core stays small.
Use when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release workflow; writing or auditing changelog entries; writing a blog post or capturing release screenshots; deciding what to do about data already on disk; or auditing prose that reads like a leaked authoring session. Covers the two-repo symlink layout, the CI-parity verification chain, the record-and-ship contract, where blog media is hosted, and the seams that are intentional.
Use when changing the PenguinHarness Web App (`packages/web`) — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a form field, writing user-facing copy, or building a popup. Covers the semantic tone tokens, the icon size/stroke/gap scale, the semantic-versus-formatting rule for explanatory text, the two-dictionary i18n contract, and the portal-panel pattern with its Esc and scroll caveats.
Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
Initialize an Agent's settings from a user requirement by writing AGENTS.md, setting identity metadata, and installing only needed Skills.
Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.