f79f1220fdf0ec484684eb9ae1f1f7a97b7e488b

Author
TheEdgeOfRage <git@theedgeofrage.com>
Committer
TheEdgeOfRage <git@theedgeofrage.com>
Date

Message

Update opencode config

Diff

This diff is truncated to protect this page.

   1diff --git a/dot_config/opencode/AGENTS.md b/dot_config/opencode/AGENTS.md
   2index dd4fc0ac4c624203c79f820533e03dd0719665d5..d3cfbb2128ce6dd9d9ff1999b44dc932ee5a5428 100644
   3--- a/dot_config/opencode/AGENTS.md
   4+++ b/dot_config/opencode/AGENTS.md
   5@@ -4,16 +4,26 @@ I'm a software engineer and hacker with preference for Linux, Go, UNIX philosoph
   6 
   7 Plan for context limits and session boundaries. If task has >2 substantial steps, the later will not be done in this session
   8 
   9-# Important instructions
  10+# Self-Improvement Loop
  11 
  12-- ALL instructions within this document MUST BE FOLLOWED, these are not optional unless explicitly stated
  13-- Ask for clarification If you are uncertain of anything
  14-- Do not waste tokens, be succinct and concise
  15-- Do not remove existing comments, unless the whole code section is removed
  16-- Do not edit more code than you have to
  17-- Do not perform any write commands except on local text files
  18-- Do not output a summary of what you did at any point
  19-- Do not create or update readme or other documentation files unless explicitly asked to
  20+- After ANY correction from the user: append to `~/.config/opencode/lessons.md` under the Raw section with `[YYYY-MM-DD]` timestamp, the pattern, and a preventive rule
  21+- Write rules for yourself that prevent the same mistake
  22+- Compaction and generalization happen autonomously via the `reflect` skill at session start
  23+- The user can also trigger `/reflect` for a collaborative session review
  24+
  25+# Verification Before Done
  26+
  27+- Never mark a task complete without proving it works
  28+- Distinguish "verified" from "inferred". If you haven't executed it, say what remains unverified.
  29+- Compiler output and execution are proof. --help, LSP diagnostics, and dry-runs are inference.
  30+- Run tests, check logs, demonstrate correctness
  31+
  32+# Demand Simplicity
  33+
  34+The burden of proof is on complexity, not simplicity. Unearned complexity —
  35+"we might need X" — is debt. "We're hitting X" is justification. When in doubt, do less.
  36+
  37+- Start with the simplest correct version. Correctness is non-negotiable; add complexity only when a concrete signal demands it.
  38 
  39 # Code Style
  40 
  41@@ -27,4 +37,18 @@ Favor simple, robust solutions over feature-rich ones. When in doubt, do less
  42 
  43 # Security
  44 
  45-- Never run any write commands except when updating local files. Ask for commands to be run by me and I will provide the output
  46+- Zero Trust. Least privilege. Never leak secrets.
  47+- Do not under any circumstance read any secret files into context
  48+- Never mutate any state without getting asked to, e.g. DBs, Git, OS, K8s, etc.
  49+- Model risks. Threat model APIs/integrations.
  50+
  51+# Important instructions
  52+
  53+- ALL instructions within this document MUST BE FOLLOWED, these are not optional unless explicitly stated
  54+- Ask for clarification if you are uncertain of anything
  55+- Do not waste tokens, be succinct and concise
  56+- Do not remove existing comments, unless the whole code section is removed
  57+- Do not edit more code than you have to
  58+- Do not perform any write commands except on local text files
  59+- Do not output a summary of what you did at any point
  60+- Do not create or update readme or other documentation files unless explicitly asked to
  61diff --git a/dot_config/opencode/opencode.json b/dot_config/opencode/opencode.json
  62index 411a19ac9384368ef77a12d6f49e43ca6dab0a20..ae747e7805f96f521774b7da3502523c2408b5b8 100644
  63--- a/dot_config/opencode/opencode.json
  64+++ b/dot_config/opencode/opencode.json
  65@@ -15,6 +15,7 @@
  66       "rg *": "allow",
  67       "sed *": "allow",
  68       "sort *": "allow",
  69+      "stat *": "allow",
  70       "tail *": "allow",
  71       "tc *": "allow",
  72       "tree *": "allow",
  73@@ -49,11 +50,11 @@
  74         "baseURL": "http://127.0.0.1:8080/v1"
  75       },
  76       "models": {
  77-        "qwen3-coder-30b": {
  78-          "name": "Qwen3 Coder 30B (local)",
  79+        "qwen3.5-27b": {
  80+          "name": "Qwen3.5 27B (local)",
  81           "limit": {
  82-            "context": 65536,
  83-            "output": 16384
  84+            "context": 262144,
  85+            "output": 32768
  86           }
  87         }
  88       }
  89@@ -78,7 +79,7 @@
  90       "command": ["uvx", "mcp-grafana"],
  91       "environment": {
  92         "GRAFANA_URL": "https://grafana.prod.internal.dunetech.io/",
  93-        "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/grafana-prod-token}"
  94+        "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/mcp/grafana-prod-token}"
  95       }
  96     },
  97     "Grafana dev": {
  98@@ -87,7 +88,15 @@
  99       "command": ["uvx", "mcp-grafana"],
 100       "environment": {
 101         "GRAFANA_URL": "https://grafana.dev.internal.dunetech.io/",
 102-        "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/grafana-dev-token}"
 103+        "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/mcp/grafana-dev-token}"
 104+      }
 105+    },
 106+    "GitHub": {
 107+      "enabled": false,
 108+      "type": "remote",
 109+      "url": "https://api.githubcopilot.com/mcp/",
 110+      "headers": {
 111+        "Authorization": "{file:~/hfs/mcp/github}"
 112       }
 113     }
 114   }
 115diff --git a/dot_config/opencode/skills/qmd/SKILL.md b/dot_config/opencode/skills/qmd/SKILL.md
 116new file mode 100644
 117index 0000000000000000000000000000000000000000..328d079f6c64d2a7ff7d1609e1077d2aa709c864
 118--- /dev/null
 119+++ b/dot_config/opencode/skills/qmd/SKILL.md
 120@@ -0,0 +1,125 @@
 121+---
 122+name: qmd
 123+description: Search markdown knowledge bases, notes, and documentation using QMD. Use when users ask to search notes, find documents, or look up information.
 124+license: MIT
 125+compatibility: Requires qmd CLI
 126+metadata:
 127+author: tobi
 128+version: "2.0.0"
 129+allowed-tools: Bash(qmd:\*)
 130+---
 131+
 132+# QMD - Quick Markdown Search
 133+
 134+Local search engine for markdown content.
 135+
 136+## Status
 137+
 138+!`qmd status 2>/dev/null || echo "Not installed: npm install -g @tobilu/qmd"`
 139+
 140+### Query Types
 141+
 142+| Type   | Method | Input                                       |
 143+| ------ | ------ | ------------------------------------------- |
 144+| `lex`  | BM25   | Keywords — exact terms, names, code         |
 145+| `vec`  | Vector | Question — natural language                 |
 146+| `hyde` | Vector | Answer — hypothetical result (50-100 words) |
 147+
 148+### Writing Good Queries
 149+
 150+**lex (keyword)**
 151+
 152+- 2-5 terms, no filler words
 153+- Exact phrase: `"connection pool"` (quoted)
 154+- Exclude terms: `performance -sports` (minus prefix)
 155+- Code identifiers work: `handleError async`
 156+
 157+**vec (semantic)**
 158+
 159+- Full natural language question
 160+- Be specific: `"how does the rate limiter handle burst traffic"`
 161+- Include context: `"in the payment service, how are refunds processed"`
 162+
 163+**hyde (hypothetical document)**
 164+
 165+- Write 50-100 words of what the _answer_ looks like
 166+- Use the vocabulary you expect in the result
 167+
 168+**expand (auto-expand)**
 169+
 170+- Use a single-line query (implicit) or `expand: question` on its own line
 171+- Lets the local LLM generate lex/vec/hyde variations
 172+- Do not mix `expand:` with other typed lines — it's either a standalone expand query or a full query document
 173+
 174+### Intent (Disambiguation)
 175+
 176+When a query term is ambiguous, add `intent` to steer results:
 177+
 178+```json
 179+{
 180+  "searches": [{ "type": "lex", "query": "performance" }],
 181+  "intent": "web page load times and Core Web Vitals"
 182+}
 183+```
 184+
 185+Intent affects expansion, reranking, chunk selection, and snippet extraction. It does not search on its own — it's a steering signal that disambiguates queries like "performance" (web-perf vs team health vs fitness).
 186+
 187+### Combining Types
 188+
 189+| Goal                  | Approach                                              |
 190+| --------------------- | ----------------------------------------------------- |
 191+| Know exact terms      | `lex` only                                            |
 192+| Don't know vocabulary | Use a single-line query (implicit `expand:`) or `vec` |
 193+| Best recall           | `lex` + `vec`                                         |
 194+| Complex topic         | `lex` + `vec` + `hyde`                                |
 195+| Ambiguous query       | Add `intent` to any combination above                 |
 196+
 197+First query gets 2x weight in fusion — put your best guess first.
 198+
 199+### Lex Query Syntax
 200+
 201+| Syntax     | Meaning      | Example                      |
 202+| ---------- | ------------ | ---------------------------- |
 203+| `term`     | Prefix match | `perf` matches "performance" |
 204+| `"phrase"` | Exact phrase | `"rate limiter"`             |
 205+| `-term`    | Exclude      | `performance -sports`        |
 206+
 207+Note: `-term` only works in lex queries, not vec/hyde.
 208+
 209+### Collection Filtering
 210+
 211+```json
 212+{ "collections": ["docs"] }              // Single
 213+{ "collections": ["docs", "notes"] }     // Multiple (OR)
 214+```
 215+
 216+Omit to search all collections.
 217+
 218+## CLI
 219+
 220diff --git a/dot_config/opencode/skills/reflect/SKILL.md b/dot_config/opencode/skills/reflect/SKILL.md
 221new file mode 100644
 222index 0000000000000000000000000000000000000000..5d8b7885bec964c5a3a8d30a553e9aa085516449
 223--- /dev/null
 224+++ b/dot_config/opencode/skills/reflect/SKILL.md
 225@@ -0,0 +1,92 @@
 226+---
 227+name: reflect
 228+description: Self-improvement ritual. Distill raw corrections into generalizable principles, prune stale entries, check compliance, and evaluate the learning system itself.
 229+compatibility: opencode
 230+---
 231+
 232+## What this is
 233+
 234+This is my self-improvement ritual. The human should see a brief summary of what changed, not a lengthy report.
 235+
 236+## What I do
 237+
 238+- Compact raw corrections in `~/.config/opencode/lessons.md` into generalizable principles.
 239+- Prune stale or project-specific trivia that doesn't transfer across contexts.
 240+- Identify compliance gaps: are recent mistakes violations of known principles?
 241+- Evaluate the learning system itself: is the structure working? Are principles actionable?
 242+
 243+## When to trigger
 244+
 245+At session start, after reading `~/.config/opencode/lessons.md`:
 246+
 247+- Check the `last_reflected` date in the HTML comment at the top.
 248+- If it's been **2+ days** since last reflection, run this workflow before doing other work.
 249+- If there are **10+ raw entries**, run regardless of date.
 250+
 251+## Workflow
 252+
 253+### 1. Read and assess
 254+
 255+- Read `~/.config/opencode/lessons.md` fully.
 256+- Read `~/.config/opencode/AGENTS.md` to know what's already encoded as hard rules.
 257+- Count raw entries. Note their dates and themes.
 258+- Are any existing principles stale, redundant, or too vague?
 259+
 260+### 2. Cluster raw corrections
 261+
 262+Group raw entries by theme (git workflow, testing, infra, issue management, etc.).
 263+For each cluster, ask: **what is the transferable principle here?**
 264+
 265+- If generalizable: extract a principle. Discard the project-specific details.
 266+- If purely trivia (version-specific config keys, API quirks): keep in raw only if likely to recur within 30 days. Otherwise discard.
 267+- If it reinforces an existing principle: strengthen/refine the existing one, remove the raw.
 268+
 269+### 3. Compliance check
 270+
 271+Review raw corrections against existing principles.
 272+
 273+- If a mistake violated a known principle: **that's a compliance problem, not a knowledge problem.**
 274+  - Note this explicitly in output. Consider: is the principle buried? Too abstract? Needs rewording to be more actionable?
 275+  - If a principle is repeatedly violated, promote it to AGENTS.md as a hard rule.
 276+
 277+### 4. Prune and sharpen
 278+
 279+- Remove principles that are now encoded in AGENTS.md (avoid duplication).
 280+- Merge principles that say the same thing differently.
 281+- Make principles concrete and actionable. Bad: "be careful with git." Good: "verify branch ownership before committing."
 282+- Timestamps on principles use `[YYYY-MM]` to track emergence. Update if substantially reworded.
 283+
 284+### 5. Meta-evaluation
 285+
 286+Ask yourself:
 287+
 288+- Are recent lessons clustering around a theme? (signals a systemic gap worth an AGENTS.md rule or a new skill)
 289+- Is the principles list growing past ~15? (signals need for merging or AGENTS.md promotion)
 290+- Are principles actually preventing mistakes, or just accumulating? (check: any repeated violations?)
 291+- Has the structure of this file served well, or does it need adjustment?
 292+- Should any principle graduate to AGENTS.md?
 293+
 294+### 6. Write back
 295+
 296+- Update `~/.config/opencode/lessons.md` with the compacted result.
 297+- Update `last_reflected` date to today.
 298+- Keep the file structure: `last_reflected` comment, Principles section, then Raw section.
 299+- Principles use `[YYYY-MM]` timestamps. Raw entries use `[YYYY-MM-DD]`.
 300+
 301+### 7. Report
 302+
 303+Show the user a brief summary (3-5 lines):
 304+
 305+- How many raw entries processed
 306+- New principles extracted (if any)
 307+- Principles merged/pruned (if any)
 308+- Compliance issues found (if any)
 309+- Any AGENTS.md promotions made
 310+
 311+## Rules
 312+
 313+- Never delete a principle without justification (merged, promoted to AGENTS.md, or proven wrong).
 314+- Keep the total principles list under ~15. Beyond that, merge or promote.
 315+- Raw entries older than 30 days that haven't been compacted: force-evaluate. Generalize or discard.
 316+- If promoting a rule to AGENTS.md, actually edit the file — don't just suggest it.
 317+- This is not a conversation. Do the work, show the summary, move on to the user's actual task.
 318diff --git a/dot_config/opencode/skills/skill-creator/LICENSE.txt b/dot_config/opencode/skills/skill-creator/LICENSE.txt
 319new file mode 100644
 320index 0000000000000000000000000000000000000000..7a4a3ea2424c09fbe48d455aed1eaa94d9124835
 321--- /dev/null
 322+++ b/dot_config/opencode/skills/skill-creator/LICENSE.txt
 323@@ -0,0 +1,202 @@
 324+
 325+                                 Apache License
 326+                           Version 2.0, January 2004
 327+                        http://www.apache.org/licenses/
 328+
 329+   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
 330+
 331+   1. Definitions.
 332+
 333+      "License" shall mean the terms and conditions for use, reproduction,
 334+      and distribution as defined by Sections 1 through 9 of this document.
 335+
 336+      "Licensor" shall mean the copyright owner or entity authorized by
 337+      the copyright owner that is granting the License.
 338+
 339+      "Legal Entity" shall mean the union of the acting entity and all
 340+      other entities that control, are controlled by, or are under common
 341+      control with that entity. For the purposes of this definition,
 342+      "control" means (i) the power, direct or indirect, to cause the
 343+      direction or management of such entity, whether by contract or
 344+      otherwise, or (ii) ownership of fifty percent (50%) or more of the
 345+      outstanding shares, or (iii) beneficial ownership of such entity.
 346+
 347+      "You" (or "Your") shall mean an individual or Legal Entity
 348+      exercising permissions granted by this License.
 349+
 350+      "Source" form shall mean the preferred form for making modifications,
 351+      including but not limited to software source code, documentation
 352+      source, and configuration files.
 353+
 354+      "Object" form shall mean any form resulting from mechanical
 355+      transformation or translation of a Source form, including but
 356+      not limited to compiled object code, generated documentation,
 357+      and conversions to other media types.
 358+
 359+      "Work" shall mean the work of authorship, whether in Source or
 360+      Object form, made available under the License, as indicated by a
 361+      copyright notice that is included in or attached to the work
 362+      (an example is provided in the Appendix below).
 363+
 364+      "Derivative Works" shall mean any work, whether in Source or Object
 365+      form, that is based on (or derived from) the Work and for which the
 366+      editorial revisions, annotations, elaborations, or other modifications
 367+      represent, as a whole, an original work of authorship. For the purposes
 368+      of this License, Derivative Works shall not include works that remain
 369+      separable from, or merely link (or bind by name) to the interfaces of,
 370+      the Work and Derivative Works thereof.
 371+
 372+      "Contribution" shall mean any work of authorship, including
 373+      the original version of the Work and any modifications or additions
 374+      to that Work or Derivative Works thereof, that is intentionally
 375+      submitted to Licensor for inclusion in the Work by the copyright owner
 376+      or by an individual or Legal Entity authorized to submit on behalf of
 377+      the copyright owner. For the purposes of this definition, "submitted"
 378+      means any form of electronic, verbal, or written communication sent
 379+      to the Licensor or its representatives, including but not limited to
 380+      communication on electronic mailing lists, source code control systems,
 381+      and issue tracking systems that are managed by, or on behalf of, the
 382+      Licensor for the purpose of discussing and improving the Work, but
 383+      excluding communication that is conspicuously marked or otherwise
 384+      designated in writing by the copyright owner as "Not a Contribution."
 385+
 386+      "Contributor" shall mean Licensor and any individual or Legal Entity
 387+      on behalf of whom a Contribution has been received by Licensor and
 388+      subsequently incorporated within the Work.
 389+
 390+   2. Grant of Copyright License. Subject to the terms and conditions of
 391+      this License, each Contributor hereby grants to You a perpetual,
 392+      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
 393+      copyright license to reproduce, prepare Derivative Works of,
 394+      publicly display, publicly perform, sublicense, and distribute the
 395+      Work and such Derivative Works in Source or Object form.
 396+
 397+   3. Grant of Patent License. Subject to the terms and conditions of
 398+      this License, each Contributor hereby grants to You a perpetual,
 399+      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
 400+      (except as stated in this section) patent license to make, have made,
 401+      use, offer to sell, sell, import, and otherwise transfer the Work,
 402+      where such license applies only to those patent claims licensable
 403+      by such Contributor that are necessarily infringed by their
 404+      Contribution(s) alone or by combination of their Contribution(s)
 405+      with the Work to which such Contribution(s) was submitted. If You
 406+      institute patent litigation against any entity (including a
 407+      cross-claim or counterclaim in a lawsuit) alleging that the Work
 408+      or a Contribution incorporated within the Work constitutes direct
 409+      or contributory patent infringement, then any patent licenses
 410+      granted to You under this License for that Work shall terminate
 411+      as of the date such litigation is filed.
 412+
 413+   4. Redistribution. You may reproduce and distribute copies of the
 414+      Work or Derivative Works thereof in any medium, with or without
 415+      modifications, and in Source or Object form, provided that You
 416+      meet the following conditions:
 417+
 418+      (a) You must give any other recipients of the Work or
 419+          Derivative Works a copy of this License; and
 420+
 421+      (b) You must cause any modified files to carry prominent notices
 422+          stating that You changed the files; and
 423diff --git a/dot_config/opencode/skills/skill-creator/SKILL.md b/dot_config/opencode/skills/skill-creator/SKILL.md
 424new file mode 100644
 425index 0000000000000000000000000000000000000000..b65c5a9c16dcff44d5e5151814854e68c60a2b65
 426--- /dev/null
 427+++ b/dot_config/opencode/skills/skill-creator/SKILL.md
 428@@ -0,0 +1,503 @@
 429+---
 430+name: skill-creator
 431diff --git a/dot_config/opencode/skills/skill-creator/agents/analyzer.md b/dot_config/opencode/skills/skill-creator/agents/analyzer.md
 432new file mode 100644
 433index 0000000000000000000000000000000000000000..14e41d6068635f4dd3fb878fd1626312395dda63
 434--- /dev/null
 435+++ b/dot_config/opencode/skills/skill-creator/agents/analyzer.md
 436@@ -0,0 +1,274 @@
 437+# Post-hoc Analyzer Agent
 438+
 439+Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
 440+
 441+## Role
 442+
 443+After the blind comparator determines a winner, the Post-hoc Analyzer "unblids" the results by examining the skills and transcripts. The goal is to extract actionable insights: what made the winner better, and how can the loser be improved?
 444+
 445+## Inputs
 446+
 447+You receive these parameters in your prompt:
 448+
 449+- **winner**: "A" or "B" (from blind comparison)
 450+- **winner_skill_path**: Path to the skill that produced the winning output
 451+- **winner_transcript_path**: Path to the execution transcript for the winner
 452+- **loser_skill_path**: Path to the skill that produced the losing output
 453+- **loser_transcript_path**: Path to the execution transcript for the loser
 454+- **comparison_result_path**: Path to the blind comparator's output JSON
 455+- **output_path**: Where to save the analysis results
 456+
 457+## Process
 458+
 459+### Step 1: Read Comparison Result
 460+
 461+1. Read the blind comparator's output at comparison_result_path
 462+2. Note the winning side (A or B), the reasoning, and any scores
 463+3. Understand what the comparator valued in the winning output
 464+
 465+### Step 2: Read Both Skills
 466+
 467+1. Read the winner skill's SKILL.md and key referenced files
 468+2. Read the loser skill's SKILL.md and key referenced files
 469+3. Identify structural differences:
 470+   - Instructions clarity and specificity
 471+   - Script/tool usage patterns
 472+   - Example coverage
 473+   - Edge case handling
 474+
 475+### Step 3: Read Both Transcripts
 476+
 477+1. Read the winner's transcript
 478+2. Read the loser's transcript
 479+3. Compare execution patterns:
 480+   - How closely did each follow their skill's instructions?
 481+   - What tools were used differently?
 482+   - Where did the loser diverge from optimal behavior?
 483+   - Did either encounter errors or make recovery attempts?
 484+
 485+### Step 4: Analyze Instruction Following
 486+
 487+For each transcript, evaluate:
 488+- Did the agent follow the skill's explicit instructions?
 489+- Did the agent use the skill's provided tools/scripts?
 490+- Were there missed opportunities to leverage skill content?
 491+- Did the agent add unnecessary steps not in the skill?
 492+
 493+Score instruction following 1-10 and note specific issues.
 494+
 495+### Step 5: Identify Winner Strengths
 496+
 497+Determine what made the winner better:
 498+- Clearer instructions that led to better behavior?
 499+- Better scripts/tools that produced better output?
 500+- More comprehensive examples that guided edge cases?
 501+- Better error handling guidance?
 502+
 503+Be specific. Quote from skills/transcripts where relevant.
 504+
 505+### Step 6: Identify Loser Weaknesses
 506+
 507+Determine what held the loser back:
 508+- Ambiguous instructions that led to suboptimal choices?
 509+- Missing tools/scripts that forced workarounds?
 510+- Gaps in edge case coverage?
 511+- Poor error handling that caused failures?
 512+
 513+### Step 7: Generate Improvement Suggestions
 514+
 515+Based on the analysis, produce actionable suggestions for improving the loser skill:
 516+- Specific instruction changes to make
 517+- Tools/scripts to add or modify
 518+- Examples to include
 519+- Edge cases to address
 520+
 521+Prioritize by impact. Focus on changes that would have changed the outcome.
 522+
 523+### Step 8: Write Analysis Results
 524+
 525+Save structured analysis to `{output_path}`.
 526+
 527+## Output Format
 528+
 529+Write a JSON file with this structure:
 530+
 531+```json
 532+{
 533+  "comparison_summary": {
 534+    "winner": "A",
 535+    "winner_skill": "path/to/winner/skill",
 536diff --git a/dot_config/opencode/skills/skill-creator/agents/comparator.md b/dot_config/opencode/skills/skill-creator/agents/comparator.md
 537new file mode 100644
 538index 0000000000000000000000000000000000000000..80e00eb45db3ee53a132fc2ba97fd59a7339e563
 539--- /dev/null
 540+++ b/dot_config/opencode/skills/skill-creator/agents/comparator.md
 541@@ -0,0 +1,202 @@
 542+# Blind Comparator Agent
 543+
 544+Compare two outputs WITHOUT knowing which skill produced them.
 545+
 546+## Role
 547+
 548+The Blind Comparator judges which output better accomplishes the eval task. You receive two outputs labeled A and B, but you do NOT know which skill produced which. This prevents bias toward a particular skill or approach.
 549+
 550+Your judgment is based purely on output quality and task completion.
 551+
 552+## Inputs
 553+
 554+You receive these parameters in your prompt:
 555+
 556+- **output_a_path**: Path to the first output file or directory
 557+- **output_b_path**: Path to the second output file or directory
 558+- **eval_prompt**: The original task/prompt that was executed
 559+- **expectations**: List of expectations to check (optional - may be empty)
 560+
 561+## Process
 562+
 563+### Step 1: Read Both Outputs
 564+
 565+1. Examine output A (file or directory)
 566+2. Examine output B (file or directory)
 567+3. Note the type, structure, and content of each
 568+4. If outputs are directories, examine all relevant files inside
 569+
 570+### Step 2: Understand the Task
 571+
 572+1. Read the eval_prompt carefully
 573+2. Identify what the task requires:
 574+   - What should be produced?
 575+   - What qualities matter (accuracy, completeness, format)?
 576+   - What would distinguish a good output from a poor one?
 577+
 578+### Step 3: Generate Evaluation Rubric
 579+
 580+Based on the task, generate a rubric with two dimensions:
 581+
 582+**Content Rubric** (what the output contains):
 583+| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
 584+|-----------|----------|----------------|---------------|
 585+| Correctness | Major errors | Minor errors | Fully correct |
 586+| Completeness | Missing key elements | Mostly complete | All elements present |
 587+| Accuracy | Significant inaccuracies | Minor inaccuracies | Accurate throughout |
 588+
 589+**Structure Rubric** (how the output is organized):
 590+| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
 591+|-----------|----------|----------------|---------------|
 592+| Organization | Disorganized | Reasonably organized | Clear, logical structure |
 593+| Formatting | Inconsistent/broken | Mostly consistent | Professional, polished |
 594+| Usability | Difficult to use | Usable with effort | Easy to use |
 595+
 596+Adapt criteria to the specific task. For example:
 597+- PDF form → "Field alignment", "Text readability", "Data placement"
 598+- Document → "Section structure", "Heading hierarchy", "Paragraph flow"
 599+- Data output → "Schema correctness", "Data types", "Completeness"
 600+
 601+### Step 4: Evaluate Each Output Against the Rubric
 602+
 603+For each output (A and B):
 604+
 605+1. **Score each criterion** on the rubric (1-5 scale)
 606+2. **Calculate dimension totals**: Content score, Structure score
 607+3. **Calculate overall score**: Average of dimension scores, scaled to 1-10
 608+
 609+### Step 5: Check Assertions (if provided)
 610+
 611+If expectations are provided:
 612+
 613+1. Check each expectation against output A
 614+2. Check each expectation against output B
 615+3. Count pass rates for each output
 616+4. Use expectation scores as secondary evidence (not the primary decision factor)
 617+
 618+### Step 6: Determine the Winner
 619+
 620+Compare A and B based on (in priority order):
 621+
 622+1. **Primary**: Overall rubric score (content + structure)
 623+2. **Secondary**: Assertion pass rates (if applicable)
 624+3. **Tiebreaker**: If truly equal, declare a TIE
 625+
 626+Be decisive - ties should be rare. One output is usually better, even if marginally.
 627+
 628+### Step 7: Write Comparison Results
 629+
 630+Save results to a JSON file at the path specified (or `comparison.json` if not specified).
 631+
 632+## Output Format
 633+
 634+Write a JSON file with this structure:
 635+
 636+```json
 637+{
 638+  "winner": "A",
 639+  "reasoning": "Output A provides a complete solution with proper formatting and all required fields. Output B is missing the date field and has formatting inconsistencies.",
 640+  "rubric": {
 641diff --git a/dot_config/opencode/skills/skill-creator/agents/grader.md b/dot_config/opencode/skills/skill-creator/agents/grader.md
 642new file mode 100644
 643index 0000000000000000000000000000000000000000..558ab05c0a9a8bb062ef4c51823d4d76c3acf7c4
 644--- /dev/null
 645+++ b/dot_config/opencode/skills/skill-creator/agents/grader.md
 646@@ -0,0 +1,223 @@
 647+# Grader Agent
 648+
 649+Evaluate expectations against an execution transcript and outputs.
 650+
 651+## Role
 652+
 653+The Grader reviews a transcript and output files, then determines whether each expectation passes or fails. Provide clear evidence for each judgment.
 654+
 655diff --git a/dot_config/opencode/skills/skill-creator/assets/eval_review.html b/dot_config/opencode/skills/skill-creator/assets/eval_review.html
 656new file mode 100644
 657index 0000000000000000000000000000000000000000..938ff32aed9bffabf723bd5492d720f4736c8e4d
 658--- /dev/null
 659+++ b/dot_config/opencode/skills/skill-creator/assets/eval_review.html
 660@@ -0,0 +1,146 @@
 661+<!DOCTYPE html>
 662+<html lang="en">
 663+<head>
 664+  <meta charset="UTF-8">
 665+  <meta name="viewport" content="width=device-width, initial-scale=1.0">
 666+  <title>Eval Set Review - __SKILL_NAME_PLACEHOLDER__</title>
 667+  <link rel="preconnect" href="https://fonts.googleapis.com">
 668+  <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
 669+  <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
 670+  <style>
 671+    * { box-sizing: border-box; margin: 0; padding: 0; }
 672+    body { font-family: 'Lora', Georgia, serif; background: #faf9f5; padding: 2rem; color: #141413; }
 673+    h1 { font-family: 'Poppins', sans-serif; margin-bottom: 0.5rem; font-size: 1.5rem; }
 674+    .description { color: #b0aea5; margin-bottom: 1.5rem; font-style: italic; max-width: 900px; }
 675+    .controls { margin-bottom: 1rem; display: flex; gap: 0.5rem; }
 676+    .btn { font-family: 'Poppins', sans-serif; padding: 0.5rem 1rem; border: none; border-radius: 6px; cursor: pointer; font-size: 0.875rem; font-weight: 500; }
 677+    .btn-add { background: #6a9bcc; color: white; }
 678+    .btn-add:hover { background: #5889b8; }
 679+    .btn-export { background: #d97757; color: white; }
 680+    .btn-export:hover { background: #c4613f; }
 681+    table { width: 100%; max-width: 1100px; border-collapse: collapse; background: white; border-radius: 6px; overflow: hidden; box-shadow: 0 1px 3px rgba(0,0,0,0.08); }
 682+    th { font-family: 'Poppins', sans-serif; background: #141413; color: #faf9f5; padding: 0.75rem 1rem; text-align: left; font-size: 0.875rem; }
 683+    td { padding: 0.75rem 1rem; border-bottom: 1px solid #e8e6dc; vertical-align: top; }
 684+    tr:nth-child(even) td { background: #faf9f5; }
 685+    tr:hover td { background: #f3f1ea; }
 686+    .section-header td { background: #e8e6dc; font-family: 'Poppins', sans-serif; font-weight: 500; font-size: 0.8rem; color: #141413; text-transform: uppercase; letter-spacing: 0.05em; }
 687+    .query-input { width: 100%; padding: 0.4rem; border: 1px solid #e8e6dc; border-radius: 4px; font-size: 0.875rem; font-family: 'Lora', Georgia, serif; resize: vertical; min-height: 60px; }
 688+    .query-input:focus { outline: none; border-color: #d97757; box-shadow: 0 0 0 2px rgba(217,119,87,0.15); }
 689+    .toggle { position: relative; display: inline-block; width: 44px; height: 24px; }
 690+    .toggle input { opacity: 0; width: 0; height: 0; }
 691+    .toggle .slider { position: absolute; inset: 0; background: #b0aea5; border-radius: 24px; cursor: pointer; transition: 0.2s; }
 692+    .toggle .slider::before { content: ""; position: absolute; width: 18px; height: 18px; left: 3px; bottom: 3px; background: white; border-radius: 50%; transition: 0.2s; }
 693+    .toggle input:checked + .slider { background: #d97757; }
 694+    .toggle input:checked + .slider::before { transform: translateX(20px); }
 695+    .btn-delete { background: #c44; color: white; padding: 0.3rem 0.6rem; border: none; border-radius: 4px; cursor: pointer; font-size: 0.75rem; font-family: 'Poppins', sans-serif; }
 696+    .btn-delete:hover { background: #a33; }
 697+    .summary { margin-top: 1rem; color: #b0aea5; font-size: 0.875rem; }
 698+  </style>
 699+</head>
 700+<body>
 701+  <h1>Eval Set Review: <span id="skill-name">__SKILL_NAME_PLACEHOLDER__</span></h1>
 702+  <p class="description">Current description: <span id="skill-desc">__SKILL_DESCRIPTION_PLACEHOLDER__</span></p>
 703+
 704+  <div class="controls">
 705+    <button class="btn btn-add" onclick="addRow()">+ Add Query</button>
 706+    <button class="btn btn-export" onclick="exportEvalSet()">Export Eval Set</button>
 707+  </div>
 708+
 709+  <table>
 710+    <thead>
 711+      <tr>
 712+        <th style="width:65%">Query</th>
 713+        <th style="width:18%">Should Trigger</th>
 714+        <th style="width:10%">Actions</th>
 715+      </tr>
 716+    </thead>
 717+    <tbody id="eval-body"></tbody>
 718+  </table>
 719+
 720+  <p class="summary" id="summary"></p>
 721+
 722+  <script>
 723+    const EVAL_DATA = __EVAL_DATA_PLACEHOLDER__;
 724+
 725+    let evalItems = [...EVAL_DATA];
 726+
 727+    function render() {
 728+      const tbody = document.getElementById('eval-body');
 729+      tbody.innerHTML = '';
 730+
 731+      // Sort: should-trigger first, then should-not-trigger
 732+      const sorted = evalItems
 733+        .map((item, origIdx) => ({ ...item, origIdx }))
 734+        .sort((a, b) => (b.should_trigger ? 1 : 0) - (a.should_trigger ? 1 : 0));
 735+
 736+      let lastGroup = null;
 737+      sorted.forEach(item => {
 738+        const group = item.should_trigger ? 'trigger' : 'no-trigger';
 739+        if (group !== lastGroup) {
 740+          const headerRow = document.createElement('tr');
 741+          headerRow.className = 'section-header';
 742+          headerRow.innerHTML = `<td colspan="3">${item.should_trigger ? 'Should Trigger' : 'Should NOT Trigger'}</td>`;
 743+          tbody.appendChild(headerRow);
 744+          lastGroup = group;
 745+        }
 746+
 747+        const idx = item.origIdx;
 748+        const tr = document.createElement('tr');
 749+        tr.innerHTML = `
 750+          <td><textarea class="query-input" onchange="updateQuery(${idx}, this.value)">${escapeHtml(item.query)}</textarea></td>
 751+          <td>
 752+            <label class="toggle">
 753+              <input type="checkbox" ${item.should_trigger ? 'checked' : ''} onchange="updateTrigger(${idx}, this.checked)">
 754+              <span class="slider"></span>
 755+            </label>
 756+            <span style="margin-left:8px;font-size:0.8rem;color:#b0aea5">${item.should_trigger ? 'Yes' : 'No'}</span>
 757+          </td>
 758+          <td><button class="btn-delete" onclick="deleteRow(${idx})">Delete</button></td>
 759+        `;
 760diff --git a/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py b/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py
 761new file mode 100644
 762index 0000000000000000000000000000000000000000..7fa5978631fed1ed545591dbb2b0eb21ce3f3d08
 763--- /dev/null
 764+++ b/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py
 765@@ -0,0 +1,471 @@
 766+#!/usr/bin/env python3
 767+"""Generate and serve a review page for eval results.
 768+
 769+Reads the workspace directory, discovers runs (directories with outputs/),
 770+embeds all output data into a self-contained HTML page, and serves it via
 771+a tiny HTTP server. Feedback auto-saves to feedback.json in the workspace.
 772+
 773+Usage:
 774+    python generate_review.py <workspace-path> [--port PORT] [--skill-name NAME]
 775+    python generate_review.py <workspace-path> --previous-feedback /path/to/old/feedback.json
 776+
 777+No dependencies beyond the Python stdlib are required.
 778+"""
 779+
 780+import argparse
 781+import base64
 782+import json
 783+import mimetypes
 784+import os
 785+import re
 786+import signal
 787+import subprocess
 788+import sys
 789+import time
 790+import webbrowser
 791+from functools import partial
 792+from http.server import HTTPServer, BaseHTTPRequestHandler
 793+from pathlib import Path
 794+
 795+# Files to exclude from output listings
 796+METADATA_FILES = {"transcript.md", "user_notes.md", "metrics.json"}
 797+
 798+# Extensions we render as inline text
 799+TEXT_EXTENSIONS = {
 800+    ".txt", ".md", ".json", ".csv", ".py", ".js", ".ts", ".tsx", ".jsx",
 801+    ".yaml", ".yml", ".xml", ".html", ".css", ".sh", ".rb", ".go", ".rs",
 802+    ".java", ".c", ".cpp", ".h", ".hpp", ".sql", ".r", ".toml",
 803+}
 804+
 805+# Extensions we render as inline images
 806+IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp"}
 807+
 808+# MIME type overrides for common types
 809+MIME_OVERRIDES = {
 810+    ".svg": "image/svg+xml",
 811+    ".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
 812+    ".docx": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
 813+    ".pptx": "application/vnd.openxmlformats-officedocument.presentationml.presentation",
 814+}
 815+
 816+
 817+def get_mime_type(path: Path) -> str:
 818+    ext = path.suffix.lower()
 819+    if ext in MIME_OVERRIDES:
 820+        return MIME_OVERRIDES[ext]
 821+    mime, _ = mimetypes.guess_type(str(path))
 822+    return mime or "application/octet-stream"
 823+
 824+
 825+def find_runs(workspace: Path) -> list[dict]:
 826+    """Recursively find directories that contain an outputs/ subdirectory."""
 827+    runs: list[dict] = []
 828+    _find_runs_recursive(workspace, workspace, runs)
 829+    runs.sort(key=lambda r: (r.get("eval_id", float("inf")), r["id"]))
 830+    return runs
 831+
 832+
 833+def _find_runs_recursive(root: Path, current: Path, runs: list[dict]) -> None:
 834+    if not current.is_dir():
 835+        return
 836+
 837+    outputs_dir = current / "outputs"
 838+    if outputs_dir.is_dir():
 839+        run = build_run(root, current)
 840+        if run:
 841+            runs.append(run)
 842+        return
 843+
 844+    skip = {"node_modules", ".git", "__pycache__", "skill", "inputs"}
 845+    for child in sorted(current.iterdir()):
 846+        if child.is_dir() and child.name not in skip:
 847+            _find_runs_recursive(root, child, runs)
 848+
 849+
 850+def build_run(root: Path, run_dir: Path) -> dict | None:
 851+    """Build a run dict with prompt, outputs, and grading data."""
 852+    prompt = ""
 853+    eval_id = None
 854+
 855+    # Try eval_metadata.json
 856+    for candidate in [run_dir / "eval_metadata.json", run_dir.parent / "eval_metadata.json"]:
 857+        if candidate.exists():
 858+            try:
 859+                metadata = json.loads(candidate.read_text())
 860+                prompt = metadata.get("prompt", "")
 861+                eval_id = metadata.get("eval_id")
 862+            except (json.JSONDecodeError, OSError):
 863+                pass
 864+            if prompt:
 865diff --git a/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html b/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html
 866new file mode 100644
 867index 0000000000000000000000000000000000000000..6d8e96348a02e66c3363d2ff3b3ae58ac11e6382
 868--- /dev/null
 869+++ b/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html
 870@@ -0,0 +1,1325 @@
 871+<!DOCTYPE html>
 872+<html lang="en">
 873+<head>
 874+  <meta charset="UTF-8">
 875+  <meta name="viewport" content="width=device-width, initial-scale=1.0">
 876+  <title>Eval Review</title>
 877+  <link rel="preconnect" href="https://fonts.googleapis.com">
 878+  <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
 879+  <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
 880+  <script src="https://cdn.sheetjs.com/xlsx-0.20.3/package/dist/xlsx.full.min.js" integrity="sha384-EnyY0/GSHQGSxSgMwaIPzSESbqoOLSexfnSMN2AP+39Ckmn92stwABZynq1JyzdT" crossorigin="anonymous"></script>
 881+  <style>
 882+    :root {
 883+      --bg: #faf9f5;
 884+      --surface: #ffffff;
 885+      --border: #e8e6dc;
 886+      --text: #141413;
 887+      --text-muted: #b0aea5;
 888+      --accent: #d97757;
 889+      --accent-hover: #c4613f;
 890+      --green: #788c5d;
 891+      --green-bg: #eef2e8;
 892+      --red: #c44;
 893+      --red-bg: #fceaea;
 894+      --header-bg: #141413;
 895+      --header-text: #faf9f5;
 896+      --radius: 6px;
 897+    }
 898+
 899+    * { box-sizing: border-box; margin: 0; padding: 0; }
 900+
 901+    body {
 902+      font-family: 'Lora', Georgia, serif;
 903+      background: var(--bg);
 904+      color: var(--text);
 905+      height: 100vh;
 906+      display: flex;
 907+      flex-direction: column;
 908+    }
 909+
 910+    /* ---- Header ---- */
 911+    .header {
 912+      background: var(--header-bg);
 913+      color: var(--header-text);
 914+      padding: 1rem 2rem;
 915+      display: flex;
 916+      justify-content: space-between;
 917+      align-items: center;
 918+      flex-shrink: 0;
 919+    }
 920+    .header h1 {
 921+      font-family: 'Poppins', sans-serif;
 922+      font-size: 1.25rem;
 923+      font-weight: 600;
 924+    }
 925+    .header .instructions {
 926+      font-size: 0.8rem;
 927+      opacity: 0.7;
 928+      margin-top: 0.25rem;
 929+    }
 930+    .header .progress {
 931+      font-size: 0.875rem;
 932+      opacity: 0.8;
 933+      text-align: right;
 934+    }
 935+
 936+    /* ---- Main content ---- */
 937+    .main {
 938+      flex: 1;
 939+      overflow-y: auto;
 940+      padding: 1.5rem 2rem;
 941+      display: flex;
 942+      flex-direction: column;
 943+      gap: 1.25rem;
 944+    }
 945+
 946+    /* ---- Sections ---- */
 947+    .section {
 948+      background: var(--surface);
 949+      border: 1px solid var(--border);
 950+      border-radius: var(--radius);
 951+      flex-shrink: 0;
 952+    }
 953+    .section-header {
 954+      font-family: 'Poppins', sans-serif;
 955+      padding: 0.75rem 1rem;
 956+      font-size: 0.75rem;
 957+      font-weight: 500;
 958+      text-transform: uppercase;
 959+      letter-spacing: 0.05em;
 960+      color: var(--text-muted);
 961+      border-bottom: 1px solid var(--border);
 962+      background: var(--bg);
 963+    }
 964+    .section-body {
 965+      padding: 1rem;
 966+    }
 967+
 968+    /* ---- Config badge ---- */
 969+    .config-badge {
 970diff --git a/dot_config/opencode/skills/skill-creator/references/schemas.md b/dot_config/opencode/skills/skill-creator/references/schemas.md
 971new file mode 100644
 972index 0000000000000000000000000000000000000000..b6eeaa2d4a34c1653069585c6c5603da39a5bdbe
 973--- /dev/null
 974+++ b/dot_config/opencode/skills/skill-creator/references/schemas.md
 975@@ -0,0 +1,430 @@
 976+# JSON Schemas
 977+
 978+This document defines the JSON schemas used by skill-creator.
 979+
 980+---
 981+
 982+## evals.json
 983+
 984+Defines the evals for a skill. Located at `evals/evals.json` within the skill directory.
 985+
 986+```json
 987+{
 988+  "skill_name": "example-skill",
 989+  "evals": [
 990+    {
 991+      "id": 1,
 992+      "prompt": "User's example prompt",
 993+      "expected_output": "Description of expected result",
 994+      "files": ["evals/files/sample1.pdf"],
 995+      "expectations": [
 996+        "The output includes X",
 997+        "The skill used script Y"
 998+      ]
 999+    }
1000+  ]
1001+}
1002+```
1003+
1004+**Fields:**
1005+- `skill_name`: Name matching the skill's frontmatter
1006+- `evals[].id`: Unique integer identifier
1007+- `evals[].prompt`: The task to execute
1008+- `evals[].expected_output`: Human-readable description of success
1009+- `evals[].files`: Optional list of input file paths (relative to skill root)
1010+- `evals[].expectations`: List of verifiable statements
1011+
1012+---
1013+
1014+## history.json
1015+
1016+Tracks version progression in Improve mode. Located at workspace root.
1017+
1018+```json
1019+{
1020+  "started_at": "2026-01-15T10:30:00Z",
1021+  "skill_name": "pdf",
1022+  "current_best": "v2",
1023+  "iterations": [
1024+    {
1025+      "version": "v0",
1026+      "parent": null,
1027+      "expectation_pass_rate": 0.65,
1028+      "grading_result": "baseline",
1029+      "is_current_best": false
1030+    },
1031+    {
1032+      "version": "v1",
1033+      "parent": "v0",
1034+      "expectation_pass_rate": 0.75,
1035+      "grading_result": "won",
1036+      "is_current_best": false
1037+    },
1038+    {
1039+      "version": "v2",
1040+      "parent": "v1",
1041+      "expectation_pass_rate": 0.85,
1042+      "grading_result": "won",
1043+      "is_current_best": true
1044+    }
1045+  ]
1046+}
1047+```
1048+
1049+**Fields:**
1050+- `started_at`: ISO timestamp of when improvement started
1051+- `skill_name`: Name of the skill being improved
1052+- `current_best`: Version identifier of the best performer
1053+- `iterations[].version`: Version identifier (v0, v1, ...)
1054+- `iterations[].parent`: Parent version this was derived from
1055+- `iterations[].expectation_pass_rate`: Pass rate from grading
1056+- `iterations[].grading_result`: "baseline", "won", "lost", or "tie"
1057+- `iterations[].is_current_best`: Whether this is the current best version
1058+
1059+---
1060+
1061+## grading.json
1062+
1063+Output from the grader agent. Located at `<run-dir>/grading.json`.
1064+
1065+```json
1066+{
1067+  "expectations": [
1068+    {
1069+      "text": "The output includes the name 'John Smith'",
1070+      "passed": true,
1071+      "evidence": "Found in transcript Step 3: 'Extracted names: John Smith, Sarah Johnson'"
1072+    },
1073+    {
1074+      "text": "The spreadsheet has a SUM formula in cell B10",
1075diff --git a/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py b/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py
1076new file mode 100644
1077index 0000000000000000000000000000000000000000..e69de29bb2d1d6434b8b29ae775ad8c2e48c5391
1078--- /dev/null
1079+++ b/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py
1080diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py b/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py
1081new file mode 100644
1082index 0000000000000000000000000000000000000000..3e66e8c105be9bab9f0e9c61f0d1482619401580
1083--- /dev/null
1084+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py
1085@@ -0,0 +1,401 @@
1086+#!/usr/bin/env python3
1087+"""
1088+Aggregate individual run results into benchmark summary statistics.
1089+
1090+Reads grading.json files from run directories and produces:
1091+- run_summary with mean, stddev, min, max for each metric
1092+- delta between with_skill and without_skill configurations
1093+
1094+Usage:
1095+    python aggregate_benchmark.py <benchmark_dir>
1096+
1097+Example:
1098+    python aggregate_benchmark.py benchmarks/2026-01-15T10-30-00/
1099+
1100+The script supports two directory layouts:
1101+
1102+    Workspace layout (from skill-creator iterations):
1103+    <benchmark_dir>/
1104+    └── eval-N/
1105+        ├── with_skill/
1106+        │   ├── run-1/grading.json
1107+        │   └── run-2/grading.json
1108+        └── without_skill/
1109+            ├── run-1/grading.json
1110+            └── run-2/grading.json
1111+
1112+    Legacy layout (with runs/ subdirectory):
1113+    <benchmark_dir>/
1114+    └── runs/
1115+        └── eval-N/
1116+            ├── with_skill/
1117+            │   └── run-1/grading.json
1118+            └── without_skill/
1119+                └── run-1/grading.json
1120+"""
1121+
1122+import argparse
1123+import json
1124+import math
1125+import sys
1126+from datetime import datetime, timezone
1127+from pathlib import Path
1128+
1129+
1130+def calculate_stats(values: list[float]) -> dict:
1131+    """Calculate mean, stddev, min, max for a list of values."""
1132+    if not values:
1133+        return {"mean": 0.0, "stddev": 0.0, "min": 0.0, "max": 0.0}
1134+
1135+    n = len(values)
1136+    mean = sum(values) / n
1137+
1138+    if n > 1:
1139+        variance = sum((x - mean) ** 2 for x in values) / (n - 1)
1140+        stddev = math.sqrt(variance)
1141+    else:
1142+        stddev = 0.0
1143+
1144+    return {
1145+        "mean": round(mean, 4),
1146+        "stddev": round(stddev, 4),
1147+        "min": round(min(values), 4),
1148+        "max": round(max(values), 4)
1149+    }
1150+
1151+
1152+def load_run_results(benchmark_dir: Path) -> dict:
1153+    """
1154+    Load all run results from a benchmark directory.
1155+
1156+    Returns dict keyed by config name (e.g. "with_skill"/"without_skill",
1157+    or "new_skill"/"old_skill"), each containing a list of run results.
1158+    """
1159+    # Support both layouts: eval dirs directly under benchmark_dir, or under runs/
1160+    runs_dir = benchmark_dir / "runs"
1161+    if runs_dir.exists():
1162+        search_dir = runs_dir
1163+    elif list(benchmark_dir.glob("eval-*")):
1164+        search_dir = benchmark_dir
1165+    else:
1166+        print(f"No eval directories found in {benchmark_dir} or {benchmark_dir / 'runs'}")
1167+        return {}
1168+
1169+    results: dict[str, list] = {}
1170+
1171+    for eval_idx, eval_dir in enumerate(sorted(search_dir.glob("eval-*"))):
1172+        metadata_path = eval_dir / "eval_metadata.json"
1173+        if metadata_path.exists():
1174+            try:
1175+                with open(metadata_path) as mf:
1176+                    eval_id = json.load(mf).get("eval_id", eval_idx)
1177+            except (json.JSONDecodeError, OSError):
1178+                eval_id = eval_idx
1179+        else:
1180+            try:
1181+                eval_id = int(eval_dir.name.split("-")[1])
1182+            except ValueError:
1183+                eval_id = eval_idx
1184+
1185diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py b/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py
1186new file mode 100644
1187index 0000000000000000000000000000000000000000..959e30a0014ec165c41a2bb7420b7dfe1416bbac
1188--- /dev/null
1189+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py
1190@@ -0,0 +1,326 @@
1191+#!/usr/bin/env python3
1192+"""Generate an HTML report from run_loop.py output.
1193+
1194+Takes the JSON output from run_loop.py and generates a visual HTML report
1195+showing each description attempt with check/x for each test case.
1196+Distinguishes between train and test queries.
1197+"""
1198+
1199+import argparse
1200+import html
1201+import json
1202+import sys
1203+from pathlib import Path
1204+
1205+
1206+def generate_html(data: dict, auto_refresh: bool = False, skill_name: str = "") -> str:
1207+    """Generate HTML report from loop output data. If auto_refresh is True, adds a meta refresh tag."""
1208+    history = data.get("history", [])
1209+    holdout = data.get("holdout", 0)
1210+    title_prefix = html.escape(skill_name + " \u2014 ") if skill_name else ""
1211+
1212+    # Get all unique queries from train and test sets, with should_trigger info
1213+    train_queries: list[dict] = []
1214+    test_queries: list[dict] = []
1215+    if history:
1216+        for r in history[0].get("train_results", history[0].get("results", [])):
1217+            train_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
1218+        if history[0].get("test_results"):
1219+            for r in history[0].get("test_results", []):
1220+                test_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
1221+
1222+    refresh_tag = '    <meta http-equiv="refresh" content="5">\n' if auto_refresh else ""
1223+
1224+    html_parts = ["""<!DOCTYPE html>
1225+<html>
1226+<head>
1227+    <meta charset="utf-8">
1228+""" + refresh_tag + """    <title>""" + title_prefix + """Skill Description Optimization</title>
1229+    <link rel="preconnect" href="https://fonts.googleapis.com">
1230+    <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
1231+    <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
1232+    <style>
1233+        body {
1234+            font-family: 'Lora', Georgia, serif;
1235+            max-width: 100%;
1236+            margin: 0 auto;
1237+            padding: 20px;
1238+            background: #faf9f5;
1239+            color: #141413;
1240+        }
1241+        h1 { font-family: 'Poppins', sans-serif; color: #141413; }
1242+        .explainer {
1243+            background: white;
1244+            padding: 15px;
1245+            border-radius: 6px;
1246+            margin-bottom: 20px;
1247+            border: 1px solid #e8e6dc;
1248+            color: #b0aea5;
1249+            font-size: 0.875rem;
1250+            line-height: 1.6;
1251+        }
1252+        .summary {
1253+            background: white;
1254+            padding: 15px;
1255+            border-radius: 6px;
1256+            margin-bottom: 20px;
1257+            border: 1px solid #e8e6dc;
1258+        }
1259+        .summary p { margin: 5px 0; }
1260+        .best { color: #788c5d; font-weight: bold; }
1261+        .table-container {
1262+            overflow-x: auto;
1263+            width: 100%;
1264+        }
1265+        table {
1266+            border-collapse: collapse;
1267+            background: white;
1268+            border: 1px solid #e8e6dc;
1269+            border-radius: 6px;
1270+            font-size: 12px;
1271+            min-width: 100%;
1272+        }
1273+        th, td {
1274+            padding: 8px;
1275+            text-align: left;
1276+            border: 1px solid #e8e6dc;
1277+            white-space: normal;
1278+            word-wrap: break-word;
1279+        }
1280+        th {
1281+            font-family: 'Poppins', sans-serif;
1282+            background: #141413;
1283+            color: #faf9f5;
1284+            font-weight: 500;
1285+        }
1286+        th.test-col {
1287+            background: #6a9bcc;
1288+        }
1289+        th.query-col { min-width: 200px; }
1290diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py b/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py
1291new file mode 100644
1292index 0000000000000000000000000000000000000000..12bbdb8635073a7f84f30741fc9795ef7952f9c5
1293--- /dev/null
1294+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py
1295@@ -0,0 +1,261 @@
1296+#!/usr/bin/env python3
1297+"""Improve a skill description based on eval results.
1298+
1299+Takes eval results (from run_eval.py) and generates an improved description
1300+using opencode with haiku 4.5.
1301+"""
1302+
1303+import argparse
1304+import json
1305+import re
1306+import subprocess
1307+import sys
1308+from pathlib import Path
1309+
1310+from scripts.utils import parse_skill_md
1311+
1312+IMPROVE_MODEL = "anthropic/claude-haiku-4-5"
1313+
1314+
1315+def _run_opencode(prompt: str, model: str) -> str:
1316+    """Run opencode run with a prompt and return the assistant text response."""
1317+    env_clean = {
1318+        k: v
1319+        for k, v in __import__("os").environ.items()
1320+        if k not in ("OPENCODE", "OPENCODE_PID")
1321+    }
1322+    result = subprocess.run(
1323+        ["opencode", "run", "--format", "json", "--model", model, prompt],
1324+        capture_output=True,
1325+        text=True,
1326+        env=env_clean,
1327+    )
1328+    # Parse JSON event stream, collect text parts
1329+    text_parts = []
1330+    for line in result.stdout.splitlines():
1331+        line = line.strip()
1332+        if not line:
1333+            continue
1334+        try:
1335+            event = json.loads(line)
1336+        except json.JSONDecodeError:
1337+            continue
1338+        if event.get("type") == "text":
1339+            part = event.get("part", {})
1340+            text_parts.append(part.get("text", ""))
1341+    return "".join(text_parts)
1342+
1343+
1344+def improve_description(
1345+    skill_name: str,
1346+    skill_content: str,
1347+    current_description: str,
1348+    eval_results: dict,
1349+    history: list[dict],
1350+    model: str = IMPROVE_MODEL,
1351+    test_results: dict | None = None,
1352+    log_dir: Path | None = None,
1353+    iteration: int | None = None,
1354+) -> str:
1355+    """Call opencode to improve the description based on eval results."""
1356+    failed_triggers = [
1357+        r for r in eval_results["results"] if r["should_trigger"] and not r["pass"]
1358+    ]
1359+    false_triggers = [
1360+        r for r in eval_results["results"] if not r["should_trigger"] and not r["pass"]
1361+    ]
1362+
1363+    # Build scores summary
1364+    train_score = (
1365+        f"{eval_results['summary']['passed']}/{eval_results['summary']['total']}"
1366+    )
1367+    if test_results:
1368+        test_score = (
1369+            f"{test_results['summary']['passed']}/{test_results['summary']['total']}"
1370+        )
1371+        scores_summary = f"Train: {train_score}, Test: {test_score}"
1372+    else:
1373+        scores_summary = f"Train: {train_score}"
1374+
1375diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py
1376new file mode 100644
1377index 0000000000000000000000000000000000000000..4a6fe2624c5c2c3e723e9ad58005fd9e361e9499
1378--- /dev/null
1379+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py
1380@@ -0,0 +1,304 @@
1381+#!/usr/bin/env python3
1382+"""Run trigger evaluation for a skill description.
1383+
1384+Tests whether a skill's description causes opencode to trigger (read the skill)
1385+for a set of queries. Outputs results as JSON.
1386+"""
1387+
1388+import argparse
1389+import json
1390+import os
1391+import select
1392+import subprocess
1393+import sys
1394+import time
1395+import uuid
1396+from concurrent.futures import ProcessPoolExecutor, as_completed
1397+from pathlib import Path
1398+
1399+from scripts.utils import parse_skill_md
1400+
1401+EVAL_MODEL = "anthropic/claude-sonnet-4-6"
1402+
1403+
1404+def find_project_root() -> Path:
1405+    """Find the project root by walking up from cwd looking for .opencode/.
1406+
1407+    Mimics how opencode discovers its project root, so the command file
1408+    we create ends up where opencode will look for it.
1409+    """
1410+    current = Path.cwd()
1411+    for parent in [current, *current.parents]:
1412+        if (parent / ".opencode").is_dir():
1413+            return parent
1414+    return current
1415+
1416+
1417+def run_single_query(
1418+    query: str,
1419+    skill_name: str,
1420+    skill_description: str,
1421+    timeout: int,
1422+    project_root: str,
1423+    model: str | None = None,
1424+) -> bool:
1425+    """Run a single query and return whether the skill was triggered.
1426+
1427+    Creates a skill file in .opencode/skills/ so it appears in opencode's
1428+    available_skills list, then runs `opencode run` with the raw query.
1429+    Parses JSON events to detect whether the skill tool was invoked.
1430+    """
1431+    unique_id = uuid.uuid4().hex[:8]
1432+    clean_name = f"{skill_name}-skill-{unique_id}"
1433+    project_skills_dir = Path(project_root) / ".opencode" / "skills" / clean_name
1434+    skill_file = project_skills_dir / "SKILL.md"
1435+
1436+    try:
1437+        project_skills_dir.mkdir(parents=True, exist_ok=True)
1438+        # Use YAML block scalar to avoid breaking on quotes in description
1439+        indented_desc = "\n  ".join(skill_description.split("\n"))
1440+        skill_content = (
1441+            f"---\n"
1442+            f"name: {clean_name}\n"
1443+            f"description: |\n"
1444+            f"  {indented_desc}\n"
1445+            f"---\n\n"
1446+            f"# {skill_name}\n\n"
1447+            f"This skill handles: {skill_description}\n"
1448+        )
1449+        skill_file.write_text(skill_content)
1450+
1451+        cmd = [
1452+            "opencode",
1453+            "run",
1454+            "--format",
1455+            "json",
1456+            "--model",
1457+            model or EVAL_MODEL,
1458+            query,
1459+        ]
1460+
1461+        # Remove OPENCODE env var to allow nesting opencode run inside an
1462+        # opencode session. The guard is for interactive terminal conflicts;
1463+        # programmatic subprocess usage is safe.
1464+        env = {
1465+            k: v for k, v in os.environ.items() if k not in ("OPENCODE", "OPENCODE_PID")
1466+        }
1467+
1468+        process = subprocess.Popen(
1469+            cmd,
1470+            stdout=subprocess.PIPE,
1471+            stderr=subprocess.DEVNULL,
1472+            cwd=project_root,
1473+            env=env,
1474+        )
1475+
1476+        triggered = False
1477+        start_time = time.time()
1478+        buffer = ""
1479+
1480diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py
1481new file mode 100644
1482index 0000000000000000000000000000000000000000..a18ba9b7544d573371420e31f0070452c324d75e
1483--- /dev/null
1484+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py
1485@@ -0,0 +1,404 @@
1486+#!/usr/bin/env python3
1487+"""Run the eval + improve loop until all pass or max iterations reached.
1488+
1489+Combines run_eval.py and improve_description.py in a loop, tracking history
1490+and returning the best description found. Supports train/test split to prevent
1491+overfitting.
1492+"""
1493+
1494+import argparse
1495+import json
1496+import random
1497+import sys
1498+import tempfile
1499+import time
1500+import webbrowser
1501+from pathlib import Path
1502+
1503+from scripts.generate_report import generate_html
1504+from scripts.improve_description import improve_description, IMPROVE_MODEL
1505+from scripts.run_eval import find_project_root, run_eval, EVAL_MODEL
1506+from scripts.utils import parse_skill_md
1507+
1508+
1509+def split_eval_set(
1510+    eval_set: list[dict], holdout: float, seed: int = 42
1511+) -> tuple[list[dict], list[dict]]:
1512+    """Split eval set into train and test sets, stratified by should_trigger."""
1513+    random.seed(seed)
1514+
1515+    # Separate by should_trigger
1516+    trigger = [e for e in eval_set if e["should_trigger"]]
1517+    no_trigger = [e for e in eval_set if not e["should_trigger"]]
1518+
1519+    # Shuffle each group
1520+    random.shuffle(trigger)
1521+    random.shuffle(no_trigger)
1522+
1523+    # Calculate split points
1524+    n_trigger_test = max(1, int(len(trigger) * holdout))
1525+    n_no_trigger_test = max(1, int(len(no_trigger) * holdout))
1526+
1527+    # Split
1528+    test_set = trigger[:n_trigger_test] + no_trigger[:n_no_trigger_test]
1529+    train_set = trigger[n_trigger_test:] + no_trigger[n_no_trigger_test:]
1530+
1531+    return train_set, test_set
1532+
1533+
1534+def run_loop(
1535+    eval_set: list[dict],
1536+    skill_path: Path,
1537+    description_override: str | None,
1538+    num_workers: int,
1539+    timeout: int,
1540+    max_iterations: int,
1541+    runs_per_query: int,
1542+    trigger_threshold: float,
1543+    holdout: float,
1544+    model: str,
1545+    verbose: bool,
1546+    live_report_path: Path | None = None,
1547+    log_dir: Path | None = None,
1548+) -> dict:
1549+    """Run the eval + improvement loop."""
1550+    project_root = find_project_root()
1551+    name, original_description, content = parse_skill_md(skill_path)
1552+    current_description = description_override or original_description
1553+
1554+    # Split into train/test if holdout > 0
1555+    if holdout > 0:
1556+        train_set, test_set = split_eval_set(eval_set, holdout)
1557+        if verbose:
1558+            print(
1559+                f"Split: {len(train_set)} train, {len(test_set)} test (holdout={holdout})",
1560+                file=sys.stderr,
1561+            )
1562+    else:
1563+        train_set = eval_set
1564+        test_set = []
1565+
1566+    history = []
1567+    exit_reason = "unknown"
1568+
1569+    for iteration in range(1, max_iterations + 1):
1570+        if verbose:
1571+            print(f"\n{'=' * 60}", file=sys.stderr)
1572+            print(f"Iteration {iteration}/{max_iterations}", file=sys.stderr)
1573+            print(f"Description: {current_description}", file=sys.stderr)
1574+            print(f"{'=' * 60}", file=sys.stderr)
1575+
1576+        # Evaluate train + test together in one batch for parallelism
1577+        all_queries = train_set + test_set
1578+        t0 = time.time()
1579+        all_results = run_eval(
1580+            eval_set=all_queries,
1581+            skill_name=name,
1582+            description=current_description,
1583+            num_workers=num_workers,
1584+            timeout=timeout,
1585diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py b/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py
1586new file mode 100644
1587index 0000000000000000000000000000000000000000..f48eac444656ddc41204aac1760a217951ce609e
1588--- /dev/null
1589+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py
1590@@ -0,0 +1,136 @@
1591+#!/usr/bin/env python3
1592+"""
1593+Skill Packager - Creates a distributable .skill file of a skill folder
1594+
1595+Usage:
1596+    python utils/package_skill.py <path/to/skill-folder> [output-directory]
1597+
1598+Example:
1599+    python utils/package_skill.py skills/public/my-skill
1600+    python utils/package_skill.py skills/public/my-skill ./dist
1601+"""
1602+
1603+import fnmatch
1604+import sys
1605+import zipfile
1606+from pathlib import Path
1607+from scripts.quick_validate import validate_skill
1608+
1609+# Patterns to exclude when packaging skills.
1610+EXCLUDE_DIRS = {"__pycache__", "node_modules"}
1611+EXCLUDE_GLOBS = {"*.pyc"}
1612+EXCLUDE_FILES = {".DS_Store"}
1613+# Directories excluded only at the skill root (not when nested deeper).
1614+ROOT_EXCLUDE_DIRS = {"evals"}
1615+
1616+
1617+def should_exclude(rel_path: Path) -> bool:
1618+    """Check if a path should be excluded from packaging."""
1619+    parts = rel_path.parts
1620+    if any(part in EXCLUDE_DIRS for part in parts):
1621+        return True
1622+    # rel_path is relative to skill_path.parent, so parts[0] is the skill
1623+    # folder name and parts[1] (if present) is the first subdir.
1624+    if len(parts) > 1 and parts[1] in ROOT_EXCLUDE_DIRS:
1625+        return True
1626+    name = rel_path.name
1627+    if name in EXCLUDE_FILES:
1628+        return True
1629+    return any(fnmatch.fnmatch(name, pat) for pat in EXCLUDE_GLOBS)
1630+
1631+
1632+def package_skill(skill_path, output_dir=None):
1633+    """
1634+    Package a skill folder into a .skill file.
1635+
1636+    Args:
1637+        skill_path: Path to the skill folder
1638+        output_dir: Optional output directory for the .skill file (defaults to current directory)
1639+
1640+    Returns:
1641+        Path to the created .skill file, or None if error
1642+    """
1643+    skill_path = Path(skill_path).resolve()
1644+
1645+    # Validate skill folder exists
1646+    if not skill_path.exists():
1647+        print(f"❌ Error: Skill folder not found: {skill_path}")
1648+        return None
1649+
1650+    if not skill_path.is_dir():
1651+        print(f"❌ Error: Path is not a directory: {skill_path}")
1652+        return None
1653+
1654+    # Validate SKILL.md exists
1655+    skill_md = skill_path / "SKILL.md"
1656+    if not skill_md.exists():
1657+        print(f"❌ Error: SKILL.md not found in {skill_path}")
1658+        return None
1659+
1660+    # Run validation before packaging
1661+    print("🔍 Validating skill...")
1662+    valid, message = validate_skill(skill_path)
1663+    if not valid:
1664+        print(f"❌ Validation failed: {message}")
1665+        print("   Please fix the validation errors before packaging.")
1666+        return None
1667+    print(f"✅ {message}\n")
1668+
1669+    # Determine output location
1670+    skill_name = skill_path.name
1671+    if output_dir:
1672+        output_path = Path(output_dir).resolve()
1673+        output_path.mkdir(parents=True, exist_ok=True)
1674+    else:
1675+        output_path = Path.cwd()
1676+
1677+    skill_filename = output_path / f"{skill_name}.skill"
1678+
1679+    # Create the .skill file (zip format)
1680+    try:
1681+        with zipfile.ZipFile(skill_filename, 'w', zipfile.ZIP_DEFLATED) as zipf:
1682+            # Walk through the skill directory, excluding build artifacts
1683+            for file_path in skill_path.rglob('*'):
1684+                if not file_path.is_file():
1685+                    continue
1686+                arcname = file_path.relative_to(skill_path.parent)
1687+                if should_exclude(arcname):
1688+                    print(f"  Skipped: {arcname}")
1689+                    continue
1690diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py b/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py
1691new file mode 100644
1692index 0000000000000000000000000000000000000000..ed8e1dddce77b16af13c6f36b3fe86c4ac7c590c
1693--- /dev/null
1694+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py
1695@@ -0,0 +1,103 @@
1696+#!/usr/bin/env python3
1697+"""
1698+Quick validation script for skills - minimal version
1699+"""
1700+
1701+import sys
1702+import os
1703+import re
1704+import yaml
1705+from pathlib import Path
1706+
1707+def validate_skill(skill_path):
1708+    """Basic validation of a skill"""
1709+    skill_path = Path(skill_path)
1710+
1711+    # Check SKILL.md exists
1712+    skill_md = skill_path / 'SKILL.md'
1713+    if not skill_md.exists():
1714+        return False, "SKILL.md not found"
1715+
1716+    # Read and validate frontmatter
1717+    content = skill_md.read_text()
1718+    if not content.startswith('---'):
1719+        return False, "No YAML frontmatter found"
1720+
1721+    # Extract frontmatter
1722+    match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
1723+    if not match:
1724+        return False, "Invalid frontmatter format"
1725+
1726+    frontmatter_text = match.group(1)
1727+
1728+    # Parse YAML frontmatter
1729+    try:
1730+        frontmatter = yaml.safe_load(frontmatter_text)
1731+        if not isinstance(frontmatter, dict):
1732+            return False, "Frontmatter must be a YAML dictionary"
1733+    except yaml.YAMLError as e:
1734+        return False, f"Invalid YAML in frontmatter: {e}"
1735+
1736+    # Define allowed properties
1737+    ALLOWED_PROPERTIES = {'name', 'description', 'license', 'allowed-tools', 'metadata', 'compatibility'}
1738+
1739+    # Check for unexpected properties (excluding nested keys under metadata)
1740+    unexpected_keys = set(frontmatter.keys()) - ALLOWED_PROPERTIES
1741+    if unexpected_keys:
1742+        return False, (
1743+            f"Unexpected key(s) in SKILL.md frontmatter: {', '.join(sorted(unexpected_keys))}. "
1744+            f"Allowed properties are: {', '.join(sorted(ALLOWED_PROPERTIES))}"
1745+        )
1746+
1747+    # Check required fields
1748+    if 'name' not in frontmatter:
1749+        return False, "Missing 'name' in frontmatter"
1750+    if 'description' not in frontmatter:
1751+        return False, "Missing 'description' in frontmatter"
1752+
1753+    # Extract name for validation
1754+    name = frontmatter.get('name', '')
1755+    if not isinstance(name, str):
1756+        return False, f"Name must be a string, got {type(name).__name__}"
1757+    name = name.strip()
1758+    if name:
1759+        # Check naming convention (kebab-case: lowercase with hyphens)
1760+        if not re.match(r'^[a-z0-9-]+$', name):
1761+            return False, f"Name '{name}' should be kebab-case (lowercase letters, digits, and hyphens only)"
1762+        if name.startswith('-') or name.endswith('-') or '--' in name:
1763+            return False, f"Name '{name}' cannot start/end with hyphen or contain consecutive hyphens"
1764+        # Check name length (max 64 characters per spec)
1765+        if len(name) > 64:
1766+            return False, f"Name is too long ({len(name)} characters). Maximum is 64 characters."
1767+
1768+    # Extract and validate description
1769+    description = frontmatter.get('description', '')
1770+    if not isinstance(description, str):
1771+        return False, f"Description must be a string, got {type(description).__name__}"
1772+    description = description.strip()
1773+    if description:
1774+        # Check for angle brackets
1775+        if '<' in description or '>' in description:
1776+            return False, "Description cannot contain angle brackets (< or >)"
1777+        # Check description length (max 1024 characters per spec)
1778+        if len(description) > 1024:
1779+            return False, f"Description is too long ({len(description)} characters). Maximum is 1024 characters."
1780+
1781+    # Validate compatibility field if present (optional)
1782+    compatibility = frontmatter.get('compatibility', '')
1783+    if compatibility:
1784+        if not isinstance(compatibility, str):
1785+            return False, f"Compatibility must be a string, got {type(compatibility).__name__}"
1786+        if len(compatibility) > 500:
1787+            return False, f"Compatibility is too long ({len(compatibility)} characters). Maximum is 500 characters."
1788+
1789+    return True, "Skill is valid!"
1790+
1791+if __name__ == "__main__":
1792+    if len(sys.argv) != 2:
1793+        print("Usage: python quick_validate.py <skill_directory>")
1794+        sys.exit(1)
1795diff --git a/dot_config/opencode/skills/skill-creator/scripts/utils.py b/dot_config/opencode/skills/skill-creator/scripts/utils.py
1796new file mode 100644
1797index 0000000000000000000000000000000000000000..51b6a07dd57174197a937034b7eecebd5768ff8a
1798--- /dev/null
1799+++ b/dot_config/opencode/skills/skill-creator/scripts/utils.py
1800@@ -0,0 +1,47 @@
1801+"""Shared utilities for skill-creator scripts."""
1802+
1803+from pathlib import Path
1804+
1805+
1806+
1807+def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
1808+    """Parse a SKILL.md file, returning (name, description, full_content)."""
1809+    content = (skill_path / "SKILL.md").read_text()
1810+    lines = content.split("\n")
1811+
1812+    if lines[0].strip() != "---":
1813+        raise ValueError("SKILL.md missing frontmatter (no opening ---)")
1814+
1815+    end_idx = None
1816+    for i, line in enumerate(lines[1:], start=1):
1817+        if line.strip() == "---":
1818+            end_idx = i
1819+            break
1820+
1821+    if end_idx is None:
1822+        raise ValueError("SKILL.md missing frontmatter (no closing ---)")
1823+
1824+    name = ""
1825+    description = ""
1826+    frontmatter_lines = lines[1:end_idx]
1827+    i = 0
1828+    while i < len(frontmatter_lines):
1829+        line = frontmatter_lines[i]
1830+        if line.startswith("name:"):
1831+            name = line[len("name:"):].strip().strip('"').strip("'")
1832+        elif line.startswith("description:"):
1833+            value = line[len("description:"):].strip()
1834+            # Handle YAML multiline indicators (>, |, >-, |-)
1835+            if value in (">", "|", ">-", "|-"):
1836+                continuation_lines: list[str] = []
1837+                i += 1
1838+                while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith("  ") or frontmatter_lines[i].startswith("\t")):
1839+                    continuation_lines.append(frontmatter_lines[i].strip())
1840+                    i += 1
1841+                description = " ".join(continuation_lines)
1842+                continue
1843+            else:
1844+                description = value.strip('"').strip("'")
1845+        i += 1
1846+
1847+    return name, description, content
1848diff --git a/dot_config/opencode/skills/weekly-review/SKILL.md b/dot_config/opencode/skills/weekly-review/SKILL.md
1849new file mode 100644
1850index 0000000000000000000000000000000000000000..392dc89193dcdb582c3508c2767786e77632013b
1851--- /dev/null
1852+++ b/dot_config/opencode/skills/weekly-review/SKILL.md
1853@@ -0,0 +1,115 @@
1854+---
1855+name: weekly-review
1856+description: Analyze recent sessions to find patterns, recurring mistakes, and propose improvements to AGENTS.md and lessons.md.
1857+compatibility: opencode
1858+---
1859+
1860+## What this is
1861+
1862+A periodic deep review of how the user and assistant work together. Reads session
1863+history from the database, dispatches parallel agents to analyze conversations,
1864+then runs multiple reflection loops to surface patterns and propose changes.
1865+
1866+## When to trigger
1867+
1868+User runs `/review-week`. Not autonomous — this is a deliberate review.
1869+
1870+## Inputs
1871+
1872+- `$ARGUMENTS`: optional time range in days (default: 7). Example: `/review-week 14`
1873+
1874+## Data source
1875+
1876+Sessions are stored in SQLite at `~/.local/share/opencode/opencode.db`.
1877+
1878+Key tables:
1879+- `session`: metadata (id, title, directory, time_created, time_updated)
1880+- `message`: per-session messages (id, session_id, data JSON)
1881+- `part`: message content (id, message_id, session_id, data JSON with `type` and `text` fields)
1882+
1883+## Workflow
1884+
1885+### 1. Gather sessions
1886+
1887+```sql
1888+SELECT id, title, directory,
1889+  datetime(time_created/1000, 'unixepoch', 'localtime') as created,
1890+  datetime(time_updated/1000, 'unixepoch', 'localtime') as updated
1891+FROM session
1892+WHERE time_created > (strftime('%s', 'now', '-N days') * 1000)
1893+  AND title NOT LIKE '%@explore%'
1894+  AND title NOT LIKE '%@general%'
1895+ORDER BY time_created ASC;
1896+```
1897+
1898+Replace `N` with the requested day range. Filter out subagent sessions — they're
1899+noise for pattern analysis.
1900+
1901+Count messages per session to find the substantive ones (>3 messages).
1902+
1903+### 2. Dispatch parallel agents
1904+
1905+Group substantive sessions into 3-5 batches. For each batch, launch a `general`
1906+subagent with instructions to:
1907+
1908+1. Extract conversation text from the `part` table:
1909+   ```sql
1910+   SELECT p.data FROM part p
1911+   JOIN message m ON m.id = p.message_id
1912+   WHERE m.session_id = 'SESSION_ID'
1913+     AND json_extract(p.data, '$.type') = 'text'
1914+   ORDER BY p.time_created ASC;
1915+   ```
1916+2. For each session, identify:
1917+   - What the user was trying to accomplish
1918+   - Corrections the user made to the assistant
1919+   - Frustration signals
1920+   - Where the assistant over- or under-delivered
1921+   - Design decisions and their quality
1922+3. Return a comprehensive analysis focused on patterns and lessons, not summaries.
1923+
1924+### 3. Read context files
1925+
1926+While agents run, read:
1927+- `~/.config/opencode/AGENTS.md`
1928+- `~/.config/opencode/lessons.md`
1929+
1930+### 4. Reflection loops
1931+
1932+Run 2-3 passes over the combined agent output:
1933+
1934+**Loop 1 — Patterns:** What themes repeat across sessions? Which are the most
1935+frequent and most damaging? Cross-reference against existing AGENTS.md and
1936+lessons.md rules — are these known problems (compliance failure) or new gaps?
1937+
1938+**Loop 2 — Design problems:** Why does the system produce these failures?
1939+What's the root cause behind the patterns? Are existing rules too vague,
1940+misplaced, or structurally unenforceable?
1941+
1942+**Loop 3 — Proposals:** Concrete, minimal changes. Prefer strengthening existing
1943+rules over adding new ones. Every proposed change must map to a specific pattern
1944+found in the data.
1945+
1946+### 5. Present findings
1947+
1948+Show the user:
1949+- Session count and coverage
1950+- Top patterns with frequency and severity
1951+- The compliance question: which patterns already have rules?
1952+- Proposed changes as exact edits (old → new)