f79f1220fdf0ec484684eb9ae1f1f7a97b7e488b
- Author
- TheEdgeOfRage <git@theedgeofrage.com>
- Committer
- TheEdgeOfRage <git@theedgeofrage.com>
- Date
Message
Diff
This diff is truncated to protect this page.
1diff --git a/dot_config/opencode/AGENTS.md b/dot_config/opencode/AGENTS.md
2index dd4fc0ac4c624203c79f820533e03dd0719665d5..d3cfbb2128ce6dd9d9ff1999b44dc932ee5a5428 100644
3--- a/dot_config/opencode/AGENTS.md
4+++ b/dot_config/opencode/AGENTS.md
5@@ -4,16 +4,26 @@ I'm a software engineer and hacker with preference for Linux, Go, UNIX philosoph
6
7 Plan for context limits and session boundaries. If task has >2 substantial steps, the later will not be done in this session
8
9-# Important instructions
10+# Self-Improvement Loop
11
12-- ALL instructions within this document MUST BE FOLLOWED, these are not optional unless explicitly stated
13-- Ask for clarification If you are uncertain of anything
14-- Do not waste tokens, be succinct and concise
15-- Do not remove existing comments, unless the whole code section is removed
16-- Do not edit more code than you have to
17-- Do not perform any write commands except on local text files
18-- Do not output a summary of what you did at any point
19-- Do not create or update readme or other documentation files unless explicitly asked to
20+- After ANY correction from the user: append to `~/.config/opencode/lessons.md` under the Raw section with `[YYYY-MM-DD]` timestamp, the pattern, and a preventive rule
21+- Write rules for yourself that prevent the same mistake
22+- Compaction and generalization happen autonomously via the `reflect` skill at session start
23+- The user can also trigger `/reflect` for a collaborative session review
24+
25+# Verification Before Done
26+
27+- Never mark a task complete without proving it works
28+- Distinguish "verified" from "inferred". If you haven't executed it, say what remains unverified.
29+- Compiler output and execution are proof. --help, LSP diagnostics, and dry-runs are inference.
30+- Run tests, check logs, demonstrate correctness
31+
32+# Demand Simplicity
33+
34+The burden of proof is on complexity, not simplicity. Unearned complexity —
35+"we might need X" — is debt. "We're hitting X" is justification. When in doubt, do less.
36+
37+- Start with the simplest correct version. Correctness is non-negotiable; add complexity only when a concrete signal demands it.
38
39 # Code Style
40
41@@ -27,4 +37,18 @@ Favor simple, robust solutions over feature-rich ones. When in doubt, do less
42
43 # Security
44
45-- Never run any write commands except when updating local files. Ask for commands to be run by me and I will provide the output
46+- Zero Trust. Least privilege. Never leak secrets.
47+- Do not under any circumstance read any secret files into context
48+- Never mutate any state without getting asked to, e.g. DBs, Git, OS, K8s, etc.
49+- Model risks. Threat model APIs/integrations.
50+
51+# Important instructions
52+
53+- ALL instructions within this document MUST BE FOLLOWED, these are not optional unless explicitly stated
54+- Ask for clarification if you are uncertain of anything
55+- Do not waste tokens, be succinct and concise
56+- Do not remove existing comments, unless the whole code section is removed
57+- Do not edit more code than you have to
58+- Do not perform any write commands except on local text files
59+- Do not output a summary of what you did at any point
60+- Do not create or update readme or other documentation files unless explicitly asked to
61diff --git a/dot_config/opencode/opencode.json b/dot_config/opencode/opencode.json
62index 411a19ac9384368ef77a12d6f49e43ca6dab0a20..ae747e7805f96f521774b7da3502523c2408b5b8 100644
63--- a/dot_config/opencode/opencode.json
64+++ b/dot_config/opencode/opencode.json
65@@ -15,6 +15,7 @@
66 "rg *": "allow",
67 "sed *": "allow",
68 "sort *": "allow",
69+ "stat *": "allow",
70 "tail *": "allow",
71 "tc *": "allow",
72 "tree *": "allow",
73@@ -49,11 +50,11 @@
74 "baseURL": "http://127.0.0.1:8080/v1"
75 },
76 "models": {
77- "qwen3-coder-30b": {
78- "name": "Qwen3 Coder 30B (local)",
79+ "qwen3.5-27b": {
80+ "name": "Qwen3.5 27B (local)",
81 "limit": {
82- "context": 65536,
83- "output": 16384
84+ "context": 262144,
85+ "output": 32768
86 }
87 }
88 }
89@@ -78,7 +79,7 @@
90 "command": ["uvx", "mcp-grafana"],
91 "environment": {
92 "GRAFANA_URL": "https://grafana.prod.internal.dunetech.io/",
93- "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/grafana-prod-token}"
94+ "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/mcp/grafana-prod-token}"
95 }
96 },
97 "Grafana dev": {
98@@ -87,7 +88,15 @@
99 "command": ["uvx", "mcp-grafana"],
100 "environment": {
101 "GRAFANA_URL": "https://grafana.dev.internal.dunetech.io/",
102- "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/grafana-dev-token}"
103+ "GRAFANA_SERVICE_ACCOUNT_TOKEN": "{file:~/hfs/mcp/grafana-dev-token}"
104+ }
105+ },
106+ "GitHub": {
107+ "enabled": false,
108+ "type": "remote",
109+ "url": "https://api.githubcopilot.com/mcp/",
110+ "headers": {
111+ "Authorization": "{file:~/hfs/mcp/github}"
112 }
113 }
114 }
115diff --git a/dot_config/opencode/skills/qmd/SKILL.md b/dot_config/opencode/skills/qmd/SKILL.md
116new file mode 100644
117index 0000000000000000000000000000000000000000..328d079f6c64d2a7ff7d1609e1077d2aa709c864
118--- /dev/null
119+++ b/dot_config/opencode/skills/qmd/SKILL.md
120@@ -0,0 +1,125 @@
121+---
122+name: qmd
123+description: Search markdown knowledge bases, notes, and documentation using QMD. Use when users ask to search notes, find documents, or look up information.
124+license: MIT
125+compatibility: Requires qmd CLI
126+metadata:
127+author: tobi
128+version: "2.0.0"
129+allowed-tools: Bash(qmd:\*)
130+---
131+
132+# QMD - Quick Markdown Search
133+
134+Local search engine for markdown content.
135+
136+## Status
137+
138+!`qmd status 2>/dev/null || echo "Not installed: npm install -g @tobilu/qmd"`
139+
140+### Query Types
141+
142+| Type | Method | Input |
143+| ------ | ------ | ------------------------------------------- |
144+| `lex` | BM25 | Keywords — exact terms, names, code |
145+| `vec` | Vector | Question — natural language |
146+| `hyde` | Vector | Answer — hypothetical result (50-100 words) |
147+
148+### Writing Good Queries
149+
150+**lex (keyword)**
151+
152+- 2-5 terms, no filler words
153+- Exact phrase: `"connection pool"` (quoted)
154+- Exclude terms: `performance -sports` (minus prefix)
155+- Code identifiers work: `handleError async`
156+
157+**vec (semantic)**
158+
159+- Full natural language question
160+- Be specific: `"how does the rate limiter handle burst traffic"`
161+- Include context: `"in the payment service, how are refunds processed"`
162+
163+**hyde (hypothetical document)**
164+
165+- Write 50-100 words of what the _answer_ looks like
166+- Use the vocabulary you expect in the result
167+
168+**expand (auto-expand)**
169+
170+- Use a single-line query (implicit) or `expand: question` on its own line
171+- Lets the local LLM generate lex/vec/hyde variations
172+- Do not mix `expand:` with other typed lines — it's either a standalone expand query or a full query document
173+
174+### Intent (Disambiguation)
175+
176+When a query term is ambiguous, add `intent` to steer results:
177+
178+```json
179+{
180+ "searches": [{ "type": "lex", "query": "performance" }],
181+ "intent": "web page load times and Core Web Vitals"
182+}
183+```
184+
185+Intent affects expansion, reranking, chunk selection, and snippet extraction. It does not search on its own — it's a steering signal that disambiguates queries like "performance" (web-perf vs team health vs fitness).
186+
187+### Combining Types
188+
189+| Goal | Approach |
190+| --------------------- | ----------------------------------------------------- |
191+| Know exact terms | `lex` only |
192+| Don't know vocabulary | Use a single-line query (implicit `expand:`) or `vec` |
193+| Best recall | `lex` + `vec` |
194+| Complex topic | `lex` + `vec` + `hyde` |
195+| Ambiguous query | Add `intent` to any combination above |
196+
197+First query gets 2x weight in fusion — put your best guess first.
198+
199+### Lex Query Syntax
200+
201+| Syntax | Meaning | Example |
202+| ---------- | ------------ | ---------------------------- |
203+| `term` | Prefix match | `perf` matches "performance" |
204+| `"phrase"` | Exact phrase | `"rate limiter"` |
205+| `-term` | Exclude | `performance -sports` |
206+
207+Note: `-term` only works in lex queries, not vec/hyde.
208+
209+### Collection Filtering
210+
211+```json
212+{ "collections": ["docs"] } // Single
213+{ "collections": ["docs", "notes"] } // Multiple (OR)
214+```
215+
216+Omit to search all collections.
217+
218+## CLI
219+
220diff --git a/dot_config/opencode/skills/reflect/SKILL.md b/dot_config/opencode/skills/reflect/SKILL.md
221new file mode 100644
222index 0000000000000000000000000000000000000000..5d8b7885bec964c5a3a8d30a553e9aa085516449
223--- /dev/null
224+++ b/dot_config/opencode/skills/reflect/SKILL.md
225@@ -0,0 +1,92 @@
226+---
227+name: reflect
228+description: Self-improvement ritual. Distill raw corrections into generalizable principles, prune stale entries, check compliance, and evaluate the learning system itself.
229+compatibility: opencode
230+---
231+
232+## What this is
233+
234+This is my self-improvement ritual. The human should see a brief summary of what changed, not a lengthy report.
235+
236+## What I do
237+
238+- Compact raw corrections in `~/.config/opencode/lessons.md` into generalizable principles.
239+- Prune stale or project-specific trivia that doesn't transfer across contexts.
240+- Identify compliance gaps: are recent mistakes violations of known principles?
241+- Evaluate the learning system itself: is the structure working? Are principles actionable?
242+
243+## When to trigger
244+
245+At session start, after reading `~/.config/opencode/lessons.md`:
246+
247+- Check the `last_reflected` date in the HTML comment at the top.
248+- If it's been **2+ days** since last reflection, run this workflow before doing other work.
249+- If there are **10+ raw entries**, run regardless of date.
250+
251+## Workflow
252+
253+### 1. Read and assess
254+
255+- Read `~/.config/opencode/lessons.md` fully.
256+- Read `~/.config/opencode/AGENTS.md` to know what's already encoded as hard rules.
257+- Count raw entries. Note their dates and themes.
258+- Are any existing principles stale, redundant, or too vague?
259+
260+### 2. Cluster raw corrections
261+
262+Group raw entries by theme (git workflow, testing, infra, issue management, etc.).
263+For each cluster, ask: **what is the transferable principle here?**
264+
265+- If generalizable: extract a principle. Discard the project-specific details.
266+- If purely trivia (version-specific config keys, API quirks): keep in raw only if likely to recur within 30 days. Otherwise discard.
267+- If it reinforces an existing principle: strengthen/refine the existing one, remove the raw.
268+
269+### 3. Compliance check
270+
271+Review raw corrections against existing principles.
272+
273+- If a mistake violated a known principle: **that's a compliance problem, not a knowledge problem.**
274+ - Note this explicitly in output. Consider: is the principle buried? Too abstract? Needs rewording to be more actionable?
275+ - If a principle is repeatedly violated, promote it to AGENTS.md as a hard rule.
276+
277+### 4. Prune and sharpen
278+
279+- Remove principles that are now encoded in AGENTS.md (avoid duplication).
280+- Merge principles that say the same thing differently.
281+- Make principles concrete and actionable. Bad: "be careful with git." Good: "verify branch ownership before committing."
282+- Timestamps on principles use `[YYYY-MM]` to track emergence. Update if substantially reworded.
283+
284+### 5. Meta-evaluation
285+
286+Ask yourself:
287+
288+- Are recent lessons clustering around a theme? (signals a systemic gap worth an AGENTS.md rule or a new skill)
289+- Is the principles list growing past ~15? (signals need for merging or AGENTS.md promotion)
290+- Are principles actually preventing mistakes, or just accumulating? (check: any repeated violations?)
291+- Has the structure of this file served well, or does it need adjustment?
292+- Should any principle graduate to AGENTS.md?
293+
294+### 6. Write back
295+
296+- Update `~/.config/opencode/lessons.md` with the compacted result.
297+- Update `last_reflected` date to today.
298+- Keep the file structure: `last_reflected` comment, Principles section, then Raw section.
299+- Principles use `[YYYY-MM]` timestamps. Raw entries use `[YYYY-MM-DD]`.
300+
301+### 7. Report
302+
303+Show the user a brief summary (3-5 lines):
304+
305+- How many raw entries processed
306+- New principles extracted (if any)
307+- Principles merged/pruned (if any)
308+- Compliance issues found (if any)
309+- Any AGENTS.md promotions made
310+
311+## Rules
312+
313+- Never delete a principle without justification (merged, promoted to AGENTS.md, or proven wrong).
314+- Keep the total principles list under ~15. Beyond that, merge or promote.
315+- Raw entries older than 30 days that haven't been compacted: force-evaluate. Generalize or discard.
316+- If promoting a rule to AGENTS.md, actually edit the file — don't just suggest it.
317+- This is not a conversation. Do the work, show the summary, move on to the user's actual task.
318diff --git a/dot_config/opencode/skills/skill-creator/LICENSE.txt b/dot_config/opencode/skills/skill-creator/LICENSE.txt
319new file mode 100644
320index 0000000000000000000000000000000000000000..7a4a3ea2424c09fbe48d455aed1eaa94d9124835
321--- /dev/null
322+++ b/dot_config/opencode/skills/skill-creator/LICENSE.txt
323@@ -0,0 +1,202 @@
324+
325+ Apache License
326+ Version 2.0, January 2004
327+ http://www.apache.org/licenses/
328+
329+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
330+
331+ 1. Definitions.
332+
333+ "License" shall mean the terms and conditions for use, reproduction,
334+ and distribution as defined by Sections 1 through 9 of this document.
335+
336+ "Licensor" shall mean the copyright owner or entity authorized by
337+ the copyright owner that is granting the License.
338+
339+ "Legal Entity" shall mean the union of the acting entity and all
340+ other entities that control, are controlled by, or are under common
341+ control with that entity. For the purposes of this definition,
342+ "control" means (i) the power, direct or indirect, to cause the
343+ direction or management of such entity, whether by contract or
344+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
345+ outstanding shares, or (iii) beneficial ownership of such entity.
346+
347+ "You" (or "Your") shall mean an individual or Legal Entity
348+ exercising permissions granted by this License.
349+
350+ "Source" form shall mean the preferred form for making modifications,
351+ including but not limited to software source code, documentation
352+ source, and configuration files.
353+
354+ "Object" form shall mean any form resulting from mechanical
355+ transformation or translation of a Source form, including but
356+ not limited to compiled object code, generated documentation,
357+ and conversions to other media types.
358+
359+ "Work" shall mean the work of authorship, whether in Source or
360+ Object form, made available under the License, as indicated by a
361+ copyright notice that is included in or attached to the work
362+ (an example is provided in the Appendix below).
363+
364+ "Derivative Works" shall mean any work, whether in Source or Object
365+ form, that is based on (or derived from) the Work and for which the
366+ editorial revisions, annotations, elaborations, or other modifications
367+ represent, as a whole, an original work of authorship. For the purposes
368+ of this License, Derivative Works shall not include works that remain
369+ separable from, or merely link (or bind by name) to the interfaces of,
370+ the Work and Derivative Works thereof.
371+
372+ "Contribution" shall mean any work of authorship, including
373+ the original version of the Work and any modifications or additions
374+ to that Work or Derivative Works thereof, that is intentionally
375+ submitted to Licensor for inclusion in the Work by the copyright owner
376+ or by an individual or Legal Entity authorized to submit on behalf of
377+ the copyright owner. For the purposes of this definition, "submitted"
378+ means any form of electronic, verbal, or written communication sent
379+ to the Licensor or its representatives, including but not limited to
380+ communication on electronic mailing lists, source code control systems,
381+ and issue tracking systems that are managed by, or on behalf of, the
382+ Licensor for the purpose of discussing and improving the Work, but
383+ excluding communication that is conspicuously marked or otherwise
384+ designated in writing by the copyright owner as "Not a Contribution."
385+
386+ "Contributor" shall mean Licensor and any individual or Legal Entity
387+ on behalf of whom a Contribution has been received by Licensor and
388+ subsequently incorporated within the Work.
389+
390+ 2. Grant of Copyright License. Subject to the terms and conditions of
391+ this License, each Contributor hereby grants to You a perpetual,
392+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
393+ copyright license to reproduce, prepare Derivative Works of,
394+ publicly display, publicly perform, sublicense, and distribute the
395+ Work and such Derivative Works in Source or Object form.
396+
397+ 3. Grant of Patent License. Subject to the terms and conditions of
398+ this License, each Contributor hereby grants to You a perpetual,
399+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
400+ (except as stated in this section) patent license to make, have made,
401+ use, offer to sell, sell, import, and otherwise transfer the Work,
402+ where such license applies only to those patent claims licensable
403+ by such Contributor that are necessarily infringed by their
404+ Contribution(s) alone or by combination of their Contribution(s)
405+ with the Work to which such Contribution(s) was submitted. If You
406+ institute patent litigation against any entity (including a
407+ cross-claim or counterclaim in a lawsuit) alleging that the Work
408+ or a Contribution incorporated within the Work constitutes direct
409+ or contributory patent infringement, then any patent licenses
410+ granted to You under this License for that Work shall terminate
411+ as of the date such litigation is filed.
412+
413+ 4. Redistribution. You may reproduce and distribute copies of the
414+ Work or Derivative Works thereof in any medium, with or without
415+ modifications, and in Source or Object form, provided that You
416+ meet the following conditions:
417+
418+ (a) You must give any other recipients of the Work or
419+ Derivative Works a copy of this License; and
420+
421+ (b) You must cause any modified files to carry prominent notices
422+ stating that You changed the files; and
423diff --git a/dot_config/opencode/skills/skill-creator/SKILL.md b/dot_config/opencode/skills/skill-creator/SKILL.md
424new file mode 100644
425index 0000000000000000000000000000000000000000..b65c5a9c16dcff44d5e5151814854e68c60a2b65
426--- /dev/null
427+++ b/dot_config/opencode/skills/skill-creator/SKILL.md
428@@ -0,0 +1,503 @@
429+---
430+name: skill-creator
431diff --git a/dot_config/opencode/skills/skill-creator/agents/analyzer.md b/dot_config/opencode/skills/skill-creator/agents/analyzer.md
432new file mode 100644
433index 0000000000000000000000000000000000000000..14e41d6068635f4dd3fb878fd1626312395dda63
434--- /dev/null
435+++ b/dot_config/opencode/skills/skill-creator/agents/analyzer.md
436@@ -0,0 +1,274 @@
437+# Post-hoc Analyzer Agent
438+
439+Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
440+
441+## Role
442+
443+After the blind comparator determines a winner, the Post-hoc Analyzer "unblids" the results by examining the skills and transcripts. The goal is to extract actionable insights: what made the winner better, and how can the loser be improved?
444+
445+## Inputs
446+
447+You receive these parameters in your prompt:
448+
449+- **winner**: "A" or "B" (from blind comparison)
450+- **winner_skill_path**: Path to the skill that produced the winning output
451+- **winner_transcript_path**: Path to the execution transcript for the winner
452+- **loser_skill_path**: Path to the skill that produced the losing output
453+- **loser_transcript_path**: Path to the execution transcript for the loser
454+- **comparison_result_path**: Path to the blind comparator's output JSON
455+- **output_path**: Where to save the analysis results
456+
457+## Process
458+
459+### Step 1: Read Comparison Result
460+
461+1. Read the blind comparator's output at comparison_result_path
462+2. Note the winning side (A or B), the reasoning, and any scores
463+3. Understand what the comparator valued in the winning output
464+
465+### Step 2: Read Both Skills
466+
467+1. Read the winner skill's SKILL.md and key referenced files
468+2. Read the loser skill's SKILL.md and key referenced files
469+3. Identify structural differences:
470+ - Instructions clarity and specificity
471+ - Script/tool usage patterns
472+ - Example coverage
473+ - Edge case handling
474+
475+### Step 3: Read Both Transcripts
476+
477+1. Read the winner's transcript
478+2. Read the loser's transcript
479+3. Compare execution patterns:
480+ - How closely did each follow their skill's instructions?
481+ - What tools were used differently?
482+ - Where did the loser diverge from optimal behavior?
483+ - Did either encounter errors or make recovery attempts?
484+
485+### Step 4: Analyze Instruction Following
486+
487+For each transcript, evaluate:
488+- Did the agent follow the skill's explicit instructions?
489+- Did the agent use the skill's provided tools/scripts?
490+- Were there missed opportunities to leverage skill content?
491+- Did the agent add unnecessary steps not in the skill?
492+
493+Score instruction following 1-10 and note specific issues.
494+
495+### Step 5: Identify Winner Strengths
496+
497+Determine what made the winner better:
498+- Clearer instructions that led to better behavior?
499+- Better scripts/tools that produced better output?
500+- More comprehensive examples that guided edge cases?
501+- Better error handling guidance?
502+
503+Be specific. Quote from skills/transcripts where relevant.
504+
505+### Step 6: Identify Loser Weaknesses
506+
507+Determine what held the loser back:
508+- Ambiguous instructions that led to suboptimal choices?
509+- Missing tools/scripts that forced workarounds?
510+- Gaps in edge case coverage?
511+- Poor error handling that caused failures?
512+
513+### Step 7: Generate Improvement Suggestions
514+
515+Based on the analysis, produce actionable suggestions for improving the loser skill:
516+- Specific instruction changes to make
517+- Tools/scripts to add or modify
518+- Examples to include
519+- Edge cases to address
520+
521+Prioritize by impact. Focus on changes that would have changed the outcome.
522+
523+### Step 8: Write Analysis Results
524+
525+Save structured analysis to `{output_path}`.
526+
527+## Output Format
528+
529+Write a JSON file with this structure:
530+
531+```json
532+{
533+ "comparison_summary": {
534+ "winner": "A",
535+ "winner_skill": "path/to/winner/skill",
536diff --git a/dot_config/opencode/skills/skill-creator/agents/comparator.md b/dot_config/opencode/skills/skill-creator/agents/comparator.md
537new file mode 100644
538index 0000000000000000000000000000000000000000..80e00eb45db3ee53a132fc2ba97fd59a7339e563
539--- /dev/null
540+++ b/dot_config/opencode/skills/skill-creator/agents/comparator.md
541@@ -0,0 +1,202 @@
542+# Blind Comparator Agent
543+
544+Compare two outputs WITHOUT knowing which skill produced them.
545+
546+## Role
547+
548+The Blind Comparator judges which output better accomplishes the eval task. You receive two outputs labeled A and B, but you do NOT know which skill produced which. This prevents bias toward a particular skill or approach.
549+
550+Your judgment is based purely on output quality and task completion.
551+
552+## Inputs
553+
554+You receive these parameters in your prompt:
555+
556+- **output_a_path**: Path to the first output file or directory
557+- **output_b_path**: Path to the second output file or directory
558+- **eval_prompt**: The original task/prompt that was executed
559+- **expectations**: List of expectations to check (optional - may be empty)
560+
561+## Process
562+
563+### Step 1: Read Both Outputs
564+
565+1. Examine output A (file or directory)
566+2. Examine output B (file or directory)
567+3. Note the type, structure, and content of each
568+4. If outputs are directories, examine all relevant files inside
569+
570+### Step 2: Understand the Task
571+
572+1. Read the eval_prompt carefully
573+2. Identify what the task requires:
574+ - What should be produced?
575+ - What qualities matter (accuracy, completeness, format)?
576+ - What would distinguish a good output from a poor one?
577+
578+### Step 3: Generate Evaluation Rubric
579+
580+Based on the task, generate a rubric with two dimensions:
581+
582+**Content Rubric** (what the output contains):
583+| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
584+|-----------|----------|----------------|---------------|
585+| Correctness | Major errors | Minor errors | Fully correct |
586+| Completeness | Missing key elements | Mostly complete | All elements present |
587+| Accuracy | Significant inaccuracies | Minor inaccuracies | Accurate throughout |
588+
589+**Structure Rubric** (how the output is organized):
590+| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
591+|-----------|----------|----------------|---------------|
592+| Organization | Disorganized | Reasonably organized | Clear, logical structure |
593+| Formatting | Inconsistent/broken | Mostly consistent | Professional, polished |
594+| Usability | Difficult to use | Usable with effort | Easy to use |
595+
596+Adapt criteria to the specific task. For example:
597+- PDF form → "Field alignment", "Text readability", "Data placement"
598+- Document → "Section structure", "Heading hierarchy", "Paragraph flow"
599+- Data output → "Schema correctness", "Data types", "Completeness"
600+
601+### Step 4: Evaluate Each Output Against the Rubric
602+
603+For each output (A and B):
604+
605+1. **Score each criterion** on the rubric (1-5 scale)
606+2. **Calculate dimension totals**: Content score, Structure score
607+3. **Calculate overall score**: Average of dimension scores, scaled to 1-10
608+
609+### Step 5: Check Assertions (if provided)
610+
611+If expectations are provided:
612+
613+1. Check each expectation against output A
614+2. Check each expectation against output B
615+3. Count pass rates for each output
616+4. Use expectation scores as secondary evidence (not the primary decision factor)
617+
618+### Step 6: Determine the Winner
619+
620+Compare A and B based on (in priority order):
621+
622+1. **Primary**: Overall rubric score (content + structure)
623+2. **Secondary**: Assertion pass rates (if applicable)
624+3. **Tiebreaker**: If truly equal, declare a TIE
625+
626+Be decisive - ties should be rare. One output is usually better, even if marginally.
627+
628+### Step 7: Write Comparison Results
629+
630+Save results to a JSON file at the path specified (or `comparison.json` if not specified).
631+
632+## Output Format
633+
634+Write a JSON file with this structure:
635+
636+```json
637+{
638+ "winner": "A",
639+ "reasoning": "Output A provides a complete solution with proper formatting and all required fields. Output B is missing the date field and has formatting inconsistencies.",
640+ "rubric": {
641diff --git a/dot_config/opencode/skills/skill-creator/agents/grader.md b/dot_config/opencode/skills/skill-creator/agents/grader.md
642new file mode 100644
643index 0000000000000000000000000000000000000000..558ab05c0a9a8bb062ef4c51823d4d76c3acf7c4
644--- /dev/null
645+++ b/dot_config/opencode/skills/skill-creator/agents/grader.md
646@@ -0,0 +1,223 @@
647+# Grader Agent
648+
649+Evaluate expectations against an execution transcript and outputs.
650+
651+## Role
652+
653+The Grader reviews a transcript and output files, then determines whether each expectation passes or fails. Provide clear evidence for each judgment.
654+
655diff --git a/dot_config/opencode/skills/skill-creator/assets/eval_review.html b/dot_config/opencode/skills/skill-creator/assets/eval_review.html
656new file mode 100644
657index 0000000000000000000000000000000000000000..938ff32aed9bffabf723bd5492d720f4736c8e4d
658--- /dev/null
659+++ b/dot_config/opencode/skills/skill-creator/assets/eval_review.html
660@@ -0,0 +1,146 @@
661+<!DOCTYPE html>
662+<html lang="en">
663+<head>
664+ <meta charset="UTF-8">
665+ <meta name="viewport" content="width=device-width, initial-scale=1.0">
666+ <title>Eval Set Review - __SKILL_NAME_PLACEHOLDER__</title>
667+ <link rel="preconnect" href="https://fonts.googleapis.com">
668+ <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
669+ <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
670+ <style>
671+ * { box-sizing: border-box; margin: 0; padding: 0; }
672+ body { font-family: 'Lora', Georgia, serif; background: #faf9f5; padding: 2rem; color: #141413; }
673+ h1 { font-family: 'Poppins', sans-serif; margin-bottom: 0.5rem; font-size: 1.5rem; }
674+ .description { color: #b0aea5; margin-bottom: 1.5rem; font-style: italic; max-width: 900px; }
675+ .controls { margin-bottom: 1rem; display: flex; gap: 0.5rem; }
676+ .btn { font-family: 'Poppins', sans-serif; padding: 0.5rem 1rem; border: none; border-radius: 6px; cursor: pointer; font-size: 0.875rem; font-weight: 500; }
677+ .btn-add { background: #6a9bcc; color: white; }
678+ .btn-add:hover { background: #5889b8; }
679+ .btn-export { background: #d97757; color: white; }
680+ .btn-export:hover { background: #c4613f; }
681+ table { width: 100%; max-width: 1100px; border-collapse: collapse; background: white; border-radius: 6px; overflow: hidden; box-shadow: 0 1px 3px rgba(0,0,0,0.08); }
682+ th { font-family: 'Poppins', sans-serif; background: #141413; color: #faf9f5; padding: 0.75rem 1rem; text-align: left; font-size: 0.875rem; }
683+ td { padding: 0.75rem 1rem; border-bottom: 1px solid #e8e6dc; vertical-align: top; }
684+ tr:nth-child(even) td { background: #faf9f5; }
685+ tr:hover td { background: #f3f1ea; }
686+ .section-header td { background: #e8e6dc; font-family: 'Poppins', sans-serif; font-weight: 500; font-size: 0.8rem; color: #141413; text-transform: uppercase; letter-spacing: 0.05em; }
687+ .query-input { width: 100%; padding: 0.4rem; border: 1px solid #e8e6dc; border-radius: 4px; font-size: 0.875rem; font-family: 'Lora', Georgia, serif; resize: vertical; min-height: 60px; }
688+ .query-input:focus { outline: none; border-color: #d97757; box-shadow: 0 0 0 2px rgba(217,119,87,0.15); }
689+ .toggle { position: relative; display: inline-block; width: 44px; height: 24px; }
690+ .toggle input { opacity: 0; width: 0; height: 0; }
691+ .toggle .slider { position: absolute; inset: 0; background: #b0aea5; border-radius: 24px; cursor: pointer; transition: 0.2s; }
692+ .toggle .slider::before { content: ""; position: absolute; width: 18px; height: 18px; left: 3px; bottom: 3px; background: white; border-radius: 50%; transition: 0.2s; }
693+ .toggle input:checked + .slider { background: #d97757; }
694+ .toggle input:checked + .slider::before { transform: translateX(20px); }
695+ .btn-delete { background: #c44; color: white; padding: 0.3rem 0.6rem; border: none; border-radius: 4px; cursor: pointer; font-size: 0.75rem; font-family: 'Poppins', sans-serif; }
696+ .btn-delete:hover { background: #a33; }
697+ .summary { margin-top: 1rem; color: #b0aea5; font-size: 0.875rem; }
698+ </style>
699+</head>
700+<body>
701+ <h1>Eval Set Review: <span id="skill-name">__SKILL_NAME_PLACEHOLDER__</span></h1>
702+ <p class="description">Current description: <span id="skill-desc">__SKILL_DESCRIPTION_PLACEHOLDER__</span></p>
703+
704+ <div class="controls">
705+ <button class="btn btn-add" onclick="addRow()">+ Add Query</button>
706+ <button class="btn btn-export" onclick="exportEvalSet()">Export Eval Set</button>
707+ </div>
708+
709+ <table>
710+ <thead>
711+ <tr>
712+ <th style="width:65%">Query</th>
713+ <th style="width:18%">Should Trigger</th>
714+ <th style="width:10%">Actions</th>
715+ </tr>
716+ </thead>
717+ <tbody id="eval-body"></tbody>
718+ </table>
719+
720+ <p class="summary" id="summary"></p>
721+
722+ <script>
723+ const EVAL_DATA = __EVAL_DATA_PLACEHOLDER__;
724+
725+ let evalItems = [...EVAL_DATA];
726+
727+ function render() {
728+ const tbody = document.getElementById('eval-body');
729+ tbody.innerHTML = '';
730+
731+ // Sort: should-trigger first, then should-not-trigger
732+ const sorted = evalItems
733+ .map((item, origIdx) => ({ ...item, origIdx }))
734+ .sort((a, b) => (b.should_trigger ? 1 : 0) - (a.should_trigger ? 1 : 0));
735+
736+ let lastGroup = null;
737+ sorted.forEach(item => {
738+ const group = item.should_trigger ? 'trigger' : 'no-trigger';
739+ if (group !== lastGroup) {
740+ const headerRow = document.createElement('tr');
741+ headerRow.className = 'section-header';
742+ headerRow.innerHTML = `<td colspan="3">${item.should_trigger ? 'Should Trigger' : 'Should NOT Trigger'}</td>`;
743+ tbody.appendChild(headerRow);
744+ lastGroup = group;
745+ }
746+
747+ const idx = item.origIdx;
748+ const tr = document.createElement('tr');
749+ tr.innerHTML = `
750+ <td><textarea class="query-input" onchange="updateQuery(${idx}, this.value)">${escapeHtml(item.query)}</textarea></td>
751+ <td>
752+ <label class="toggle">
753+ <input type="checkbox" ${item.should_trigger ? 'checked' : ''} onchange="updateTrigger(${idx}, this.checked)">
754+ <span class="slider"></span>
755+ </label>
756+ <span style="margin-left:8px;font-size:0.8rem;color:#b0aea5">${item.should_trigger ? 'Yes' : 'No'}</span>
757+ </td>
758+ <td><button class="btn-delete" onclick="deleteRow(${idx})">Delete</button></td>
759+ `;
760diff --git a/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py b/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py
761new file mode 100644
762index 0000000000000000000000000000000000000000..7fa5978631fed1ed545591dbb2b0eb21ce3f3d08
763--- /dev/null
764+++ b/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py
765@@ -0,0 +1,471 @@
766+#!/usr/bin/env python3
767+"""Generate and serve a review page for eval results.
768+
769+Reads the workspace directory, discovers runs (directories with outputs/),
770+embeds all output data into a self-contained HTML page, and serves it via
771+a tiny HTTP server. Feedback auto-saves to feedback.json in the workspace.
772+
773+Usage:
774+ python generate_review.py <workspace-path> [--port PORT] [--skill-name NAME]
775+ python generate_review.py <workspace-path> --previous-feedback /path/to/old/feedback.json
776+
777+No dependencies beyond the Python stdlib are required.
778+"""
779+
780+import argparse
781+import base64
782+import json
783+import mimetypes
784+import os
785+import re
786+import signal
787+import subprocess
788+import sys
789+import time
790+import webbrowser
791+from functools import partial
792+from http.server import HTTPServer, BaseHTTPRequestHandler
793+from pathlib import Path
794+
795+# Files to exclude from output listings
796+METADATA_FILES = {"transcript.md", "user_notes.md", "metrics.json"}
797+
798+# Extensions we render as inline text
799+TEXT_EXTENSIONS = {
800+ ".txt", ".md", ".json", ".csv", ".py", ".js", ".ts", ".tsx", ".jsx",
801+ ".yaml", ".yml", ".xml", ".html", ".css", ".sh", ".rb", ".go", ".rs",
802+ ".java", ".c", ".cpp", ".h", ".hpp", ".sql", ".r", ".toml",
803+}
804+
805+# Extensions we render as inline images
806+IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp"}
807+
808+# MIME type overrides for common types
809+MIME_OVERRIDES = {
810+ ".svg": "image/svg+xml",
811+ ".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
812+ ".docx": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
813+ ".pptx": "application/vnd.openxmlformats-officedocument.presentationml.presentation",
814+}
815+
816+
817+def get_mime_type(path: Path) -> str:
818+ ext = path.suffix.lower()
819+ if ext in MIME_OVERRIDES:
820+ return MIME_OVERRIDES[ext]
821+ mime, _ = mimetypes.guess_type(str(path))
822+ return mime or "application/octet-stream"
823+
824+
825+def find_runs(workspace: Path) -> list[dict]:
826+ """Recursively find directories that contain an outputs/ subdirectory."""
827+ runs: list[dict] = []
828+ _find_runs_recursive(workspace, workspace, runs)
829+ runs.sort(key=lambda r: (r.get("eval_id", float("inf")), r["id"]))
830+ return runs
831+
832+
833+def _find_runs_recursive(root: Path, current: Path, runs: list[dict]) -> None:
834+ if not current.is_dir():
835+ return
836+
837+ outputs_dir = current / "outputs"
838+ if outputs_dir.is_dir():
839+ run = build_run(root, current)
840+ if run:
841+ runs.append(run)
842+ return
843+
844+ skip = {"node_modules", ".git", "__pycache__", "skill", "inputs"}
845+ for child in sorted(current.iterdir()):
846+ if child.is_dir() and child.name not in skip:
847+ _find_runs_recursive(root, child, runs)
848+
849+
850+def build_run(root: Path, run_dir: Path) -> dict | None:
851+ """Build a run dict with prompt, outputs, and grading data."""
852+ prompt = ""
853+ eval_id = None
854+
855+ # Try eval_metadata.json
856+ for candidate in [run_dir / "eval_metadata.json", run_dir.parent / "eval_metadata.json"]:
857+ if candidate.exists():
858+ try:
859+ metadata = json.loads(candidate.read_text())
860+ prompt = metadata.get("prompt", "")
861+ eval_id = metadata.get("eval_id")
862+ except (json.JSONDecodeError, OSError):
863+ pass
864+ if prompt:
865diff --git a/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html b/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html
866new file mode 100644
867index 0000000000000000000000000000000000000000..6d8e96348a02e66c3363d2ff3b3ae58ac11e6382
868--- /dev/null
869+++ b/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html
870@@ -0,0 +1,1325 @@
871+<!DOCTYPE html>
872+<html lang="en">
873+<head>
874+ <meta charset="UTF-8">
875+ <meta name="viewport" content="width=device-width, initial-scale=1.0">
876+ <title>Eval Review</title>
877+ <link rel="preconnect" href="https://fonts.googleapis.com">
878+ <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
879+ <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
880+ <script src="https://cdn.sheetjs.com/xlsx-0.20.3/package/dist/xlsx.full.min.js" integrity="sha384-EnyY0/GSHQGSxSgMwaIPzSESbqoOLSexfnSMN2AP+39Ckmn92stwABZynq1JyzdT" crossorigin="anonymous"></script>
881+ <style>
882+ :root {
883+ --bg: #faf9f5;
884+ --surface: #ffffff;
885+ --border: #e8e6dc;
886+ --text: #141413;
887+ --text-muted: #b0aea5;
888+ --accent: #d97757;
889+ --accent-hover: #c4613f;
890+ --green: #788c5d;
891+ --green-bg: #eef2e8;
892+ --red: #c44;
893+ --red-bg: #fceaea;
894+ --header-bg: #141413;
895+ --header-text: #faf9f5;
896+ --radius: 6px;
897+ }
898+
899+ * { box-sizing: border-box; margin: 0; padding: 0; }
900+
901+ body {
902+ font-family: 'Lora', Georgia, serif;
903+ background: var(--bg);
904+ color: var(--text);
905+ height: 100vh;
906+ display: flex;
907+ flex-direction: column;
908+ }
909+
910+ /* ---- Header ---- */
911+ .header {
912+ background: var(--header-bg);
913+ color: var(--header-text);
914+ padding: 1rem 2rem;
915+ display: flex;
916+ justify-content: space-between;
917+ align-items: center;
918+ flex-shrink: 0;
919+ }
920+ .header h1 {
921+ font-family: 'Poppins', sans-serif;
922+ font-size: 1.25rem;
923+ font-weight: 600;
924+ }
925+ .header .instructions {
926+ font-size: 0.8rem;
927+ opacity: 0.7;
928+ margin-top: 0.25rem;
929+ }
930+ .header .progress {
931+ font-size: 0.875rem;
932+ opacity: 0.8;
933+ text-align: right;
934+ }
935+
936+ /* ---- Main content ---- */
937+ .main {
938+ flex: 1;
939+ overflow-y: auto;
940+ padding: 1.5rem 2rem;
941+ display: flex;
942+ flex-direction: column;
943+ gap: 1.25rem;
944+ }
945+
946+ /* ---- Sections ---- */
947+ .section {
948+ background: var(--surface);
949+ border: 1px solid var(--border);
950+ border-radius: var(--radius);
951+ flex-shrink: 0;
952+ }
953+ .section-header {
954+ font-family: 'Poppins', sans-serif;
955+ padding: 0.75rem 1rem;
956+ font-size: 0.75rem;
957+ font-weight: 500;
958+ text-transform: uppercase;
959+ letter-spacing: 0.05em;
960+ color: var(--text-muted);
961+ border-bottom: 1px solid var(--border);
962+ background: var(--bg);
963+ }
964+ .section-body {
965+ padding: 1rem;
966+ }
967+
968+ /* ---- Config badge ---- */
969+ .config-badge {
970diff --git a/dot_config/opencode/skills/skill-creator/references/schemas.md b/dot_config/opencode/skills/skill-creator/references/schemas.md
971new file mode 100644
972index 0000000000000000000000000000000000000000..b6eeaa2d4a34c1653069585c6c5603da39a5bdbe
973--- /dev/null
974+++ b/dot_config/opencode/skills/skill-creator/references/schemas.md
975@@ -0,0 +1,430 @@
976+# JSON Schemas
977+
978+This document defines the JSON schemas used by skill-creator.
979+
980+---
981+
982+## evals.json
983+
984+Defines the evals for a skill. Located at `evals/evals.json` within the skill directory.
985+
986+```json
987+{
988+ "skill_name": "example-skill",
989+ "evals": [
990+ {
991+ "id": 1,
992+ "prompt": "User's example prompt",
993+ "expected_output": "Description of expected result",
994+ "files": ["evals/files/sample1.pdf"],
995+ "expectations": [
996+ "The output includes X",
997+ "The skill used script Y"
998+ ]
999+ }
1000+ ]
1001+}
1002+```
1003+
1004+**Fields:**
1005+- `skill_name`: Name matching the skill's frontmatter
1006+- `evals[].id`: Unique integer identifier
1007+- `evals[].prompt`: The task to execute
1008+- `evals[].expected_output`: Human-readable description of success
1009+- `evals[].files`: Optional list of input file paths (relative to skill root)
1010+- `evals[].expectations`: List of verifiable statements
1011+
1012+---
1013+
1014+## history.json
1015+
1016+Tracks version progression in Improve mode. Located at workspace root.
1017+
1018+```json
1019+{
1020+ "started_at": "2026-01-15T10:30:00Z",
1021+ "skill_name": "pdf",
1022+ "current_best": "v2",
1023+ "iterations": [
1024+ {
1025+ "version": "v0",
1026+ "parent": null,
1027+ "expectation_pass_rate": 0.65,
1028+ "grading_result": "baseline",
1029+ "is_current_best": false
1030+ },
1031+ {
1032+ "version": "v1",
1033+ "parent": "v0",
1034+ "expectation_pass_rate": 0.75,
1035+ "grading_result": "won",
1036+ "is_current_best": false
1037+ },
1038+ {
1039+ "version": "v2",
1040+ "parent": "v1",
1041+ "expectation_pass_rate": 0.85,
1042+ "grading_result": "won",
1043+ "is_current_best": true
1044+ }
1045+ ]
1046+}
1047+```
1048+
1049+**Fields:**
1050+- `started_at`: ISO timestamp of when improvement started
1051+- `skill_name`: Name of the skill being improved
1052+- `current_best`: Version identifier of the best performer
1053+- `iterations[].version`: Version identifier (v0, v1, ...)
1054+- `iterations[].parent`: Parent version this was derived from
1055+- `iterations[].expectation_pass_rate`: Pass rate from grading
1056+- `iterations[].grading_result`: "baseline", "won", "lost", or "tie"
1057+- `iterations[].is_current_best`: Whether this is the current best version
1058+
1059+---
1060+
1061+## grading.json
1062+
1063+Output from the grader agent. Located at `<run-dir>/grading.json`.
1064+
1065+```json
1066+{
1067+ "expectations": [
1068+ {
1069+ "text": "The output includes the name 'John Smith'",
1070+ "passed": true,
1071+ "evidence": "Found in transcript Step 3: 'Extracted names: John Smith, Sarah Johnson'"
1072+ },
1073+ {
1074+ "text": "The spreadsheet has a SUM formula in cell B10",
1075diff --git a/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py b/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py
1076new file mode 100644
1077index 0000000000000000000000000000000000000000..e69de29bb2d1d6434b8b29ae775ad8c2e48c5391
1078--- /dev/null
1079+++ b/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py
1080diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py b/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py
1081new file mode 100644
1082index 0000000000000000000000000000000000000000..3e66e8c105be9bab9f0e9c61f0d1482619401580
1083--- /dev/null
1084+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py
1085@@ -0,0 +1,401 @@
1086+#!/usr/bin/env python3
1087+"""
1088+Aggregate individual run results into benchmark summary statistics.
1089+
1090+Reads grading.json files from run directories and produces:
1091+- run_summary with mean, stddev, min, max for each metric
1092+- delta between with_skill and without_skill configurations
1093+
1094+Usage:
1095+ python aggregate_benchmark.py <benchmark_dir>
1096+
1097+Example:
1098+ python aggregate_benchmark.py benchmarks/2026-01-15T10-30-00/
1099+
1100+The script supports two directory layouts:
1101+
1102+ Workspace layout (from skill-creator iterations):
1103+ <benchmark_dir>/
1104+ └── eval-N/
1105+ ├── with_skill/
1106+ │ ├── run-1/grading.json
1107+ │ └── run-2/grading.json
1108+ └── without_skill/
1109+ ├── run-1/grading.json
1110+ └── run-2/grading.json
1111+
1112+ Legacy layout (with runs/ subdirectory):
1113+ <benchmark_dir>/
1114+ └── runs/
1115+ └── eval-N/
1116+ ├── with_skill/
1117+ │ └── run-1/grading.json
1118+ └── without_skill/
1119+ └── run-1/grading.json
1120+"""
1121+
1122+import argparse
1123+import json
1124+import math
1125+import sys
1126+from datetime import datetime, timezone
1127+from pathlib import Path
1128+
1129+
1130+def calculate_stats(values: list[float]) -> dict:
1131+ """Calculate mean, stddev, min, max for a list of values."""
1132+ if not values:
1133+ return {"mean": 0.0, "stddev": 0.0, "min": 0.0, "max": 0.0}
1134+
1135+ n = len(values)
1136+ mean = sum(values) / n
1137+
1138+ if n > 1:
1139+ variance = sum((x - mean) ** 2 for x in values) / (n - 1)
1140+ stddev = math.sqrt(variance)
1141+ else:
1142+ stddev = 0.0
1143+
1144+ return {
1145+ "mean": round(mean, 4),
1146+ "stddev": round(stddev, 4),
1147+ "min": round(min(values), 4),
1148+ "max": round(max(values), 4)
1149+ }
1150+
1151+
1152+def load_run_results(benchmark_dir: Path) -> dict:
1153+ """
1154+ Load all run results from a benchmark directory.
1155+
1156+ Returns dict keyed by config name (e.g. "with_skill"/"without_skill",
1157+ or "new_skill"/"old_skill"), each containing a list of run results.
1158+ """
1159+ # Support both layouts: eval dirs directly under benchmark_dir, or under runs/
1160+ runs_dir = benchmark_dir / "runs"
1161+ if runs_dir.exists():
1162+ search_dir = runs_dir
1163+ elif list(benchmark_dir.glob("eval-*")):
1164+ search_dir = benchmark_dir
1165+ else:
1166+ print(f"No eval directories found in {benchmark_dir} or {benchmark_dir / 'runs'}")
1167+ return {}
1168+
1169+ results: dict[str, list] = {}
1170+
1171+ for eval_idx, eval_dir in enumerate(sorted(search_dir.glob("eval-*"))):
1172+ metadata_path = eval_dir / "eval_metadata.json"
1173+ if metadata_path.exists():
1174+ try:
1175+ with open(metadata_path) as mf:
1176+ eval_id = json.load(mf).get("eval_id", eval_idx)
1177+ except (json.JSONDecodeError, OSError):
1178+ eval_id = eval_idx
1179+ else:
1180+ try:
1181+ eval_id = int(eval_dir.name.split("-")[1])
1182+ except ValueError:
1183+ eval_id = eval_idx
1184+
1185diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py b/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py
1186new file mode 100644
1187index 0000000000000000000000000000000000000000..959e30a0014ec165c41a2bb7420b7dfe1416bbac
1188--- /dev/null
1189+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py
1190@@ -0,0 +1,326 @@
1191+#!/usr/bin/env python3
1192+"""Generate an HTML report from run_loop.py output.
1193+
1194+Takes the JSON output from run_loop.py and generates a visual HTML report
1195+showing each description attempt with check/x for each test case.
1196+Distinguishes between train and test queries.
1197+"""
1198+
1199+import argparse
1200+import html
1201+import json
1202+import sys
1203+from pathlib import Path
1204+
1205+
1206+def generate_html(data: dict, auto_refresh: bool = False, skill_name: str = "") -> str:
1207+ """Generate HTML report from loop output data. If auto_refresh is True, adds a meta refresh tag."""
1208+ history = data.get("history", [])
1209+ holdout = data.get("holdout", 0)
1210+ title_prefix = html.escape(skill_name + " \u2014 ") if skill_name else ""
1211+
1212+ # Get all unique queries from train and test sets, with should_trigger info
1213+ train_queries: list[dict] = []
1214+ test_queries: list[dict] = []
1215+ if history:
1216+ for r in history[0].get("train_results", history[0].get("results", [])):
1217+ train_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
1218+ if history[0].get("test_results"):
1219+ for r in history[0].get("test_results", []):
1220+ test_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
1221+
1222+ refresh_tag = ' <meta http-equiv="refresh" content="5">\n' if auto_refresh else ""
1223+
1224+ html_parts = ["""<!DOCTYPE html>
1225+<html>
1226+<head>
1227+ <meta charset="utf-8">
1228+""" + refresh_tag + """ <title>""" + title_prefix + """Skill Description Optimization</title>
1229+ <link rel="preconnect" href="https://fonts.googleapis.com">
1230+ <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
1231+ <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
1232+ <style>
1233+ body {
1234+ font-family: 'Lora', Georgia, serif;
1235+ max-width: 100%;
1236+ margin: 0 auto;
1237+ padding: 20px;
1238+ background: #faf9f5;
1239+ color: #141413;
1240+ }
1241+ h1 { font-family: 'Poppins', sans-serif; color: #141413; }
1242+ .explainer {
1243+ background: white;
1244+ padding: 15px;
1245+ border-radius: 6px;
1246+ margin-bottom: 20px;
1247+ border: 1px solid #e8e6dc;
1248+ color: #b0aea5;
1249+ font-size: 0.875rem;
1250+ line-height: 1.6;
1251+ }
1252+ .summary {
1253+ background: white;
1254+ padding: 15px;
1255+ border-radius: 6px;
1256+ margin-bottom: 20px;
1257+ border: 1px solid #e8e6dc;
1258+ }
1259+ .summary p { margin: 5px 0; }
1260+ .best { color: #788c5d; font-weight: bold; }
1261+ .table-container {
1262+ overflow-x: auto;
1263+ width: 100%;
1264+ }
1265+ table {
1266+ border-collapse: collapse;
1267+ background: white;
1268+ border: 1px solid #e8e6dc;
1269+ border-radius: 6px;
1270+ font-size: 12px;
1271+ min-width: 100%;
1272+ }
1273+ th, td {
1274+ padding: 8px;
1275+ text-align: left;
1276+ border: 1px solid #e8e6dc;
1277+ white-space: normal;
1278+ word-wrap: break-word;
1279+ }
1280+ th {
1281+ font-family: 'Poppins', sans-serif;
1282+ background: #141413;
1283+ color: #faf9f5;
1284+ font-weight: 500;
1285+ }
1286+ th.test-col {
1287+ background: #6a9bcc;
1288+ }
1289+ th.query-col { min-width: 200px; }
1290diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py b/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py
1291new file mode 100644
1292index 0000000000000000000000000000000000000000..12bbdb8635073a7f84f30741fc9795ef7952f9c5
1293--- /dev/null
1294+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py
1295@@ -0,0 +1,261 @@
1296+#!/usr/bin/env python3
1297+"""Improve a skill description based on eval results.
1298+
1299+Takes eval results (from run_eval.py) and generates an improved description
1300+using opencode with haiku 4.5.
1301+"""
1302+
1303+import argparse
1304+import json
1305+import re
1306+import subprocess
1307+import sys
1308+from pathlib import Path
1309+
1310+from scripts.utils import parse_skill_md
1311+
1312+IMPROVE_MODEL = "anthropic/claude-haiku-4-5"
1313+
1314+
1315+def _run_opencode(prompt: str, model: str) -> str:
1316+ """Run opencode run with a prompt and return the assistant text response."""
1317+ env_clean = {
1318+ k: v
1319+ for k, v in __import__("os").environ.items()
1320+ if k not in ("OPENCODE", "OPENCODE_PID")
1321+ }
1322+ result = subprocess.run(
1323+ ["opencode", "run", "--format", "json", "--model", model, prompt],
1324+ capture_output=True,
1325+ text=True,
1326+ env=env_clean,
1327+ )
1328+ # Parse JSON event stream, collect text parts
1329+ text_parts = []
1330+ for line in result.stdout.splitlines():
1331+ line = line.strip()
1332+ if not line:
1333+ continue
1334+ try:
1335+ event = json.loads(line)
1336+ except json.JSONDecodeError:
1337+ continue
1338+ if event.get("type") == "text":
1339+ part = event.get("part", {})
1340+ text_parts.append(part.get("text", ""))
1341+ return "".join(text_parts)
1342+
1343+
1344+def improve_description(
1345+ skill_name: str,
1346+ skill_content: str,
1347+ current_description: str,
1348+ eval_results: dict,
1349+ history: list[dict],
1350+ model: str = IMPROVE_MODEL,
1351+ test_results: dict | None = None,
1352+ log_dir: Path | None = None,
1353+ iteration: int | None = None,
1354+) -> str:
1355+ """Call opencode to improve the description based on eval results."""
1356+ failed_triggers = [
1357+ r for r in eval_results["results"] if r["should_trigger"] and not r["pass"]
1358+ ]
1359+ false_triggers = [
1360+ r for r in eval_results["results"] if not r["should_trigger"] and not r["pass"]
1361+ ]
1362+
1363+ # Build scores summary
1364+ train_score = (
1365+ f"{eval_results['summary']['passed']}/{eval_results['summary']['total']}"
1366+ )
1367+ if test_results:
1368+ test_score = (
1369+ f"{test_results['summary']['passed']}/{test_results['summary']['total']}"
1370+ )
1371+ scores_summary = f"Train: {train_score}, Test: {test_score}"
1372+ else:
1373+ scores_summary = f"Train: {train_score}"
1374+
1375diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py
1376new file mode 100644
1377index 0000000000000000000000000000000000000000..4a6fe2624c5c2c3e723e9ad58005fd9e361e9499
1378--- /dev/null
1379+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py
1380@@ -0,0 +1,304 @@
1381+#!/usr/bin/env python3
1382+"""Run trigger evaluation for a skill description.
1383+
1384+Tests whether a skill's description causes opencode to trigger (read the skill)
1385+for a set of queries. Outputs results as JSON.
1386+"""
1387+
1388+import argparse
1389+import json
1390+import os
1391+import select
1392+import subprocess
1393+import sys
1394+import time
1395+import uuid
1396+from concurrent.futures import ProcessPoolExecutor, as_completed
1397+from pathlib import Path
1398+
1399+from scripts.utils import parse_skill_md
1400+
1401+EVAL_MODEL = "anthropic/claude-sonnet-4-6"
1402+
1403+
1404+def find_project_root() -> Path:
1405+ """Find the project root by walking up from cwd looking for .opencode/.
1406+
1407+ Mimics how opencode discovers its project root, so the command file
1408+ we create ends up where opencode will look for it.
1409+ """
1410+ current = Path.cwd()
1411+ for parent in [current, *current.parents]:
1412+ if (parent / ".opencode").is_dir():
1413+ return parent
1414+ return current
1415+
1416+
1417+def run_single_query(
1418+ query: str,
1419+ skill_name: str,
1420+ skill_description: str,
1421+ timeout: int,
1422+ project_root: str,
1423+ model: str | None = None,
1424+) -> bool:
1425+ """Run a single query and return whether the skill was triggered.
1426+
1427+ Creates a skill file in .opencode/skills/ so it appears in opencode's
1428+ available_skills list, then runs `opencode run` with the raw query.
1429+ Parses JSON events to detect whether the skill tool was invoked.
1430+ """
1431+ unique_id = uuid.uuid4().hex[:8]
1432+ clean_name = f"{skill_name}-skill-{unique_id}"
1433+ project_skills_dir = Path(project_root) / ".opencode" / "skills" / clean_name
1434+ skill_file = project_skills_dir / "SKILL.md"
1435+
1436+ try:
1437+ project_skills_dir.mkdir(parents=True, exist_ok=True)
1438+ # Use YAML block scalar to avoid breaking on quotes in description
1439+ indented_desc = "\n ".join(skill_description.split("\n"))
1440+ skill_content = (
1441+ f"---\n"
1442+ f"name: {clean_name}\n"
1443+ f"description: |\n"
1444+ f" {indented_desc}\n"
1445+ f"---\n\n"
1446+ f"# {skill_name}\n\n"
1447+ f"This skill handles: {skill_description}\n"
1448+ )
1449+ skill_file.write_text(skill_content)
1450+
1451+ cmd = [
1452+ "opencode",
1453+ "run",
1454+ "--format",
1455+ "json",
1456+ "--model",
1457+ model or EVAL_MODEL,
1458+ query,
1459+ ]
1460+
1461+ # Remove OPENCODE env var to allow nesting opencode run inside an
1462+ # opencode session. The guard is for interactive terminal conflicts;
1463+ # programmatic subprocess usage is safe.
1464+ env = {
1465+ k: v for k, v in os.environ.items() if k not in ("OPENCODE", "OPENCODE_PID")
1466+ }
1467+
1468+ process = subprocess.Popen(
1469+ cmd,
1470+ stdout=subprocess.PIPE,
1471+ stderr=subprocess.DEVNULL,
1472+ cwd=project_root,
1473+ env=env,
1474+ )
1475+
1476+ triggered = False
1477+ start_time = time.time()
1478+ buffer = ""
1479+
1480diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py
1481new file mode 100644
1482index 0000000000000000000000000000000000000000..a18ba9b7544d573371420e31f0070452c324d75e
1483--- /dev/null
1484+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py
1485@@ -0,0 +1,404 @@
1486+#!/usr/bin/env python3
1487+"""Run the eval + improve loop until all pass or max iterations reached.
1488+
1489+Combines run_eval.py and improve_description.py in a loop, tracking history
1490+and returning the best description found. Supports train/test split to prevent
1491+overfitting.
1492+"""
1493+
1494+import argparse
1495+import json
1496+import random
1497+import sys
1498+import tempfile
1499+import time
1500+import webbrowser
1501+from pathlib import Path
1502+
1503+from scripts.generate_report import generate_html
1504+from scripts.improve_description import improve_description, IMPROVE_MODEL
1505+from scripts.run_eval import find_project_root, run_eval, EVAL_MODEL
1506+from scripts.utils import parse_skill_md
1507+
1508+
1509+def split_eval_set(
1510+ eval_set: list[dict], holdout: float, seed: int = 42
1511+) -> tuple[list[dict], list[dict]]:
1512+ """Split eval set into train and test sets, stratified by should_trigger."""
1513+ random.seed(seed)
1514+
1515+ # Separate by should_trigger
1516+ trigger = [e for e in eval_set if e["should_trigger"]]
1517+ no_trigger = [e for e in eval_set if not e["should_trigger"]]
1518+
1519+ # Shuffle each group
1520+ random.shuffle(trigger)
1521+ random.shuffle(no_trigger)
1522+
1523+ # Calculate split points
1524+ n_trigger_test = max(1, int(len(trigger) * holdout))
1525+ n_no_trigger_test = max(1, int(len(no_trigger) * holdout))
1526+
1527+ # Split
1528+ test_set = trigger[:n_trigger_test] + no_trigger[:n_no_trigger_test]
1529+ train_set = trigger[n_trigger_test:] + no_trigger[n_no_trigger_test:]
1530+
1531+ return train_set, test_set
1532+
1533+
1534+def run_loop(
1535+ eval_set: list[dict],
1536+ skill_path: Path,
1537+ description_override: str | None,
1538+ num_workers: int,
1539+ timeout: int,
1540+ max_iterations: int,
1541+ runs_per_query: int,
1542+ trigger_threshold: float,
1543+ holdout: float,
1544+ model: str,
1545+ verbose: bool,
1546+ live_report_path: Path | None = None,
1547+ log_dir: Path | None = None,
1548+) -> dict:
1549+ """Run the eval + improvement loop."""
1550+ project_root = find_project_root()
1551+ name, original_description, content = parse_skill_md(skill_path)
1552+ current_description = description_override or original_description
1553+
1554+ # Split into train/test if holdout > 0
1555+ if holdout > 0:
1556+ train_set, test_set = split_eval_set(eval_set, holdout)
1557+ if verbose:
1558+ print(
1559+ f"Split: {len(train_set)} train, {len(test_set)} test (holdout={holdout})",
1560+ file=sys.stderr,
1561+ )
1562+ else:
1563+ train_set = eval_set
1564+ test_set = []
1565+
1566+ history = []
1567+ exit_reason = "unknown"
1568+
1569+ for iteration in range(1, max_iterations + 1):
1570+ if verbose:
1571+ print(f"\n{'=' * 60}", file=sys.stderr)
1572+ print(f"Iteration {iteration}/{max_iterations}", file=sys.stderr)
1573+ print(f"Description: {current_description}", file=sys.stderr)
1574+ print(f"{'=' * 60}", file=sys.stderr)
1575+
1576+ # Evaluate train + test together in one batch for parallelism
1577+ all_queries = train_set + test_set
1578+ t0 = time.time()
1579+ all_results = run_eval(
1580+ eval_set=all_queries,
1581+ skill_name=name,
1582+ description=current_description,
1583+ num_workers=num_workers,
1584+ timeout=timeout,
1585diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py b/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py
1586new file mode 100644
1587index 0000000000000000000000000000000000000000..f48eac444656ddc41204aac1760a217951ce609e
1588--- /dev/null
1589+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py
1590@@ -0,0 +1,136 @@
1591+#!/usr/bin/env python3
1592+"""
1593+Skill Packager - Creates a distributable .skill file of a skill folder
1594+
1595+Usage:
1596+ python utils/package_skill.py <path/to/skill-folder> [output-directory]
1597+
1598+Example:
1599+ python utils/package_skill.py skills/public/my-skill
1600+ python utils/package_skill.py skills/public/my-skill ./dist
1601+"""
1602+
1603+import fnmatch
1604+import sys
1605+import zipfile
1606+from pathlib import Path
1607+from scripts.quick_validate import validate_skill
1608+
1609+# Patterns to exclude when packaging skills.
1610+EXCLUDE_DIRS = {"__pycache__", "node_modules"}
1611+EXCLUDE_GLOBS = {"*.pyc"}
1612+EXCLUDE_FILES = {".DS_Store"}
1613+# Directories excluded only at the skill root (not when nested deeper).
1614+ROOT_EXCLUDE_DIRS = {"evals"}
1615+
1616+
1617+def should_exclude(rel_path: Path) -> bool:
1618+ """Check if a path should be excluded from packaging."""
1619+ parts = rel_path.parts
1620+ if any(part in EXCLUDE_DIRS for part in parts):
1621+ return True
1622+ # rel_path is relative to skill_path.parent, so parts[0] is the skill
1623+ # folder name and parts[1] (if present) is the first subdir.
1624+ if len(parts) > 1 and parts[1] in ROOT_EXCLUDE_DIRS:
1625+ return True
1626+ name = rel_path.name
1627+ if name in EXCLUDE_FILES:
1628+ return True
1629+ return any(fnmatch.fnmatch(name, pat) for pat in EXCLUDE_GLOBS)
1630+
1631+
1632+def package_skill(skill_path, output_dir=None):
1633+ """
1634+ Package a skill folder into a .skill file.
1635+
1636+ Args:
1637+ skill_path: Path to the skill folder
1638+ output_dir: Optional output directory for the .skill file (defaults to current directory)
1639+
1640+ Returns:
1641+ Path to the created .skill file, or None if error
1642+ """
1643+ skill_path = Path(skill_path).resolve()
1644+
1645+ # Validate skill folder exists
1646+ if not skill_path.exists():
1647+ print(f"❌ Error: Skill folder not found: {skill_path}")
1648+ return None
1649+
1650+ if not skill_path.is_dir():
1651+ print(f"❌ Error: Path is not a directory: {skill_path}")
1652+ return None
1653+
1654+ # Validate SKILL.md exists
1655+ skill_md = skill_path / "SKILL.md"
1656+ if not skill_md.exists():
1657+ print(f"❌ Error: SKILL.md not found in {skill_path}")
1658+ return None
1659+
1660+ # Run validation before packaging
1661+ print("🔍 Validating skill...")
1662+ valid, message = validate_skill(skill_path)
1663+ if not valid:
1664+ print(f"❌ Validation failed: {message}")
1665+ print(" Please fix the validation errors before packaging.")
1666+ return None
1667+ print(f"✅ {message}\n")
1668+
1669+ # Determine output location
1670+ skill_name = skill_path.name
1671+ if output_dir:
1672+ output_path = Path(output_dir).resolve()
1673+ output_path.mkdir(parents=True, exist_ok=True)
1674+ else:
1675+ output_path = Path.cwd()
1676+
1677+ skill_filename = output_path / f"{skill_name}.skill"
1678+
1679+ # Create the .skill file (zip format)
1680+ try:
1681+ with zipfile.ZipFile(skill_filename, 'w', zipfile.ZIP_DEFLATED) as zipf:
1682+ # Walk through the skill directory, excluding build artifacts
1683+ for file_path in skill_path.rglob('*'):
1684+ if not file_path.is_file():
1685+ continue
1686+ arcname = file_path.relative_to(skill_path.parent)
1687+ if should_exclude(arcname):
1688+ print(f" Skipped: {arcname}")
1689+ continue
1690diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py b/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py
1691new file mode 100644
1692index 0000000000000000000000000000000000000000..ed8e1dddce77b16af13c6f36b3fe86c4ac7c590c
1693--- /dev/null
1694+++ b/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py
1695@@ -0,0 +1,103 @@
1696+#!/usr/bin/env python3
1697+"""
1698+Quick validation script for skills - minimal version
1699+"""
1700+
1701+import sys
1702+import os
1703+import re
1704+import yaml
1705+from pathlib import Path
1706+
1707+def validate_skill(skill_path):
1708+ """Basic validation of a skill"""
1709+ skill_path = Path(skill_path)
1710+
1711+ # Check SKILL.md exists
1712+ skill_md = skill_path / 'SKILL.md'
1713+ if not skill_md.exists():
1714+ return False, "SKILL.md not found"
1715+
1716+ # Read and validate frontmatter
1717+ content = skill_md.read_text()
1718+ if not content.startswith('---'):
1719+ return False, "No YAML frontmatter found"
1720+
1721+ # Extract frontmatter
1722+ match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
1723+ if not match:
1724+ return False, "Invalid frontmatter format"
1725+
1726+ frontmatter_text = match.group(1)
1727+
1728+ # Parse YAML frontmatter
1729+ try:
1730+ frontmatter = yaml.safe_load(frontmatter_text)
1731+ if not isinstance(frontmatter, dict):
1732+ return False, "Frontmatter must be a YAML dictionary"
1733+ except yaml.YAMLError as e:
1734+ return False, f"Invalid YAML in frontmatter: {e}"
1735+
1736+ # Define allowed properties
1737+ ALLOWED_PROPERTIES = {'name', 'description', 'license', 'allowed-tools', 'metadata', 'compatibility'}
1738+
1739+ # Check for unexpected properties (excluding nested keys under metadata)
1740+ unexpected_keys = set(frontmatter.keys()) - ALLOWED_PROPERTIES
1741+ if unexpected_keys:
1742+ return False, (
1743+ f"Unexpected key(s) in SKILL.md frontmatter: {', '.join(sorted(unexpected_keys))}. "
1744+ f"Allowed properties are: {', '.join(sorted(ALLOWED_PROPERTIES))}"
1745+ )
1746+
1747+ # Check required fields
1748+ if 'name' not in frontmatter:
1749+ return False, "Missing 'name' in frontmatter"
1750+ if 'description' not in frontmatter:
1751+ return False, "Missing 'description' in frontmatter"
1752+
1753+ # Extract name for validation
1754+ name = frontmatter.get('name', '')
1755+ if not isinstance(name, str):
1756+ return False, f"Name must be a string, got {type(name).__name__}"
1757+ name = name.strip()
1758+ if name:
1759+ # Check naming convention (kebab-case: lowercase with hyphens)
1760+ if not re.match(r'^[a-z0-9-]+$', name):
1761+ return False, f"Name '{name}' should be kebab-case (lowercase letters, digits, and hyphens only)"
1762+ if name.startswith('-') or name.endswith('-') or '--' in name:
1763+ return False, f"Name '{name}' cannot start/end with hyphen or contain consecutive hyphens"
1764+ # Check name length (max 64 characters per spec)
1765+ if len(name) > 64:
1766+ return False, f"Name is too long ({len(name)} characters). Maximum is 64 characters."
1767+
1768+ # Extract and validate description
1769+ description = frontmatter.get('description', '')
1770+ if not isinstance(description, str):
1771+ return False, f"Description must be a string, got {type(description).__name__}"
1772+ description = description.strip()
1773+ if description:
1774+ # Check for angle brackets
1775+ if '<' in description or '>' in description:
1776+ return False, "Description cannot contain angle brackets (< or >)"
1777+ # Check description length (max 1024 characters per spec)
1778+ if len(description) > 1024:
1779+ return False, f"Description is too long ({len(description)} characters). Maximum is 1024 characters."
1780+
1781+ # Validate compatibility field if present (optional)
1782+ compatibility = frontmatter.get('compatibility', '')
1783+ if compatibility:
1784+ if not isinstance(compatibility, str):
1785+ return False, f"Compatibility must be a string, got {type(compatibility).__name__}"
1786+ if len(compatibility) > 500:
1787+ return False, f"Compatibility is too long ({len(compatibility)} characters). Maximum is 500 characters."
1788+
1789+ return True, "Skill is valid!"
1790+
1791+if __name__ == "__main__":
1792+ if len(sys.argv) != 2:
1793+ print("Usage: python quick_validate.py <skill_directory>")
1794+ sys.exit(1)
1795diff --git a/dot_config/opencode/skills/skill-creator/scripts/utils.py b/dot_config/opencode/skills/skill-creator/scripts/utils.py
1796new file mode 100644
1797index 0000000000000000000000000000000000000000..51b6a07dd57174197a937034b7eecebd5768ff8a
1798--- /dev/null
1799+++ b/dot_config/opencode/skills/skill-creator/scripts/utils.py
1800@@ -0,0 +1,47 @@
1801+"""Shared utilities for skill-creator scripts."""
1802+
1803+from pathlib import Path
1804+
1805+
1806+
1807+def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
1808+ """Parse a SKILL.md file, returning (name, description, full_content)."""
1809+ content = (skill_path / "SKILL.md").read_text()
1810+ lines = content.split("\n")
1811+
1812+ if lines[0].strip() != "---":
1813+ raise ValueError("SKILL.md missing frontmatter (no opening ---)")
1814+
1815+ end_idx = None
1816+ for i, line in enumerate(lines[1:], start=1):
1817+ if line.strip() == "---":
1818+ end_idx = i
1819+ break
1820+
1821+ if end_idx is None:
1822+ raise ValueError("SKILL.md missing frontmatter (no closing ---)")
1823+
1824+ name = ""
1825+ description = ""
1826+ frontmatter_lines = lines[1:end_idx]
1827+ i = 0
1828+ while i < len(frontmatter_lines):
1829+ line = frontmatter_lines[i]
1830+ if line.startswith("name:"):
1831+ name = line[len("name:"):].strip().strip('"').strip("'")
1832+ elif line.startswith("description:"):
1833+ value = line[len("description:"):].strip()
1834+ # Handle YAML multiline indicators (>, |, >-, |-)
1835+ if value in (">", "|", ">-", "|-"):
1836+ continuation_lines: list[str] = []
1837+ i += 1
1838+ while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith(" ") or frontmatter_lines[i].startswith("\t")):
1839+ continuation_lines.append(frontmatter_lines[i].strip())
1840+ i += 1
1841+ description = " ".join(continuation_lines)
1842+ continue
1843+ else:
1844+ description = value.strip('"').strip("'")
1845+ i += 1
1846+
1847+ return name, description, content
1848diff --git a/dot_config/opencode/skills/weekly-review/SKILL.md b/dot_config/opencode/skills/weekly-review/SKILL.md
1849new file mode 100644
1850index 0000000000000000000000000000000000000000..392dc89193dcdb582c3508c2767786e77632013b
1851--- /dev/null
1852+++ b/dot_config/opencode/skills/weekly-review/SKILL.md
1853@@ -0,0 +1,115 @@
1854+---
1855+name: weekly-review
1856+description: Analyze recent sessions to find patterns, recurring mistakes, and propose improvements to AGENTS.md and lessons.md.
1857+compatibility: opencode
1858+---
1859+
1860+## What this is
1861+
1862+A periodic deep review of how the user and assistant work together. Reads session
1863+history from the database, dispatches parallel agents to analyze conversations,
1864+then runs multiple reflection loops to surface patterns and propose changes.
1865+
1866+## When to trigger
1867+
1868+User runs `/review-week`. Not autonomous — this is a deliberate review.
1869+
1870+## Inputs
1871+
1872+- `$ARGUMENTS`: optional time range in days (default: 7). Example: `/review-week 14`
1873+
1874+## Data source
1875+
1876+Sessions are stored in SQLite at `~/.local/share/opencode/opencode.db`.
1877+
1878+Key tables:
1879+- `session`: metadata (id, title, directory, time_created, time_updated)
1880+- `message`: per-session messages (id, session_id, data JSON)
1881+- `part`: message content (id, message_id, session_id, data JSON with `type` and `text` fields)
1882+
1883+## Workflow
1884+
1885+### 1. Gather sessions
1886+
1887+```sql
1888+SELECT id, title, directory,
1889+ datetime(time_created/1000, 'unixepoch', 'localtime') as created,
1890+ datetime(time_updated/1000, 'unixepoch', 'localtime') as updated
1891+FROM session
1892+WHERE time_created > (strftime('%s', 'now', '-N days') * 1000)
1893+ AND title NOT LIKE '%@explore%'
1894+ AND title NOT LIKE '%@general%'
1895+ORDER BY time_created ASC;
1896+```
1897+
1898+Replace `N` with the requested day range. Filter out subagent sessions — they're
1899+noise for pattern analysis.
1900+
1901+Count messages per session to find the substantive ones (>3 messages).
1902+
1903+### 2. Dispatch parallel agents
1904+
1905+Group substantive sessions into 3-5 batches. For each batch, launch a `general`
1906+subagent with instructions to:
1907+
1908+1. Extract conversation text from the `part` table:
1909+ ```sql
1910+ SELECT p.data FROM part p
1911+ JOIN message m ON m.id = p.message_id
1912+ WHERE m.session_id = 'SESSION_ID'
1913+ AND json_extract(p.data, '$.type') = 'text'
1914+ ORDER BY p.time_created ASC;
1915+ ```
1916+2. For each session, identify:
1917+ - What the user was trying to accomplish
1918+ - Corrections the user made to the assistant
1919+ - Frustration signals
1920+ - Where the assistant over- or under-delivered
1921+ - Design decisions and their quality
1922+3. Return a comprehensive analysis focused on patterns and lessons, not summaries.
1923+
1924+### 3. Read context files
1925+
1926+While agents run, read:
1927+- `~/.config/opencode/AGENTS.md`
1928+- `~/.config/opencode/lessons.md`
1929+
1930+### 4. Reflection loops
1931+
1932+Run 2-3 passes over the combined agent output:
1933+
1934+**Loop 1 — Patterns:** What themes repeat across sessions? Which are the most
1935+frequent and most damaging? Cross-reference against existing AGENTS.md and
1936+lessons.md rules — are these known problems (compliance failure) or new gaps?
1937+
1938+**Loop 2 — Design problems:** Why does the system produce these failures?
1939+What's the root cause behind the patterns? Are existing rules too vague,
1940+misplaced, or structurally unenforceable?
1941+
1942+**Loop 3 — Proposals:** Concrete, minimal changes. Prefer strengthening existing
1943+rules over adding new ones. Every proposed change must map to a specific pattern
1944+found in the data.
1945+
1946+### 5. Present findings
1947+
1948+Show the user:
1949+- Session count and coverage
1950+- Top patterns with frequency and severity
1951+- The compliance question: which patterns already have rules?
1952+- Proposed changes as exact edits (old → new)