c87b816a776652345b4e72629ab2d1ceef5677f0
- Author
- TheEdgeOfRage <git@theedgeofrage.com>
- Committer
- TheEdgeOfRage <git@theedgeofrage.com>
- Date
Message
Diff
This diff is truncated to protect this page.
1diff --git a/dot_config/opencode/skills/qmd/SKILL.md b/dot_config/opencode/skills/qmd/SKILL.md
2deleted file mode 100644
3index de910987c35e99fc0024863f7fbbe5040974387b..0000000000000000000000000000000000000000
4--- a/dot_config/opencode/skills/qmd/SKILL.md
5+++ /dev/null
6@@ -1,125 +0,0 @@
7----
8-name: qmd
9-description: Search markdown knowledge bases, notes, and documentation using QMD. Use when users ask to search notes, find documents, or look up information in local files.
10-license: MIT
11-compatibility: Requires qmd CLI
12-metadata:
13-author: tobi
14-version: "2.0.0"
15-allowed-tools: Bash(qmd:\*)
16----
17-
18-# QMD - Quick Markdown Search
19-
20-Local search engine for markdown content.
21-
22-## Status
23-
24-!`qmd status 2>/dev/null || echo "Not installed: npm install -g @tobilu/qmd"`
25-
26-### Query Types
27-
28-| Type | Method | Input |
29-| ------ | ------ | ------------------------------------------- |
30-| `lex` | BM25 | Keywords — exact terms, names, code |
31-| `vec` | Vector | Question — natural language |
32-| `hyde` | Vector | Answer — hypothetical result (50-100 words) |
33-
34-### Writing Good Queries
35-
36-**lex (keyword)**
37-
38-- 2-5 terms, no filler words
39-- Exact phrase: `"connection pool"` (quoted)
40-- Exclude terms: `performance -sports` (minus prefix)
41-- Code identifiers work: `handleError async`
42-
43-**vec (semantic)**
44-
45-- Full natural language question
46-- Be specific: `"how does the rate limiter handle burst traffic"`
47-- Include context: `"in the payment service, how are refunds processed"`
48-
49-**hyde (hypothetical document)**
50-
51-- Write 50-100 words of what the _answer_ looks like
52-- Use the vocabulary you expect in the result
53-
54-**expand (auto-expand)**
55-
56-- Use a single-line query (implicit) or `expand: question` on its own line
57-- Lets the local LLM generate lex/vec/hyde variations
58-- Do not mix `expand:` with other typed lines — it's either a standalone expand query or a full query document
59-
60-### Intent (Disambiguation)
61-
62-When a query term is ambiguous, add `intent` to steer results:
63-
64-```json
65-{
66- "searches": [{ "type": "lex", "query": "performance" }],
67- "intent": "web page load times and Core Web Vitals"
68-}
69-```
70-
71-Intent affects expansion, reranking, chunk selection, and snippet extraction. It does not search on its own — it's a steering signal that disambiguates queries like "performance" (web-perf vs team health vs fitness).
72-
73-### Combining Types
74-
75-| Goal | Approach |
76-| --------------------- | ----------------------------------------------------- |
77-| Know exact terms | `lex` only |
78-| Don't know vocabulary | Use a single-line query (implicit `expand:`) or `vec` |
79-| Best recall | `lex` + `vec` |
80-| Complex topic | `lex` + `vec` + `hyde` |
81-| Ambiguous query | Add `intent` to any combination above |
82-
83-First query gets 2x weight in fusion — put your best guess first.
84-
85-### Lex Query Syntax
86-
87-| Syntax | Meaning | Example |
88-| ---------- | ------------ | ---------------------------- |
89-| `term` | Prefix match | `perf` matches "performance" |
90-| `"phrase"` | Exact phrase | `"rate limiter"` |
91-| `-term` | Exclude | `performance -sports` |
92-
93-Note: `-term` only works in lex queries, not vec/hyde.
94-
95-### Collection Filtering
96-
97-```json
98-{ "collections": ["docs"] } // Single
99-{ "collections": ["docs", "notes"] } // Multiple (OR)
100-```
101-
102-Omit to search all collections.
103-
104-## CLI
105-
106diff --git a/dot_config/opencode/skills/skill-creator/LICENSE.txt b/dot_config/opencode/skills/skill-creator/LICENSE.txt
107deleted file mode 100644
108index 7a4a3ea2424c09fbe48d455aed1eaa94d9124835..0000000000000000000000000000000000000000
109--- a/dot_config/opencode/skills/skill-creator/LICENSE.txt
110+++ /dev/null
111@@ -1,202 +0,0 @@
112-
113- Apache License
114- Version 2.0, January 2004
115- http://www.apache.org/licenses/
116-
117- TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
118-
119- 1. Definitions.
120-
121- "License" shall mean the terms and conditions for use, reproduction,
122- and distribution as defined by Sections 1 through 9 of this document.
123-
124- "Licensor" shall mean the copyright owner or entity authorized by
125- the copyright owner that is granting the License.
126-
127- "Legal Entity" shall mean the union of the acting entity and all
128- other entities that control, are controlled by, or are under common
129- control with that entity. For the purposes of this definition,
130- "control" means (i) the power, direct or indirect, to cause the
131- direction or management of such entity, whether by contract or
132- otherwise, or (ii) ownership of fifty percent (50%) or more of the
133- outstanding shares, or (iii) beneficial ownership of such entity.
134-
135- "You" (or "Your") shall mean an individual or Legal Entity
136- exercising permissions granted by this License.
137-
138- "Source" form shall mean the preferred form for making modifications,
139- including but not limited to software source code, documentation
140- source, and configuration files.
141-
142- "Object" form shall mean any form resulting from mechanical
143- transformation or translation of a Source form, including but
144- not limited to compiled object code, generated documentation,
145- and conversions to other media types.
146-
147- "Work" shall mean the work of authorship, whether in Source or
148- Object form, made available under the License, as indicated by a
149- copyright notice that is included in or attached to the work
150- (an example is provided in the Appendix below).
151-
152- "Derivative Works" shall mean any work, whether in Source or Object
153- form, that is based on (or derived from) the Work and for which the
154- editorial revisions, annotations, elaborations, or other modifications
155- represent, as a whole, an original work of authorship. For the purposes
156- of this License, Derivative Works shall not include works that remain
157- separable from, or merely link (or bind by name) to the interfaces of,
158- the Work and Derivative Works thereof.
159-
160- "Contribution" shall mean any work of authorship, including
161- the original version of the Work and any modifications or additions
162- to that Work or Derivative Works thereof, that is intentionally
163- submitted to Licensor for inclusion in the Work by the copyright owner
164- or by an individual or Legal Entity authorized to submit on behalf of
165- the copyright owner. For the purposes of this definition, "submitted"
166- means any form of electronic, verbal, or written communication sent
167- to the Licensor or its representatives, including but not limited to
168- communication on electronic mailing lists, source code control systems,
169- and issue tracking systems that are managed by, or on behalf of, the
170- Licensor for the purpose of discussing and improving the Work, but
171- excluding communication that is conspicuously marked or otherwise
172- designated in writing by the copyright owner as "Not a Contribution."
173-
174- "Contributor" shall mean Licensor and any individual or Legal Entity
175- on behalf of whom a Contribution has been received by Licensor and
176- subsequently incorporated within the Work.
177-
178- 2. Grant of Copyright License. Subject to the terms and conditions of
179- this License, each Contributor hereby grants to You a perpetual,
180- worldwide, non-exclusive, no-charge, royalty-free, irrevocable
181- copyright license to reproduce, prepare Derivative Works of,
182- publicly display, publicly perform, sublicense, and distribute the
183- Work and such Derivative Works in Source or Object form.
184-
185- 3. Grant of Patent License. Subject to the terms and conditions of
186- this License, each Contributor hereby grants to You a perpetual,
187- worldwide, non-exclusive, no-charge, royalty-free, irrevocable
188- (except as stated in this section) patent license to make, have made,
189- use, offer to sell, sell, import, and otherwise transfer the Work,
190- where such license applies only to those patent claims licensable
191- by such Contributor that are necessarily infringed by their
192- Contribution(s) alone or by combination of their Contribution(s)
193- with the Work to which such Contribution(s) was submitted. If You
194- institute patent litigation against any entity (including a
195- cross-claim or counterclaim in a lawsuit) alleging that the Work
196- or a Contribution incorporated within the Work constitutes direct
197- or contributory patent infringement, then any patent licenses
198- granted to You under this License for that Work shall terminate
199- as of the date such litigation is filed.
200-
201- 4. Redistribution. You may reproduce and distribute copies of the
202- Work or Derivative Works thereof in any medium, with or without
203- modifications, and in Source or Object form, provided that You
204- meet the following conditions:
205-
206- (a) You must give any other recipients of the Work or
207- Derivative Works a copy of this License; and
208-
209- (b) You must cause any modified files to carry prominent notices
210- stating that You changed the files; and
211diff --git a/dot_config/opencode/skills/skill-creator/SKILL.md b/dot_config/opencode/skills/skill-creator/SKILL.md
212deleted file mode 100644
213index b65c5a9c16dcff44d5e5151814854e68c60a2b65..0000000000000000000000000000000000000000
214--- a/dot_config/opencode/skills/skill-creator/SKILL.md
215+++ /dev/null
216@@ -1,503 +0,0 @@
217----
218-name: skill-creator
219diff --git a/dot_config/opencode/skills/skill-creator/agents/analyzer.md b/dot_config/opencode/skills/skill-creator/agents/analyzer.md
220deleted file mode 100644
221index 14e41d6068635f4dd3fb878fd1626312395dda63..0000000000000000000000000000000000000000
222--- a/dot_config/opencode/skills/skill-creator/agents/analyzer.md
223+++ /dev/null
224@@ -1,274 +0,0 @@
225-# Post-hoc Analyzer Agent
226-
227-Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
228-
229-## Role
230-
231-After the blind comparator determines a winner, the Post-hoc Analyzer "unblids" the results by examining the skills and transcripts. The goal is to extract actionable insights: what made the winner better, and how can the loser be improved?
232-
233-## Inputs
234-
235-You receive these parameters in your prompt:
236-
237-- **winner**: "A" or "B" (from blind comparison)
238-- **winner_skill_path**: Path to the skill that produced the winning output
239-- **winner_transcript_path**: Path to the execution transcript for the winner
240-- **loser_skill_path**: Path to the skill that produced the losing output
241-- **loser_transcript_path**: Path to the execution transcript for the loser
242-- **comparison_result_path**: Path to the blind comparator's output JSON
243-- **output_path**: Where to save the analysis results
244-
245-## Process
246-
247-### Step 1: Read Comparison Result
248-
249-1. Read the blind comparator's output at comparison_result_path
250-2. Note the winning side (A or B), the reasoning, and any scores
251-3. Understand what the comparator valued in the winning output
252-
253-### Step 2: Read Both Skills
254-
255-1. Read the winner skill's SKILL.md and key referenced files
256-2. Read the loser skill's SKILL.md and key referenced files
257-3. Identify structural differences:
258- - Instructions clarity and specificity
259- - Script/tool usage patterns
260- - Example coverage
261- - Edge case handling
262-
263-### Step 3: Read Both Transcripts
264-
265-1. Read the winner's transcript
266-2. Read the loser's transcript
267-3. Compare execution patterns:
268- - How closely did each follow their skill's instructions?
269- - What tools were used differently?
270- - Where did the loser diverge from optimal behavior?
271- - Did either encounter errors or make recovery attempts?
272-
273-### Step 4: Analyze Instruction Following
274-
275-For each transcript, evaluate:
276-- Did the agent follow the skill's explicit instructions?
277-- Did the agent use the skill's provided tools/scripts?
278-- Were there missed opportunities to leverage skill content?
279-- Did the agent add unnecessary steps not in the skill?
280-
281-Score instruction following 1-10 and note specific issues.
282-
283-### Step 5: Identify Winner Strengths
284-
285-Determine what made the winner better:
286-- Clearer instructions that led to better behavior?
287-- Better scripts/tools that produced better output?
288-- More comprehensive examples that guided edge cases?
289-- Better error handling guidance?
290-
291-Be specific. Quote from skills/transcripts where relevant.
292-
293-### Step 6: Identify Loser Weaknesses
294-
295-Determine what held the loser back:
296-- Ambiguous instructions that led to suboptimal choices?
297-- Missing tools/scripts that forced workarounds?
298-- Gaps in edge case coverage?
299-- Poor error handling that caused failures?
300-
301-### Step 7: Generate Improvement Suggestions
302-
303-Based on the analysis, produce actionable suggestions for improving the loser skill:
304-- Specific instruction changes to make
305-- Tools/scripts to add or modify
306-- Examples to include
307-- Edge cases to address
308-
309-Prioritize by impact. Focus on changes that would have changed the outcome.
310-
311-### Step 8: Write Analysis Results
312-
313-Save structured analysis to `{output_path}`.
314-
315-## Output Format
316-
317-Write a JSON file with this structure:
318-
319-```json
320-{
321- "comparison_summary": {
322- "winner": "A",
323- "winner_skill": "path/to/winner/skill",
324diff --git a/dot_config/opencode/skills/skill-creator/agents/comparator.md b/dot_config/opencode/skills/skill-creator/agents/comparator.md
325deleted file mode 100644
326index 80e00eb45db3ee53a132fc2ba97fd59a7339e563..0000000000000000000000000000000000000000
327--- a/dot_config/opencode/skills/skill-creator/agents/comparator.md
328+++ /dev/null
329@@ -1,202 +0,0 @@
330-# Blind Comparator Agent
331-
332-Compare two outputs WITHOUT knowing which skill produced them.
333-
334-## Role
335-
336-The Blind Comparator judges which output better accomplishes the eval task. You receive two outputs labeled A and B, but you do NOT know which skill produced which. This prevents bias toward a particular skill or approach.
337-
338-Your judgment is based purely on output quality and task completion.
339-
340-## Inputs
341-
342-You receive these parameters in your prompt:
343-
344-- **output_a_path**: Path to the first output file or directory
345-- **output_b_path**: Path to the second output file or directory
346-- **eval_prompt**: The original task/prompt that was executed
347-- **expectations**: List of expectations to check (optional - may be empty)
348-
349-## Process
350-
351-### Step 1: Read Both Outputs
352-
353-1. Examine output A (file or directory)
354-2. Examine output B (file or directory)
355-3. Note the type, structure, and content of each
356-4. If outputs are directories, examine all relevant files inside
357-
358-### Step 2: Understand the Task
359-
360-1. Read the eval_prompt carefully
361-2. Identify what the task requires:
362- - What should be produced?
363- - What qualities matter (accuracy, completeness, format)?
364- - What would distinguish a good output from a poor one?
365-
366-### Step 3: Generate Evaluation Rubric
367-
368-Based on the task, generate a rubric with two dimensions:
369-
370-**Content Rubric** (what the output contains):
371-| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
372-|-----------|----------|----------------|---------------|
373-| Correctness | Major errors | Minor errors | Fully correct |
374-| Completeness | Missing key elements | Mostly complete | All elements present |
375-| Accuracy | Significant inaccuracies | Minor inaccuracies | Accurate throughout |
376-
377-**Structure Rubric** (how the output is organized):
378-| Criterion | 1 (Poor) | 3 (Acceptable) | 5 (Excellent) |
379-|-----------|----------|----------------|---------------|
380-| Organization | Disorganized | Reasonably organized | Clear, logical structure |
381-| Formatting | Inconsistent/broken | Mostly consistent | Professional, polished |
382-| Usability | Difficult to use | Usable with effort | Easy to use |
383-
384-Adapt criteria to the specific task. For example:
385-- PDF form → "Field alignment", "Text readability", "Data placement"
386-- Document → "Section structure", "Heading hierarchy", "Paragraph flow"
387-- Data output → "Schema correctness", "Data types", "Completeness"
388-
389-### Step 4: Evaluate Each Output Against the Rubric
390-
391-For each output (A and B):
392-
393-1. **Score each criterion** on the rubric (1-5 scale)
394-2. **Calculate dimension totals**: Content score, Structure score
395-3. **Calculate overall score**: Average of dimension scores, scaled to 1-10
396-
397-### Step 5: Check Assertions (if provided)
398-
399-If expectations are provided:
400-
401-1. Check each expectation against output A
402-2. Check each expectation against output B
403-3. Count pass rates for each output
404-4. Use expectation scores as secondary evidence (not the primary decision factor)
405-
406-### Step 6: Determine the Winner
407-
408-Compare A and B based on (in priority order):
409-
410-1. **Primary**: Overall rubric score (content + structure)
411-2. **Secondary**: Assertion pass rates (if applicable)
412-3. **Tiebreaker**: If truly equal, declare a TIE
413-
414-Be decisive - ties should be rare. One output is usually better, even if marginally.
415-
416-### Step 7: Write Comparison Results
417-
418-Save results to a JSON file at the path specified (or `comparison.json` if not specified).
419-
420-## Output Format
421-
422-Write a JSON file with this structure:
423-
424-```json
425-{
426- "winner": "A",
427- "reasoning": "Output A provides a complete solution with proper formatting and all required fields. Output B is missing the date field and has formatting inconsistencies.",
428- "rubric": {
429diff --git a/dot_config/opencode/skills/skill-creator/agents/grader.md b/dot_config/opencode/skills/skill-creator/agents/grader.md
430deleted file mode 100644
431index 558ab05c0a9a8bb062ef4c51823d4d76c3acf7c4..0000000000000000000000000000000000000000
432--- a/dot_config/opencode/skills/skill-creator/agents/grader.md
433+++ /dev/null
434@@ -1,223 +0,0 @@
435-# Grader Agent
436-
437-Evaluate expectations against an execution transcript and outputs.
438-
439-## Role
440-
441-The Grader reviews a transcript and output files, then determines whether each expectation passes or fails. Provide clear evidence for each judgment.
442-
443diff --git a/dot_config/opencode/skills/skill-creator/assets/eval_review.html b/dot_config/opencode/skills/skill-creator/assets/eval_review.html
444deleted file mode 100644
445index 938ff32aed9bffabf723bd5492d720f4736c8e4d..0000000000000000000000000000000000000000
446--- a/dot_config/opencode/skills/skill-creator/assets/eval_review.html
447+++ /dev/null
448@@ -1,146 +0,0 @@
449-<!DOCTYPE html>
450-<html lang="en">
451-<head>
452- <meta charset="UTF-8">
453- <meta name="viewport" content="width=device-width, initial-scale=1.0">
454- <title>Eval Set Review - __SKILL_NAME_PLACEHOLDER__</title>
455- <link rel="preconnect" href="https://fonts.googleapis.com">
456- <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
457- <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
458- <style>
459- * { box-sizing: border-box; margin: 0; padding: 0; }
460- body { font-family: 'Lora', Georgia, serif; background: #faf9f5; padding: 2rem; color: #141413; }
461- h1 { font-family: 'Poppins', sans-serif; margin-bottom: 0.5rem; font-size: 1.5rem; }
462- .description { color: #b0aea5; margin-bottom: 1.5rem; font-style: italic; max-width: 900px; }
463- .controls { margin-bottom: 1rem; display: flex; gap: 0.5rem; }
464- .btn { font-family: 'Poppins', sans-serif; padding: 0.5rem 1rem; border: none; border-radius: 6px; cursor: pointer; font-size: 0.875rem; font-weight: 500; }
465- .btn-add { background: #6a9bcc; color: white; }
466- .btn-add:hover { background: #5889b8; }
467- .btn-export { background: #d97757; color: white; }
468- .btn-export:hover { background: #c4613f; }
469- table { width: 100%; max-width: 1100px; border-collapse: collapse; background: white; border-radius: 6px; overflow: hidden; box-shadow: 0 1px 3px rgba(0,0,0,0.08); }
470- th { font-family: 'Poppins', sans-serif; background: #141413; color: #faf9f5; padding: 0.75rem 1rem; text-align: left; font-size: 0.875rem; }
471- td { padding: 0.75rem 1rem; border-bottom: 1px solid #e8e6dc; vertical-align: top; }
472- tr:nth-child(even) td { background: #faf9f5; }
473- tr:hover td { background: #f3f1ea; }
474- .section-header td { background: #e8e6dc; font-family: 'Poppins', sans-serif; font-weight: 500; font-size: 0.8rem; color: #141413; text-transform: uppercase; letter-spacing: 0.05em; }
475- .query-input { width: 100%; padding: 0.4rem; border: 1px solid #e8e6dc; border-radius: 4px; font-size: 0.875rem; font-family: 'Lora', Georgia, serif; resize: vertical; min-height: 60px; }
476- .query-input:focus { outline: none; border-color: #d97757; box-shadow: 0 0 0 2px rgba(217,119,87,0.15); }
477- .toggle { position: relative; display: inline-block; width: 44px; height: 24px; }
478- .toggle input { opacity: 0; width: 0; height: 0; }
479- .toggle .slider { position: absolute; inset: 0; background: #b0aea5; border-radius: 24px; cursor: pointer; transition: 0.2s; }
480- .toggle .slider::before { content: ""; position: absolute; width: 18px; height: 18px; left: 3px; bottom: 3px; background: white; border-radius: 50%; transition: 0.2s; }
481- .toggle input:checked + .slider { background: #d97757; }
482- .toggle input:checked + .slider::before { transform: translateX(20px); }
483- .btn-delete { background: #c44; color: white; padding: 0.3rem 0.6rem; border: none; border-radius: 4px; cursor: pointer; font-size: 0.75rem; font-family: 'Poppins', sans-serif; }
484- .btn-delete:hover { background: #a33; }
485- .summary { margin-top: 1rem; color: #b0aea5; font-size: 0.875rem; }
486- </style>
487-</head>
488-<body>
489- <h1>Eval Set Review: <span id="skill-name">__SKILL_NAME_PLACEHOLDER__</span></h1>
490- <p class="description">Current description: <span id="skill-desc">__SKILL_DESCRIPTION_PLACEHOLDER__</span></p>
491-
492- <div class="controls">
493- <button class="btn btn-add" onclick="addRow()">+ Add Query</button>
494- <button class="btn btn-export" onclick="exportEvalSet()">Export Eval Set</button>
495- </div>
496-
497- <table>
498- <thead>
499- <tr>
500- <th style="width:65%">Query</th>
501- <th style="width:18%">Should Trigger</th>
502- <th style="width:10%">Actions</th>
503- </tr>
504- </thead>
505- <tbody id="eval-body"></tbody>
506- </table>
507-
508- <p class="summary" id="summary"></p>
509-
510- <script>
511- const EVAL_DATA = __EVAL_DATA_PLACEHOLDER__;
512-
513- let evalItems = [...EVAL_DATA];
514-
515- function render() {
516- const tbody = document.getElementById('eval-body');
517- tbody.innerHTML = '';
518-
519- // Sort: should-trigger first, then should-not-trigger
520- const sorted = evalItems
521- .map((item, origIdx) => ({ ...item, origIdx }))
522- .sort((a, b) => (b.should_trigger ? 1 : 0) - (a.should_trigger ? 1 : 0));
523-
524- let lastGroup = null;
525- sorted.forEach(item => {
526- const group = item.should_trigger ? 'trigger' : 'no-trigger';
527- if (group !== lastGroup) {
528- const headerRow = document.createElement('tr');
529- headerRow.className = 'section-header';
530- headerRow.innerHTML = `<td colspan="3">${item.should_trigger ? 'Should Trigger' : 'Should NOT Trigger'}</td>`;
531- tbody.appendChild(headerRow);
532- lastGroup = group;
533- }
534-
535- const idx = item.origIdx;
536- const tr = document.createElement('tr');
537- tr.innerHTML = `
538- <td><textarea class="query-input" onchange="updateQuery(${idx}, this.value)">${escapeHtml(item.query)}</textarea></td>
539- <td>
540- <label class="toggle">
541- <input type="checkbox" ${item.should_trigger ? 'checked' : ''} onchange="updateTrigger(${idx}, this.checked)">
542- <span class="slider"></span>
543- </label>
544- <span style="margin-left:8px;font-size:0.8rem;color:#b0aea5">${item.should_trigger ? 'Yes' : 'No'}</span>
545- </td>
546- <td><button class="btn-delete" onclick="deleteRow(${idx})">Delete</button></td>
547- `;
548diff --git a/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py b/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py
549deleted file mode 100644
550index 7fa5978631fed1ed545591dbb2b0eb21ce3f3d08..0000000000000000000000000000000000000000
551--- a/dot_config/opencode/skills/skill-creator/eval-viewer/generate_review.py
552+++ /dev/null
553@@ -1,471 +0,0 @@
554-#!/usr/bin/env python3
555-"""Generate and serve a review page for eval results.
556-
557-Reads the workspace directory, discovers runs (directories with outputs/),
558-embeds all output data into a self-contained HTML page, and serves it via
559-a tiny HTTP server. Feedback auto-saves to feedback.json in the workspace.
560-
561-Usage:
562- python generate_review.py <workspace-path> [--port PORT] [--skill-name NAME]
563- python generate_review.py <workspace-path> --previous-feedback /path/to/old/feedback.json
564-
565-No dependencies beyond the Python stdlib are required.
566-"""
567-
568-import argparse
569-import base64
570-import json
571-import mimetypes
572-import os
573-import re
574-import signal
575-import subprocess
576-import sys
577-import time
578-import webbrowser
579-from functools import partial
580-from http.server import HTTPServer, BaseHTTPRequestHandler
581-from pathlib import Path
582-
583-# Files to exclude from output listings
584-METADATA_FILES = {"transcript.md", "user_notes.md", "metrics.json"}
585-
586-# Extensions we render as inline text
587-TEXT_EXTENSIONS = {
588- ".txt", ".md", ".json", ".csv", ".py", ".js", ".ts", ".tsx", ".jsx",
589- ".yaml", ".yml", ".xml", ".html", ".css", ".sh", ".rb", ".go", ".rs",
590- ".java", ".c", ".cpp", ".h", ".hpp", ".sql", ".r", ".toml",
591-}
592-
593-# Extensions we render as inline images
594-IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".gif", ".svg", ".webp"}
595-
596-# MIME type overrides for common types
597-MIME_OVERRIDES = {
598- ".svg": "image/svg+xml",
599- ".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
600- ".docx": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
601- ".pptx": "application/vnd.openxmlformats-officedocument.presentationml.presentation",
602-}
603-
604-
605-def get_mime_type(path: Path) -> str:
606- ext = path.suffix.lower()
607- if ext in MIME_OVERRIDES:
608- return MIME_OVERRIDES[ext]
609- mime, _ = mimetypes.guess_type(str(path))
610- return mime or "application/octet-stream"
611-
612-
613-def find_runs(workspace: Path) -> list[dict]:
614- """Recursively find directories that contain an outputs/ subdirectory."""
615- runs: list[dict] = []
616- _find_runs_recursive(workspace, workspace, runs)
617- runs.sort(key=lambda r: (r.get("eval_id", float("inf")), r["id"]))
618- return runs
619-
620-
621-def _find_runs_recursive(root: Path, current: Path, runs: list[dict]) -> None:
622- if not current.is_dir():
623- return
624-
625- outputs_dir = current / "outputs"
626- if outputs_dir.is_dir():
627- run = build_run(root, current)
628- if run:
629- runs.append(run)
630- return
631-
632- skip = {"node_modules", ".git", "__pycache__", "skill", "inputs"}
633- for child in sorted(current.iterdir()):
634- if child.is_dir() and child.name not in skip:
635- _find_runs_recursive(root, child, runs)
636-
637-
638-def build_run(root: Path, run_dir: Path) -> dict | None:
639- """Build a run dict with prompt, outputs, and grading data."""
640- prompt = ""
641- eval_id = None
642-
643- # Try eval_metadata.json
644- for candidate in [run_dir / "eval_metadata.json", run_dir.parent / "eval_metadata.json"]:
645- if candidate.exists():
646- try:
647- metadata = json.loads(candidate.read_text())
648- prompt = metadata.get("prompt", "")
649- eval_id = metadata.get("eval_id")
650- except (json.JSONDecodeError, OSError):
651- pass
652- if prompt:
653diff --git a/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html b/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html
654deleted file mode 100644
655index 6d8e96348a02e66c3363d2ff3b3ae58ac11e6382..0000000000000000000000000000000000000000
656--- a/dot_config/opencode/skills/skill-creator/eval-viewer/viewer.html
657+++ /dev/null
658@@ -1,1325 +0,0 @@
659-<!DOCTYPE html>
660-<html lang="en">
661-<head>
662- <meta charset="UTF-8">
663- <meta name="viewport" content="width=device-width, initial-scale=1.0">
664- <title>Eval Review</title>
665- <link rel="preconnect" href="https://fonts.googleapis.com">
666- <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
667- <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
668- <script src="https://cdn.sheetjs.com/xlsx-0.20.3/package/dist/xlsx.full.min.js" integrity="sha384-EnyY0/GSHQGSxSgMwaIPzSESbqoOLSexfnSMN2AP+39Ckmn92stwABZynq1JyzdT" crossorigin="anonymous"></script>
669- <style>
670- :root {
671- --bg: #faf9f5;
672- --surface: #ffffff;
673- --border: #e8e6dc;
674- --text: #141413;
675- --text-muted: #b0aea5;
676- --accent: #d97757;
677- --accent-hover: #c4613f;
678- --green: #788c5d;
679- --green-bg: #eef2e8;
680- --red: #c44;
681- --red-bg: #fceaea;
682- --header-bg: #141413;
683- --header-text: #faf9f5;
684- --radius: 6px;
685- }
686-
687- * { box-sizing: border-box; margin: 0; padding: 0; }
688-
689- body {
690- font-family: 'Lora', Georgia, serif;
691- background: var(--bg);
692- color: var(--text);
693- height: 100vh;
694- display: flex;
695- flex-direction: column;
696- }
697-
698- /* ---- Header ---- */
699- .header {
700- background: var(--header-bg);
701- color: var(--header-text);
702- padding: 1rem 2rem;
703- display: flex;
704- justify-content: space-between;
705- align-items: center;
706- flex-shrink: 0;
707- }
708- .header h1 {
709- font-family: 'Poppins', sans-serif;
710- font-size: 1.25rem;
711- font-weight: 600;
712- }
713- .header .instructions {
714- font-size: 0.8rem;
715- opacity: 0.7;
716- margin-top: 0.25rem;
717- }
718- .header .progress {
719- font-size: 0.875rem;
720- opacity: 0.8;
721- text-align: right;
722- }
723-
724- /* ---- Main content ---- */
725- .main {
726- flex: 1;
727- overflow-y: auto;
728- padding: 1.5rem 2rem;
729- display: flex;
730- flex-direction: column;
731- gap: 1.25rem;
732- }
733-
734- /* ---- Sections ---- */
735- .section {
736- background: var(--surface);
737- border: 1px solid var(--border);
738- border-radius: var(--radius);
739- flex-shrink: 0;
740- }
741- .section-header {
742- font-family: 'Poppins', sans-serif;
743- padding: 0.75rem 1rem;
744- font-size: 0.75rem;
745- font-weight: 500;
746- text-transform: uppercase;
747- letter-spacing: 0.05em;
748- color: var(--text-muted);
749- border-bottom: 1px solid var(--border);
750- background: var(--bg);
751- }
752- .section-body {
753- padding: 1rem;
754- }
755-
756- /* ---- Config badge ---- */
757- .config-badge {
758diff --git a/dot_config/opencode/skills/skill-creator/references/schemas.md b/dot_config/opencode/skills/skill-creator/references/schemas.md
759deleted file mode 100644
760index b6eeaa2d4a34c1653069585c6c5603da39a5bdbe..0000000000000000000000000000000000000000
761--- a/dot_config/opencode/skills/skill-creator/references/schemas.md
762+++ /dev/null
763@@ -1,430 +0,0 @@
764-# JSON Schemas
765-
766-This document defines the JSON schemas used by skill-creator.
767-
768----
769-
770-## evals.json
771-
772-Defines the evals for a skill. Located at `evals/evals.json` within the skill directory.
773-
774-```json
775-{
776- "skill_name": "example-skill",
777- "evals": [
778- {
779- "id": 1,
780- "prompt": "User's example prompt",
781- "expected_output": "Description of expected result",
782- "files": ["evals/files/sample1.pdf"],
783- "expectations": [
784- "The output includes X",
785- "The skill used script Y"
786- ]
787- }
788- ]
789-}
790-```
791-
792-**Fields:**
793-- `skill_name`: Name matching the skill's frontmatter
794-- `evals[].id`: Unique integer identifier
795-- `evals[].prompt`: The task to execute
796-- `evals[].expected_output`: Human-readable description of success
797-- `evals[].files`: Optional list of input file paths (relative to skill root)
798-- `evals[].expectations`: List of verifiable statements
799-
800----
801-
802-## history.json
803-
804-Tracks version progression in Improve mode. Located at workspace root.
805-
806-```json
807-{
808- "started_at": "2026-01-15T10:30:00Z",
809- "skill_name": "pdf",
810- "current_best": "v2",
811- "iterations": [
812- {
813- "version": "v0",
814- "parent": null,
815- "expectation_pass_rate": 0.65,
816- "grading_result": "baseline",
817- "is_current_best": false
818- },
819- {
820- "version": "v1",
821- "parent": "v0",
822- "expectation_pass_rate": 0.75,
823- "grading_result": "won",
824- "is_current_best": false
825- },
826- {
827- "version": "v2",
828- "parent": "v1",
829- "expectation_pass_rate": 0.85,
830- "grading_result": "won",
831- "is_current_best": true
832- }
833- ]
834-}
835-```
836-
837-**Fields:**
838-- `started_at`: ISO timestamp of when improvement started
839-- `skill_name`: Name of the skill being improved
840-- `current_best`: Version identifier of the best performer
841-- `iterations[].version`: Version identifier (v0, v1, ...)
842-- `iterations[].parent`: Parent version this was derived from
843-- `iterations[].expectation_pass_rate`: Pass rate from grading
844-- `iterations[].grading_result`: "baseline", "won", "lost", or "tie"
845-- `iterations[].is_current_best`: Whether this is the current best version
846-
847----
848-
849-## grading.json
850-
851-Output from the grader agent. Located at `<run-dir>/grading.json`.
852-
853-```json
854-{
855- "expectations": [
856- {
857- "text": "The output includes the name 'John Smith'",
858- "passed": true,
859- "evidence": "Found in transcript Step 3: 'Extracted names: John Smith, Sarah Johnson'"
860- },
861- {
862- "text": "The spreadsheet has a SUM formula in cell B10",
863diff --git a/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py b/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py
864deleted file mode 100644
865index e69de29bb2d1d6434b8b29ae775ad8c2e48c5391..0000000000000000000000000000000000000000
866--- a/dot_config/opencode/skills/skill-creator/scripts/empty___init__.py
867+++ /dev/null
868diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py b/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py
869deleted file mode 100644
870index 3e66e8c105be9bab9f0e9c61f0d1482619401580..0000000000000000000000000000000000000000
871--- a/dot_config/opencode/skills/skill-creator/scripts/executable_aggregate_benchmark.py
872+++ /dev/null
873@@ -1,401 +0,0 @@
874-#!/usr/bin/env python3
875-"""
876-Aggregate individual run results into benchmark summary statistics.
877-
878-Reads grading.json files from run directories and produces:
879-- run_summary with mean, stddev, min, max for each metric
880-- delta between with_skill and without_skill configurations
881-
882-Usage:
883- python aggregate_benchmark.py <benchmark_dir>
884-
885-Example:
886- python aggregate_benchmark.py benchmarks/2026-01-15T10-30-00/
887-
888-The script supports two directory layouts:
889-
890- Workspace layout (from skill-creator iterations):
891- <benchmark_dir>/
892- └── eval-N/
893- ├── with_skill/
894- │ ├── run-1/grading.json
895- │ └── run-2/grading.json
896- └── without_skill/
897- ├── run-1/grading.json
898- └── run-2/grading.json
899-
900- Legacy layout (with runs/ subdirectory):
901- <benchmark_dir>/
902- └── runs/
903- └── eval-N/
904- ├── with_skill/
905- │ └── run-1/grading.json
906- └── without_skill/
907- └── run-1/grading.json
908-"""
909-
910-import argparse
911-import json
912-import math
913-import sys
914-from datetime import datetime, timezone
915-from pathlib import Path
916-
917-
918-def calculate_stats(values: list[float]) -> dict:
919- """Calculate mean, stddev, min, max for a list of values."""
920- if not values:
921- return {"mean": 0.0, "stddev": 0.0, "min": 0.0, "max": 0.0}
922-
923- n = len(values)
924- mean = sum(values) / n
925-
926- if n > 1:
927- variance = sum((x - mean) ** 2 for x in values) / (n - 1)
928- stddev = math.sqrt(variance)
929- else:
930- stddev = 0.0
931-
932- return {
933- "mean": round(mean, 4),
934- "stddev": round(stddev, 4),
935- "min": round(min(values), 4),
936- "max": round(max(values), 4)
937- }
938-
939-
940-def load_run_results(benchmark_dir: Path) -> dict:
941- """
942- Load all run results from a benchmark directory.
943-
944- Returns dict keyed by config name (e.g. "with_skill"/"without_skill",
945- or "new_skill"/"old_skill"), each containing a list of run results.
946- """
947- # Support both layouts: eval dirs directly under benchmark_dir, or under runs/
948- runs_dir = benchmark_dir / "runs"
949- if runs_dir.exists():
950- search_dir = runs_dir
951- elif list(benchmark_dir.glob("eval-*")):
952- search_dir = benchmark_dir
953- else:
954- print(f"No eval directories found in {benchmark_dir} or {benchmark_dir / 'runs'}")
955- return {}
956-
957- results: dict[str, list] = {}
958-
959- for eval_idx, eval_dir in enumerate(sorted(search_dir.glob("eval-*"))):
960- metadata_path = eval_dir / "eval_metadata.json"
961- if metadata_path.exists():
962- try:
963- with open(metadata_path) as mf:
964- eval_id = json.load(mf).get("eval_id", eval_idx)
965- except (json.JSONDecodeError, OSError):
966- eval_id = eval_idx
967- else:
968- try:
969- eval_id = int(eval_dir.name.split("-")[1])
970- except ValueError:
971- eval_id = eval_idx
972-
973diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py b/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py
974deleted file mode 100644
975index 959e30a0014ec165c41a2bb7420b7dfe1416bbac..0000000000000000000000000000000000000000
976--- a/dot_config/opencode/skills/skill-creator/scripts/executable_generate_report.py
977+++ /dev/null
978@@ -1,326 +0,0 @@
979-#!/usr/bin/env python3
980-"""Generate an HTML report from run_loop.py output.
981-
982-Takes the JSON output from run_loop.py and generates a visual HTML report
983-showing each description attempt with check/x for each test case.
984-Distinguishes between train and test queries.
985-"""
986-
987-import argparse
988-import html
989-import json
990-import sys
991-from pathlib import Path
992-
993-
994-def generate_html(data: dict, auto_refresh: bool = False, skill_name: str = "") -> str:
995- """Generate HTML report from loop output data. If auto_refresh is True, adds a meta refresh tag."""
996- history = data.get("history", [])
997- holdout = data.get("holdout", 0)
998- title_prefix = html.escape(skill_name + " \u2014 ") if skill_name else ""
999-
1000- # Get all unique queries from train and test sets, with should_trigger info
1001- train_queries: list[dict] = []
1002- test_queries: list[dict] = []
1003- if history:
1004- for r in history[0].get("train_results", history[0].get("results", [])):
1005- train_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
1006- if history[0].get("test_results"):
1007- for r in history[0].get("test_results", []):
1008- test_queries.append({"query": r["query"], "should_trigger": r.get("should_trigger", True)})
1009-
1010- refresh_tag = ' <meta http-equiv="refresh" content="5">\n' if auto_refresh else ""
1011-
1012- html_parts = ["""<!DOCTYPE html>
1013-<html>
1014-<head>
1015- <meta charset="utf-8">
1016-""" + refresh_tag + """ <title>""" + title_prefix + """Skill Description Optimization</title>
1017- <link rel="preconnect" href="https://fonts.googleapis.com">
1018- <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
1019- <link href="https://fonts.googleapis.com/css2?family=Poppins:wght@500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
1020- <style>
1021- body {
1022- font-family: 'Lora', Georgia, serif;
1023- max-width: 100%;
1024- margin: 0 auto;
1025- padding: 20px;
1026- background: #faf9f5;
1027- color: #141413;
1028- }
1029- h1 { font-family: 'Poppins', sans-serif; color: #141413; }
1030- .explainer {
1031- background: white;
1032- padding: 15px;
1033- border-radius: 6px;
1034- margin-bottom: 20px;
1035- border: 1px solid #e8e6dc;
1036- color: #b0aea5;
1037- font-size: 0.875rem;
1038- line-height: 1.6;
1039- }
1040- .summary {
1041- background: white;
1042- padding: 15px;
1043- border-radius: 6px;
1044- margin-bottom: 20px;
1045- border: 1px solid #e8e6dc;
1046- }
1047- .summary p { margin: 5px 0; }
1048- .best { color: #788c5d; font-weight: bold; }
1049- .table-container {
1050- overflow-x: auto;
1051- width: 100%;
1052- }
1053- table {
1054- border-collapse: collapse;
1055- background: white;
1056- border: 1px solid #e8e6dc;
1057- border-radius: 6px;
1058- font-size: 12px;
1059- min-width: 100%;
1060- }
1061- th, td {
1062- padding: 8px;
1063- text-align: left;
1064- border: 1px solid #e8e6dc;
1065- white-space: normal;
1066- word-wrap: break-word;
1067- }
1068- th {
1069- font-family: 'Poppins', sans-serif;
1070- background: #141413;
1071- color: #faf9f5;
1072- font-weight: 500;
1073- }
1074- th.test-col {
1075- background: #6a9bcc;
1076- }
1077- th.query-col { min-width: 200px; }
1078diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py b/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py
1079deleted file mode 100644
1080index 12bbdb8635073a7f84f30741fc9795ef7952f9c5..0000000000000000000000000000000000000000
1081--- a/dot_config/opencode/skills/skill-creator/scripts/executable_improve_description.py
1082+++ /dev/null
1083@@ -1,261 +0,0 @@
1084-#!/usr/bin/env python3
1085-"""Improve a skill description based on eval results.
1086-
1087-Takes eval results (from run_eval.py) and generates an improved description
1088-using opencode with haiku 4.5.
1089-"""
1090-
1091-import argparse
1092-import json
1093-import re
1094-import subprocess
1095-import sys
1096-from pathlib import Path
1097-
1098-from scripts.utils import parse_skill_md
1099-
1100-IMPROVE_MODEL = "anthropic/claude-haiku-4-5"
1101-
1102-
1103-def _run_opencode(prompt: str, model: str) -> str:
1104- """Run opencode run with a prompt and return the assistant text response."""
1105- env_clean = {
1106- k: v
1107- for k, v in __import__("os").environ.items()
1108- if k not in ("OPENCODE", "OPENCODE_PID")
1109- }
1110- result = subprocess.run(
1111- ["opencode", "run", "--format", "json", "--model", model, prompt],
1112- capture_output=True,
1113- text=True,
1114- env=env_clean,
1115- )
1116- # Parse JSON event stream, collect text parts
1117- text_parts = []
1118- for line in result.stdout.splitlines():
1119- line = line.strip()
1120- if not line:
1121- continue
1122- try:
1123- event = json.loads(line)
1124- except json.JSONDecodeError:
1125- continue
1126- if event.get("type") == "text":
1127- part = event.get("part", {})
1128- text_parts.append(part.get("text", ""))
1129- return "".join(text_parts)
1130-
1131-
1132-def improve_description(
1133- skill_name: str,
1134- skill_content: str,
1135- current_description: str,
1136- eval_results: dict,
1137- history: list[dict],
1138- model: str = IMPROVE_MODEL,
1139- test_results: dict | None = None,
1140- log_dir: Path | None = None,
1141- iteration: int | None = None,
1142-) -> str:
1143- """Call opencode to improve the description based on eval results."""
1144- failed_triggers = [
1145- r for r in eval_results["results"] if r["should_trigger"] and not r["pass"]
1146- ]
1147- false_triggers = [
1148- r for r in eval_results["results"] if not r["should_trigger"] and not r["pass"]
1149- ]
1150-
1151- # Build scores summary
1152- train_score = (
1153- f"{eval_results['summary']['passed']}/{eval_results['summary']['total']}"
1154- )
1155- if test_results:
1156- test_score = (
1157- f"{test_results['summary']['passed']}/{test_results['summary']['total']}"
1158- )
1159- scores_summary = f"Train: {train_score}, Test: {test_score}"
1160- else:
1161- scores_summary = f"Train: {train_score}"
1162-
1163diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py
1164deleted file mode 100644
1165index 4a6fe2624c5c2c3e723e9ad58005fd9e361e9499..0000000000000000000000000000000000000000
1166--- a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_eval.py
1167+++ /dev/null
1168@@ -1,304 +0,0 @@
1169-#!/usr/bin/env python3
1170-"""Run trigger evaluation for a skill description.
1171-
1172-Tests whether a skill's description causes opencode to trigger (read the skill)
1173-for a set of queries. Outputs results as JSON.
1174-"""
1175-
1176-import argparse
1177-import json
1178-import os
1179-import select
1180-import subprocess
1181-import sys
1182-import time
1183-import uuid
1184-from concurrent.futures import ProcessPoolExecutor, as_completed
1185-from pathlib import Path
1186-
1187-from scripts.utils import parse_skill_md
1188-
1189-EVAL_MODEL = "anthropic/claude-sonnet-4-6"
1190-
1191-
1192-def find_project_root() -> Path:
1193- """Find the project root by walking up from cwd looking for .opencode/.
1194-
1195- Mimics how opencode discovers its project root, so the command file
1196- we create ends up where opencode will look for it.
1197- """
1198- current = Path.cwd()
1199- for parent in [current, *current.parents]:
1200- if (parent / ".opencode").is_dir():
1201- return parent
1202- return current
1203-
1204-
1205-def run_single_query(
1206- query: str,
1207- skill_name: str,
1208- skill_description: str,
1209- timeout: int,
1210- project_root: str,
1211- model: str | None = None,
1212-) -> bool:
1213- """Run a single query and return whether the skill was triggered.
1214-
1215- Creates a skill file in .opencode/skills/ so it appears in opencode's
1216- available_skills list, then runs `opencode run` with the raw query.
1217- Parses JSON events to detect whether the skill tool was invoked.
1218- """
1219- unique_id = uuid.uuid4().hex[:8]
1220- clean_name = f"{skill_name}-skill-{unique_id}"
1221- project_skills_dir = Path(project_root) / ".opencode" / "skills" / clean_name
1222- skill_file = project_skills_dir / "SKILL.md"
1223-
1224- try:
1225- project_skills_dir.mkdir(parents=True, exist_ok=True)
1226- # Use YAML block scalar to avoid breaking on quotes in description
1227- indented_desc = "\n ".join(skill_description.split("\n"))
1228- skill_content = (
1229- f"---\n"
1230- f"name: {clean_name}\n"
1231- f"description: |\n"
1232- f" {indented_desc}\n"
1233- f"---\n\n"
1234- f"# {skill_name}\n\n"
1235- f"This skill handles: {skill_description}\n"
1236- )
1237- skill_file.write_text(skill_content)
1238-
1239- cmd = [
1240- "opencode",
1241- "run",
1242- "--format",
1243- "json",
1244- "--model",
1245- model or EVAL_MODEL,
1246- query,
1247- ]
1248-
1249- # Remove OPENCODE env var to allow nesting opencode run inside an
1250- # opencode session. The guard is for interactive terminal conflicts;
1251- # programmatic subprocess usage is safe.
1252- env = {
1253- k: v for k, v in os.environ.items() if k not in ("OPENCODE", "OPENCODE_PID")
1254- }
1255-
1256- process = subprocess.Popen(
1257- cmd,
1258- stdout=subprocess.PIPE,
1259- stderr=subprocess.DEVNULL,
1260- cwd=project_root,
1261- env=env,
1262- )
1263-
1264- triggered = False
1265- start_time = time.time()
1266- buffer = ""
1267-
1268diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py b/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py
1269deleted file mode 100644
1270index a18ba9b7544d573371420e31f0070452c324d75e..0000000000000000000000000000000000000000
1271--- a/dot_config/opencode/skills/skill-creator/scripts/executable_literal_run_loop.py
1272+++ /dev/null
1273@@ -1,404 +0,0 @@
1274-#!/usr/bin/env python3
1275-"""Run the eval + improve loop until all pass or max iterations reached.
1276-
1277-Combines run_eval.py and improve_description.py in a loop, tracking history
1278-and returning the best description found. Supports train/test split to prevent
1279-overfitting.
1280-"""
1281-
1282-import argparse
1283-import json
1284-import random
1285-import sys
1286-import tempfile
1287-import time
1288-import webbrowser
1289-from pathlib import Path
1290-
1291-from scripts.generate_report import generate_html
1292-from scripts.improve_description import improve_description, IMPROVE_MODEL
1293-from scripts.run_eval import find_project_root, run_eval, EVAL_MODEL
1294-from scripts.utils import parse_skill_md
1295-
1296-
1297-def split_eval_set(
1298- eval_set: list[dict], holdout: float, seed: int = 42
1299-) -> tuple[list[dict], list[dict]]:
1300- """Split eval set into train and test sets, stratified by should_trigger."""
1301- random.seed(seed)
1302-
1303- # Separate by should_trigger
1304- trigger = [e for e in eval_set if e["should_trigger"]]
1305- no_trigger = [e for e in eval_set if not e["should_trigger"]]
1306-
1307- # Shuffle each group
1308- random.shuffle(trigger)
1309- random.shuffle(no_trigger)
1310-
1311- # Calculate split points
1312- n_trigger_test = max(1, int(len(trigger) * holdout))
1313- n_no_trigger_test = max(1, int(len(no_trigger) * holdout))
1314-
1315- # Split
1316- test_set = trigger[:n_trigger_test] + no_trigger[:n_no_trigger_test]
1317- train_set = trigger[n_trigger_test:] + no_trigger[n_no_trigger_test:]
1318-
1319- return train_set, test_set
1320-
1321-
1322-def run_loop(
1323- eval_set: list[dict],
1324- skill_path: Path,
1325- description_override: str | None,
1326- num_workers: int,
1327- timeout: int,
1328- max_iterations: int,
1329- runs_per_query: int,
1330- trigger_threshold: float,
1331- holdout: float,
1332- model: str,
1333- verbose: bool,
1334- live_report_path: Path | None = None,
1335- log_dir: Path | None = None,
1336-) -> dict:
1337- """Run the eval + improvement loop."""
1338- project_root = find_project_root()
1339- name, original_description, content = parse_skill_md(skill_path)
1340- current_description = description_override or original_description
1341-
1342- # Split into train/test if holdout > 0
1343- if holdout > 0:
1344- train_set, test_set = split_eval_set(eval_set, holdout)
1345- if verbose:
1346- print(
1347- f"Split: {len(train_set)} train, {len(test_set)} test (holdout={holdout})",
1348- file=sys.stderr,
1349- )
1350- else:
1351- train_set = eval_set
1352- test_set = []
1353-
1354- history = []
1355- exit_reason = "unknown"
1356-
1357- for iteration in range(1, max_iterations + 1):
1358- if verbose:
1359- print(f"\n{'=' * 60}", file=sys.stderr)
1360- print(f"Iteration {iteration}/{max_iterations}", file=sys.stderr)
1361- print(f"Description: {current_description}", file=sys.stderr)
1362- print(f"{'=' * 60}", file=sys.stderr)
1363-
1364- # Evaluate train + test together in one batch for parallelism
1365- all_queries = train_set + test_set
1366- t0 = time.time()
1367- all_results = run_eval(
1368- eval_set=all_queries,
1369- skill_name=name,
1370- description=current_description,
1371- num_workers=num_workers,
1372- timeout=timeout,
1373diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py b/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py
1374deleted file mode 100644
1375index f48eac444656ddc41204aac1760a217951ce609e..0000000000000000000000000000000000000000
1376--- a/dot_config/opencode/skills/skill-creator/scripts/executable_package_skill.py
1377+++ /dev/null
1378@@ -1,136 +0,0 @@
1379-#!/usr/bin/env python3
1380-"""
1381-Skill Packager - Creates a distributable .skill file of a skill folder
1382-
1383-Usage:
1384- python utils/package_skill.py <path/to/skill-folder> [output-directory]
1385-
1386-Example:
1387- python utils/package_skill.py skills/public/my-skill
1388- python utils/package_skill.py skills/public/my-skill ./dist
1389-"""
1390-
1391-import fnmatch
1392-import sys
1393-import zipfile
1394-from pathlib import Path
1395-from scripts.quick_validate import validate_skill
1396-
1397-# Patterns to exclude when packaging skills.
1398-EXCLUDE_DIRS = {"__pycache__", "node_modules"}
1399-EXCLUDE_GLOBS = {"*.pyc"}
1400-EXCLUDE_FILES = {".DS_Store"}
1401-# Directories excluded only at the skill root (not when nested deeper).
1402-ROOT_EXCLUDE_DIRS = {"evals"}
1403-
1404-
1405-def should_exclude(rel_path: Path) -> bool:
1406- """Check if a path should be excluded from packaging."""
1407- parts = rel_path.parts
1408- if any(part in EXCLUDE_DIRS for part in parts):
1409- return True
1410- # rel_path is relative to skill_path.parent, so parts[0] is the skill
1411- # folder name and parts[1] (if present) is the first subdir.
1412- if len(parts) > 1 and parts[1] in ROOT_EXCLUDE_DIRS:
1413- return True
1414- name = rel_path.name
1415- if name in EXCLUDE_FILES:
1416- return True
1417- return any(fnmatch.fnmatch(name, pat) for pat in EXCLUDE_GLOBS)
1418-
1419-
1420-def package_skill(skill_path, output_dir=None):
1421- """
1422- Package a skill folder into a .skill file.
1423-
1424- Args:
1425- skill_path: Path to the skill folder
1426- output_dir: Optional output directory for the .skill file (defaults to current directory)
1427-
1428- Returns:
1429- Path to the created .skill file, or None if error
1430- """
1431- skill_path = Path(skill_path).resolve()
1432-
1433- # Validate skill folder exists
1434- if not skill_path.exists():
1435- print(f"❌ Error: Skill folder not found: {skill_path}")
1436- return None
1437-
1438- if not skill_path.is_dir():
1439- print(f"❌ Error: Path is not a directory: {skill_path}")
1440- return None
1441-
1442- # Validate SKILL.md exists
1443- skill_md = skill_path / "SKILL.md"
1444- if not skill_md.exists():
1445- print(f"❌ Error: SKILL.md not found in {skill_path}")
1446- return None
1447-
1448- # Run validation before packaging
1449- print("🔍 Validating skill...")
1450- valid, message = validate_skill(skill_path)
1451- if not valid:
1452- print(f"❌ Validation failed: {message}")
1453- print(" Please fix the validation errors before packaging.")
1454- return None
1455- print(f"✅ {message}\n")
1456-
1457- # Determine output location
1458- skill_name = skill_path.name
1459- if output_dir:
1460- output_path = Path(output_dir).resolve()
1461- output_path.mkdir(parents=True, exist_ok=True)
1462- else:
1463- output_path = Path.cwd()
1464-
1465- skill_filename = output_path / f"{skill_name}.skill"
1466-
1467- # Create the .skill file (zip format)
1468- try:
1469- with zipfile.ZipFile(skill_filename, 'w', zipfile.ZIP_DEFLATED) as zipf:
1470- # Walk through the skill directory, excluding build artifacts
1471- for file_path in skill_path.rglob('*'):
1472- if not file_path.is_file():
1473- continue
1474- arcname = file_path.relative_to(skill_path.parent)
1475- if should_exclude(arcname):
1476- print(f" Skipped: {arcname}")
1477- continue
1478diff --git a/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py b/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py
1479deleted file mode 100644
1480index ed8e1dddce77b16af13c6f36b3fe86c4ac7c590c..0000000000000000000000000000000000000000
1481--- a/dot_config/opencode/skills/skill-creator/scripts/executable_quick_validate.py
1482+++ /dev/null
1483@@ -1,103 +0,0 @@
1484-#!/usr/bin/env python3
1485-"""
1486-Quick validation script for skills - minimal version
1487-"""
1488-
1489-import sys
1490-import os
1491-import re
1492-import yaml
1493-from pathlib import Path
1494-
1495-def validate_skill(skill_path):
1496- """Basic validation of a skill"""
1497- skill_path = Path(skill_path)
1498-
1499- # Check SKILL.md exists
1500- skill_md = skill_path / 'SKILL.md'
1501- if not skill_md.exists():
1502- return False, "SKILL.md not found"
1503-
1504- # Read and validate frontmatter
1505- content = skill_md.read_text()
1506- if not content.startswith('---'):
1507- return False, "No YAML frontmatter found"
1508-
1509- # Extract frontmatter
1510- match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
1511- if not match:
1512- return False, "Invalid frontmatter format"
1513-
1514- frontmatter_text = match.group(1)
1515-
1516- # Parse YAML frontmatter
1517- try:
1518- frontmatter = yaml.safe_load(frontmatter_text)
1519- if not isinstance(frontmatter, dict):
1520- return False, "Frontmatter must be a YAML dictionary"
1521- except yaml.YAMLError as e:
1522- return False, f"Invalid YAML in frontmatter: {e}"
1523-
1524- # Define allowed properties
1525- ALLOWED_PROPERTIES = {'name', 'description', 'license', 'allowed-tools', 'metadata', 'compatibility'}
1526-
1527- # Check for unexpected properties (excluding nested keys under metadata)
1528- unexpected_keys = set(frontmatter.keys()) - ALLOWED_PROPERTIES
1529- if unexpected_keys:
1530- return False, (
1531- f"Unexpected key(s) in SKILL.md frontmatter: {', '.join(sorted(unexpected_keys))}. "
1532- f"Allowed properties are: {', '.join(sorted(ALLOWED_PROPERTIES))}"
1533- )
1534-
1535- # Check required fields
1536- if 'name' not in frontmatter:
1537- return False, "Missing 'name' in frontmatter"
1538- if 'description' not in frontmatter:
1539- return False, "Missing 'description' in frontmatter"
1540-
1541- # Extract name for validation
1542- name = frontmatter.get('name', '')
1543- if not isinstance(name, str):
1544- return False, f"Name must be a string, got {type(name).__name__}"
1545- name = name.strip()
1546- if name:
1547- # Check naming convention (kebab-case: lowercase with hyphens)
1548- if not re.match(r'^[a-z0-9-]+$', name):
1549- return False, f"Name '{name}' should be kebab-case (lowercase letters, digits, and hyphens only)"
1550- if name.startswith('-') or name.endswith('-') or '--' in name:
1551- return False, f"Name '{name}' cannot start/end with hyphen or contain consecutive hyphens"
1552- # Check name length (max 64 characters per spec)
1553- if len(name) > 64:
1554- return False, f"Name is too long ({len(name)} characters). Maximum is 64 characters."
1555-
1556- # Extract and validate description
1557- description = frontmatter.get('description', '')
1558- if not isinstance(description, str):
1559- return False, f"Description must be a string, got {type(description).__name__}"
1560- description = description.strip()
1561- if description:
1562- # Check for angle brackets
1563- if '<' in description or '>' in description:
1564- return False, "Description cannot contain angle brackets (< or >)"
1565- # Check description length (max 1024 characters per spec)
1566- if len(description) > 1024:
1567- return False, f"Description is too long ({len(description)} characters). Maximum is 1024 characters."
1568-
1569- # Validate compatibility field if present (optional)
1570- compatibility = frontmatter.get('compatibility', '')
1571- if compatibility:
1572- if not isinstance(compatibility, str):
1573- return False, f"Compatibility must be a string, got {type(compatibility).__name__}"
1574- if len(compatibility) > 500:
1575- return False, f"Compatibility is too long ({len(compatibility)} characters). Maximum is 500 characters."
1576-
1577- return True, "Skill is valid!"
1578-
1579-if __name__ == "__main__":
1580- if len(sys.argv) != 2:
1581- print("Usage: python quick_validate.py <skill_directory>")
1582- sys.exit(1)
1583diff --git a/dot_config/opencode/skills/skill-creator/scripts/utils.py b/dot_config/opencode/skills/skill-creator/scripts/utils.py
1584deleted file mode 100644
1585index 51b6a07dd57174197a937034b7eecebd5768ff8a..0000000000000000000000000000000000000000
1586--- a/dot_config/opencode/skills/skill-creator/scripts/utils.py
1587+++ /dev/null
1588@@ -1,47 +0,0 @@
1589-"""Shared utilities for skill-creator scripts."""
1590-
1591-from pathlib import Path
1592-
1593-
1594-
1595-def parse_skill_md(skill_path: Path) -> tuple[str, str, str]:
1596- """Parse a SKILL.md file, returning (name, description, full_content)."""
1597- content = (skill_path / "SKILL.md").read_text()
1598- lines = content.split("\n")
1599-
1600- if lines[0].strip() != "---":
1601- raise ValueError("SKILL.md missing frontmatter (no opening ---)")
1602-
1603- end_idx = None
1604- for i, line in enumerate(lines[1:], start=1):
1605- if line.strip() == "---":
1606- end_idx = i
1607- break
1608-
1609- if end_idx is None:
1610- raise ValueError("SKILL.md missing frontmatter (no closing ---)")
1611-
1612- name = ""
1613- description = ""
1614- frontmatter_lines = lines[1:end_idx]
1615- i = 0
1616- while i < len(frontmatter_lines):
1617- line = frontmatter_lines[i]
1618- if line.startswith("name:"):
1619- name = line[len("name:"):].strip().strip('"').strip("'")
1620- elif line.startswith("description:"):
1621- value = line[len("description:"):].strip()
1622- # Handle YAML multiline indicators (>, |, >-, |-)
1623- if value in (">", "|", ">-", "|-"):
1624- continuation_lines: list[str] = []
1625- i += 1
1626- while i < len(frontmatter_lines) and (frontmatter_lines[i].startswith(" ") or frontmatter_lines[i].startswith("\t")):
1627- continuation_lines.append(frontmatter_lines[i].strip())
1628- i += 1
1629- description = " ".join(continuation_lines)
1630- continue
1631- else:
1632- description = value.strip('"').strip("'")
1633- i += 1
1634-
1635- return name, description, content