Few-shot skill induction from user examples.
Core question:
Given only several user examples, can we induce a reusable skill module that improves heldout tasks compared with prompt-only, examples-only, and one-shot skill drafting?
Current project scope has two tracks:
- Data cleaning scripts: normalize WritingBench / PresentBench into reproducible train-example and heldout-task splits.
- Auto-skills module: compile user examples into a structured skill package with feature extraction, cross-example aggregation, validation notes, and generated
SKILL.md.
The auto-skill module must treat user examples as the only induction input. Benchmark rubrics/checklists may be used only for private data construction checks and heldout evaluation. They must never be normal skill-induction input; any future oracle or upper-bound experiment must be named and isolated separately.
src/auto_skill/ Core library code.
scripts/ CLI entrypoints; see scripts/README.md for inventory and lifecycle.
tests/ Minimal regression tests for shared logic and scripts.
notes/ Research notes and design sketches.
docs/ Stable repo/project docs for collaborators.
artifacts/ Small derived samples checked in while prototyping.
data/ Local benchmark checkouts; ignored.
- WritingBench:
../WritingBenchin the current local workspace, ordata/WritingBenchif you prefer a repo-local checkout. - PresentBench:
data/PresentBench_repo
data/ is ignored because PresentBench is large. Recreate it with:
git clone https://huggingface.co/datasets/PresentBench/PresentBench data/PresentBench_repo
git -C data/PresentBench_repo lfs install --local
git -C data/PresentBench_repo lfs pull
git -C data/PresentBench_repo lfs checkout
git clone --depth 1 --filter=blob:none --sparse https://github.com/PresentBench/PresentBench.git data/PresentBench_code
git -C data/PresentBench_code sparse-checkout set --skip-checks README.md judge.py judge_all.py scoring.py scoring_all.py utilsCurrent artifacts were built from a sibling WritingBench checkout:
git clone https://github.com/X-PLUG/WritingBench.git ../WritingBenchIf you keep WritingBench under data/, pass --writingbench-root data/WritingBench
when rebuilding splits.
Repo-local alternative:
git clone https://github.com/X-PLUG/WritingBench.git data/WritingBenchuv sync --extra dev
cp .env.example .envOfficial PresentBench judging also needs the upstream judge runtime:
uv sync --extra dev --extra presentbenchFill .env with the Bailian OpenAI-compatible endpoint, key, and model:
BAILIAN_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
BAILIAN_API_KEY=replace-with-your-api-key
BAILIAN_MODEL=qwen3.5-plus
# Optional: BAILIAN_TEMPERATURE=0.2
# Optional: BAILIAN_NUM_THREADS=16
# Optional: BAILIAN_STREAM=trueFor MIMO judge-swap or targeted MIMO example cleaning, add the corresponding OpenAI-compatible MIMO variables:
MIMO_BASE_URL=replace-with-mimo-compatible-endpoint
MIMO_API_KEY=replace-with-mimo-key
MIMO_MODEL=replace-with-mimo-modelFor the upstream PresentBench official judge, add a Gemini key. This is separate
from Bailian/MIMO because the official PresentBench judge.py currently accepts
only gemini / gemini_inline API types:
GENAI_API_KEY=replace-with-gemini-key
# Optional: GENAI_BASE_URL=replace-with-gemini-compatible-endpointThe minimal test path uses only the Python standard library:
uv run python -m unittest discover -s testsInspect benchmark structure:
uv run python scripts/data/inspect_benchmarks.py \
--writingbench-root ../WritingBench \
--presentbench-root data/PresentBench_repo \
--limit 2 \
--out-dir artifactsBuild few-shot splits:
uv run python scripts/data/build_fewshot_splits.py \
--writingbench-root ../WritingBench \
--presentbench-root data/PresentBench_repo \
--train-size 3 \
--heldout-size 2 \
--max-groups 4 \
--out artifacts/splits/fewshot_splits.jsonl \
--summary-out artifacts/splits/fewshot_split_summary.jsonFor the next non-MVP expansion, prefer a stratified medium run rather than a full benchmark clean:
uv run python scripts/data/build_fewshot_splits.py \
--writingbench-root ../WritingBench \
--presentbench-root data/PresentBench_repo \
--train-size 3 \
--heldout-size 2 \
--max-writing-groups 30 \
--max-present-groups 20 \
--out runs/expanded/fewshot_splits.30wb_20pb.jsonl \
--summary-out runs/expanded/fewshot_split_summary.30wb_20pb.jsonOnly move to full cleaning after this expanded split passes
validate_splits.py, audit_benchmark_flow.py, and a sampled example-quality
audit with acceptable failure rate.
Current expanded-run evidence is summarized in
notes/expanded_cleaning_strategy_2026-05-08.md; artifact hashes and handoff
commands are in docs/expanded_cleaning_manifest_2026-05-08.md. The short
version: Qwen3.5-Plus completed both the local 30WB/20PB clean and the
max-available clean. The max-available snapshot has 142 packs, 426 frozen train
examples, 284 heldout tasks, 426/426 latest generation success, and a passing
benchmark-flow audit. MIMO has a local audited subset pack with 15 packs / 45 frozen train examples
(runs/expanded/example_packs.30wb_20pb.mimo.sample.v1.jsonl,
benchmark-flow status=ok), but it is not a full MIMO-cleaned alternative.
To re-check the local expanded-cleaning handoff state:
uv run python scripts/ops/report_expanded_cleaning_status.py --expect-status ready
uv run python scripts/ops/report_expanded_cleaning_status.py --require-mimo-subset --expect-status ready
uv run python scripts/data/export_latest_successful_generations.py \
--generations runs/expanded/generated_desired_outputs.max_available.qwen.jsonl \
--out runs/expanded/generated_desired_outputs.max_available.qwen.latest_success.jsonl \
--expect-successes 426
uv run python scripts/ops/report_expanded_cleaning_status.py \
--splits runs/expanded/fewshot_splits.max_available.jsonl \
--packs runs/expanded/example_packs.max_available.qwen.v1.jsonl \
--private-eval runs/expanded/example_private_eval.max_available.jsonl \
--jobs runs/expanded/example_generation_jobs.max_available.jsonl \
--generated-outputs runs/expanded/generated_desired_outputs.max_available.qwen.latest_success.jsonl \
--expect-packs 142 \
--expect-train-examples 426 \
--expect-heldout-tasks 284 \
--expect-generation-jobs 426 \
--skip-mimo-subset \
--expect-status ready
uv run python scripts/metrics/validate_disagreement_taxonomy.py \
--packets runs/expanded/judge_disagreements.qwen_vs_mimo.sample4_wb.heldout1.jsonl \
--taxonomy notes/judge_disagreement_taxonomy_2026-05-09.jsonl \
--expect-status okThe taxonomy validator is a calibration handoff gate for the expanded Qwen-vs-MIMO WritingBench sample. It verifies that every sign-flip packet has exactly one provisional label and that the labels are tied to the candidate and baseline output hashes.
The split builder excludes known source-data mismatches, currently
education/CSAPP-Lectures_2015Fall/Lecture15, whose instructions ask for
Chapter 10 System-Level I/O while its judge checklist scores Chapter 8
Exceptional Control Flow.
It also writes a machine-readable selection summary with the total group
denominator, eligible groups, skipped-too-small groups, and max-group truncation.
Validate split schema before pushing:
uv run python scripts/data/validate_splits.py artifacts/splits/fewshot_splits.jsonlConstruct clean example packs and private generation/eval artifacts:
uv run python scripts/data/build_example_packs.py \
--splits artifacts/splits/fewshot_splits.jsonl \
--packs-out artifacts/packs/example_packs.jsonl \
--private-out artifacts/private/example_private_eval.jsonl \
--jobs-out artifacts/jobs/example_generation_jobs.jsonl \
--summary-out artifacts/reports/example_pack_summary.mdThis writes the local construction layer:
artifacts/packs/example_packs.jsonl: intermediate clean user-visible packs before desired outputs are applied; ignored by git.artifacts/jobs/example_generation_jobs.jsonl: private jobs for generating desired outputs/examples; ignored by git.artifacts/private/example_private_eval.jsonl: private rubric/judge/checklist metadata for construction and evaluation only.artifacts/reports/example_pack_summary.md: compact dataset summary; ignored by git.
Run desired-output generation jobs through Bailian:
uv run python scripts/data/run_generation_jobs.py \
--jobs artifacts/jobs/example_generation_jobs.jsonl \
--out artifacts/jobs/generated_desired_outputs.jsonl \
--config-prefix MIMO \
--limit 2 \
--num-threads 16 \
--timeout-seconds 600 \
--max-retries 0Use --config-prefix MIMO or --config-prefix QWEN when constructing a
parallel set of examples from another OpenAI-compatible model configured in
.env (MIMO_MODEL, MIMO_BASE_URL, MIMO_API_KEY, etc.). The default path
still uses the legacy BAILIAN_* / OPENAI_* chain.
The script prints per-job duration_seconds and total jobs/minute. Current
Qwen3.5-Plus PresentBench runs are latency-bound rather than RPM-bound; use the
timing plus observed 429/timeout rate to decide whether to increase toward
--num-threads 32.
For long sequential skill-induction runs, pass --stream to consume Qwen output
incrementally. --stream-log prints deltas to stderr and should only be used for
single-run debugging, not threaded batch generation.
The skill MVP runner writes an append-only stage ledger next to --out; after
an API failure or interruption, rerun with --resume to reuse successful stage
calls and completed skill rows. Do not globally enable Qwen thinking for skill
induction. Start with --no-enable-thinking; use
--thinking-stages ... --thinking-budget N only for targeted analysis/merge
A/B runs.
For long MVP experiments, use the thin Bash orchestrator rather than manually retyping every command:
RUN_SKILL_MVP=0 RUN_MEMORY=1 RUN_WRITINGBENCH=1 RUN_PRESENTBENCH=1 RUN_VALIDATE=1 \
PRESENTBENCH_NUM_THREADS=4 scripts/ops/run_mvp_pipeline.shIt sequences skill induction, extraction-memory update, WritingBench eval,
PresentBench surrogate eval, validation, timestamps, and per-step logs under
logs/mvp/. The Python CLIs remain the behavior source; the shell script only
manages long-running tasks.
PresentBench surrogate eval uses cell-level --num-threads: each
pack/task/mode cell runs independently, but generation then judge remains
sequential inside a cell and checkpoint writes stay in the main process.
Extraction memory is written to ignored local output
artifacts/memory/extraction_memory.v1.jsonl. It is limited to public
user-example-derived lessons from successful skill rows and must not contain
private rubrics, heldout feedback, judge traces, or scores.
Current stable MVP run:
uv run python scripts/skills/run_skill_mvp.py \
--packs artifacts/packs/example_packs.v1.jsonl \
--out runs/skill_mvp.qwen.mvp.jsonl \
--stream \
--resume \
--allow-partial \
--no-enable-thinking \
--no-leave-one-out \
--temperature 0.2 \
--json-temperature 0 \
--timeout-seconds 900 \
--max-retries 0 \
--max-tokens 8192Provider retries and parse retries are intentionally separate. --max-retries
covers SDK transport/rate-limit/server failures. --parse-max-attempts covers
complete model responses that intermittently return malformed JSON or invalid
judge scores. The default parse retry budget is 3 for run_skill_mvp.py,
run_writingbench_official_eval.py, run_heldout_eval.py,
audit_train_examples.py, run_grounding_eval.py,
run_pattern_similarity_eval.py, and run_self_consistency_metric.py.
Use --dry-run before spending API calls:
uv run python scripts/data/run_generation_jobs.py --limit 1 --dry-runApply generated outputs back to clean packs:
uv run python scripts/data/apply_generated_outputs.py \
--packs artifacts/packs/example_packs.jsonl \
--generations artifacts/jobs/generated_desired_outputs.jsonl \
--out artifacts/packs/example_packs.v1.jsonlAudit the frozen benchmark-flow boundary before using packs for induction:
uv run python scripts/data/audit_benchmark_flow.pyFor expanded cleaned-example artifacts, include the generation job/output provenance chain:
uv run python scripts/data/audit_benchmark_flow.py \
--splits runs/expanded/fewshot_splits.30wb_20pb.jsonl \
--packs runs/expanded/example_packs.30wb_20pb.qwen.v1.jsonl \
--private-eval runs/expanded/example_private_eval.30wb_20pb.jsonl \
--jobs runs/expanded/example_generation_jobs.30wb_20pb.jsonl \
--generated-outputs runs/expanded/generated_desired_outputs.30wb_20pb.qwen.jsonl
uv run python scripts/data/audit_benchmark_flow.py \
--splits runs/expanded/fewshot_splits.max_available.jsonl \
--packs runs/expanded/example_packs.max_available.qwen.v1.jsonl \
--private-eval runs/expanded/example_private_eval.max_available.jsonl \
--jobs runs/expanded/example_generation_jobs.max_available.jsonl \
--generated-outputs runs/expanded/generated_desired_outputs.max_available.qwen.latest_success.jsonlThe first stable experiment matrix should stay small:
prompt_onlyfew_shot_examples_onlyone_shot_skill_from_examplesours_no_validationours_full
teacher_refine_on_heldout is an upper bound, not a normal baseline, because it uses heldout-time feedback.
WritingBench and PresentBench private rubrics/checklists are used only at heldout evaluation time.
Avoid model monoculture in reported experiments. A Qwen-only run where the same model family cleans examples, induces skills, generates heldout outputs, and judges results is useful as an MVP smoke loop, but it is not strong evidence for the method. For paper-facing runs, separate at least the judge from the generation/induction model, and ideally report a small cross-model grid such as:
- data/example construction: Qwen3.5-Plus or a stronger independent model;
- induction/generation under test: Qwen, GPT, Claude, or other target models;
- evaluation judge: official benchmark evaluator when available, otherwise an independent judge model with judge-model sensitivity checks.
To split solver and judge in one run, set both BAILIAN_* (solver) and a second
prefix (e.g. JUDGE_*) in your .env, then pass --judge-config-prefix JUDGE
to the eval runner. Only JUDGE_MODEL is required; JUDGE_BASE_URL and
JUDGE_API_KEY fall back to BAILIAN_* when the prefixed values are absent.
# In .env (in addition to BAILIAN_*):
# JUDGE_BASE_URL=https://api.anthropic.com/v1
# JUDGE_API_KEY=...
# JUDGE_MODEL=claude-sonnet-4-6
uv run python scripts/eval/run_writingbench_official_eval.py \
--packs artifacts/packs/example_packs.v1.jsonl \
--skills runs/skill_mvp.qwen.mvp.jsonl \
--private-eval artifacts/private/example_private_eval.jsonl \
--writingbench-root ../WritingBench \
--judge-config-prefix JUDGE \
--out runs/writingbench_official_eval.qwen_solver.claude_judge.jsonl \
--summary-out runs/writingbench_official_eval.qwen_solver.claude_judge.summary.json \
--limit-heldout 1 --resumerun_heldout_eval.py, run_pattern_similarity_eval.py, and
run_self_consistency_metric.py accept the same --judge-config-prefix. Each
written eval row carries top-level solver_model and judge_model;
scripts/metrics/summarize_mvp_metrics.py aggregates these into a model_inventory
block. scripts/ops/report_experiment_readiness.py emits model_monoculture
when every visible row collapses to a single model, and also exposes
model_inventory.scored for paper-facing claims. The scored block excludes
non-success eval rows such as missing_score_artifact, so placeholder
PresentBench official rows cannot falsely break a Qwen-only scored run; in that
case the gate emits scored_model_monoculture. Metrics summaries keep
model_inventory focused on heldout eval rows and put self-consistency or other
diagnostic rows in diagnostic_model_inventory.
For WritingBench, use the source benchmark's evaluator prompt and per-criterion scoring shape:
uv run python scripts/eval/run_writingbench_official_eval.py \
--packs artifacts/packs/example_packs.v1.jsonl \
--skills runs/skill_mvp.qwen.mvp.jsonl \
--private-eval artifacts/private/example_private_eval.jsonl \
--writingbench-root ../WritingBench \
--pack-id writingbench_Finance_Business_Financial_Reports_en \
--limit-heldout 1 \
--resumeThis uses WritingBench's local prompt.py and Qwen as the judge model. It is the
right smoke path for this repo when we are intentionally evaluating with Qwen;
the official leaderboard currently recommends Claude or the WritingBench critic.
Both eval runners checkpoint rows to --out; --resume reuses successful cells
and retries non-success cells after interruptions or provider failures.
For PresentBench, the official evaluator is the upstream judge_all.py /
judge.py / scoring.py pipeline. It requires a local checkout of the official
code repo, benchmark materials, domain common judge prompts, judge weights, and
actual slides.pdf or slides.pptx artifacts under a result tree that mirrors
the benchmark data. It also requires the presentbench optional Python extra
because upstream imports Gemini, requests, and tqdm helpers. If generation
explicitly fails, upstream judge_all.py accepts slides_generation_failed.txt
and writes an all-zero score. The text-only
run_heldout_eval.py path is therefore only a Qwen
qwen_llm_rubric_surrogate smoke evaluator; it must not be reported as an
official PresentBench visual/PPT score. Check readiness before claiming an
official PresentBench run:
uv run python scripts/eval/export_presentbench_official_artifacts.py \
--eval runs/presentbench_surrogate_eval.qwen.mvp.jsonl \
--eval runs/presentbench_surrogate_eval.qwen.auto_skill.jsonl \
--mode-result-root prompt_only=../PresentBench/results/prompt_only \
--mode-result-root auto_skill=../PresentBench/results/auto_skill \
--overwriteuv run python scripts/eval/check_presentbench_official_eval_ready.py \
--packs artifacts/packs/example_packs.v1.jsonl \
--code-root data/PresentBench_code \
--mode-result-root prompt_only=../PresentBench/results/prompt_only \
--mode-result-root auto_skill=../PresentBench/results/auto_skill \
--judge-model gemini-3-flash-previewFor CI smoke checks where slide artifacts or generated PresentBench examples are
intentionally absent, add --allow-missing --allow-empty; do not use those flags
when claiming official readiness.
Readiness statuses are intentionally separated:
ready_for_official_judge: slides exist and upstreamjudge_all.pycan score them.ready_for_zero_score:slides_generation_failed.txtexists and upstream can write zero.scored: an upstream*_score.yamlis already present.
Once readiness is ready_for_official_judge or ready_for_zero_score, run the
repo-local wrapper. It loads .env, renders exact upstream judge.py commands
for the selected heldout cells in --packs, and fails before any API call when
the required Gemini key is missing. Use --all-presentbench only when the result
roots contain artifacts for the full upstream PresentBench checkout; otherwise it
would run judge_all.py over all 238 benchmark cases.
uv run python scripts/eval/run_presentbench_official_judge.py \
--code-root data/PresentBench_code \
--data-root data/PresentBench_repo \
--mode-result-root prompt_only=../PresentBench/results/prompt_only \
--mode-result-root auto_skill=../PresentBench/results/auto_skill \
--api-type gemini \
--model gemini-3-flash-preview \
--limit-heldout 2Use --dry-run --allow-missing-env to verify commands without GENAI_API_KEY.
Add --commands-out runs/presentbench_official_judge_commands.json to write a
handoff manifest containing argv lists, shell-safe commands, and dry-run
warnings such as a missing GENAI_API_KEY. Add --expect-commands 16 for the
current MVP official-score gap so the wrapper fails if the selected cell count
drifts. Do not use --allow-missing-env for a real run.
After upstream score YAMLs exist, summarize official PresentBench scores and paired deltas with explicit mode roots:
uv run python scripts/eval/summarize_presentbench_official_scores.py \
--packs artifacts/packs/example_packs.v1.jsonl \
--solver-model qwen3.5-plus \
--judge-model gemini-3-flash-preview \
--score-root prompt_only=../PresentBench/results/prompt_only \
--score-root auto_skill=../PresentBench/results/auto_skillGate the combined experiment state before reporting results:
uv run python scripts/ops/report_experiment_readiness.py --profile mvp --expect-status ready
uv run python scripts/metrics/validate_disagreement_taxonomy.py \
--packets runs/expanded/judge_disagreements.qwen_vs_mimo.sample4_wb.heldout1.jsonl \
--taxonomy notes/judge_disagreement_taxonomy_2026-05-09.jsonl \
--expect-status okBy default this uses --profile mvp, which expects the current MVP artifact
contract: skill_mvp.qwen.mvp.jsonl, 4-mode WritingBench eval, 4-mode
PresentBench surrogate eval, and no auto_skill_ours_full / official
PresentBench requirement. The command exits non-zero if the current MVP rows are
incomplete or contain non-success statuses; use --allow-not-ready when you only
want to write and inspect the report. Use explicit profiles when checking other
phases:
uv run python scripts/ops/report_experiment_readiness.py --profile smoke --limit-heldout 1
uv run python scripts/ops/report_experiment_readiness.py --profile full --expect-status not_readyAll profiles still fail on schema-invalid rows. smoke permits partial
old/small coverage, mvp requires current MVP coverage, and full requires
auto_skill_ours_full, auto_skill, and official PresentBench score rows.
--skills, --writing-eval, --present-surrogate-eval, and
--present-official-scores may be repeated when a phase is split across
multiple JSONL artifacts; the full profile default already merges the MVP
skill rows, WritingBench/PresentBench ours_full skill rows, and the
PresentBench surrogate auto_skill rows.
The MVP metrics summary is a reporting artifact, not a readiness gate. It covers:
- benchmark score and paired deltas against
prompt_only; - negative transfer rate over successfully paired heldout cells, with missing pairs reported separately;
- skill artifact metrics from generated
SKILL.mdfiles and cross-example reports; - heldout inference token usage for generation, judge, and combined model calls;
- skill self-consistency as an eval-only diagnostic.
Generate the self-consistency diagnostic only from public train examples and the
generated skill. It must not use private rubrics, checklists, heldout feedback,
official scores, or benchmark judge traces. Treat output-example similarity as
debug-only, never as a main effect metric. The summary separates
official_benchmark_scores from surrogate_debug_scores and debug_scores;
do not report PresentBench surrogate scores as official visual/PPT benchmark
results. When baseline and ablation modes live in separate eval JSONL files,
use cross_eval_score_summaries in the output for paired deltas joined across
files with the same evaluator and judge model.
uv run python scripts/metrics/run_self_consistency_metric.py \
--packs artifacts/packs/example_packs.v1.jsonl \
--skills runs/skill_mvp.qwen.mvp.jsonl \
--out runs/self_consistency.writingbench.qwen.mvp.jsonl \
--modes one_shot_skill_from_examples,auto_skill_feature_driven_no_validation
uv run python scripts/metrics/summarize_mvp_metrics.py \
--skills runs/skill_mvp.qwen.mvp.jsonl \
--skills runs/skill_mvp.qwen.ours_full.writingbench.jsonl \
--skills runs/skill_mvp.qwen.ours_full.presentbench.jsonl \
--packs artifacts/packs/example_packs.v1.jsonl \
--modes prompt_only,few_shot_examples_only,one_shot_skill_from_examples,ours_no_validation,auto_skill,examples_plus_one_shot_skill,examples_plus_feature_skill,slide_constrained_examples_plus_feature_skill,layout_plan_examples_plus_feature_skill \
--eval runs/writingbench_official_eval.qwen.five_modes.no_thinking_auto_skill.jsonl \
--eval runs/writingbench_official_eval.qwen.examples_plus_skill.heldout2.jsonl \
--eval runs/presentbench_surrogate_eval.qwen.mvp.jsonl \
--eval runs/presentbench_surrogate_eval.qwen.auto_skill.jsonl \
--eval runs/presentbench_surrogate_eval.qwen.examples_plus_skill.heldout2.jsonl \
--eval runs/presentbench_surrogate_eval.qwen.slide_constrained_examples_plus_feature.heldout2.jsonl \
--eval runs/presentbench_surrogate_eval.qwen.layout_plan_examples_plus_feature.heldout2.jsonl \
--eval runs/writingbench_official_eval.mimo_judge.five_modes.heldout2.jsonl \
--eval runs/writingbench_official_eval.mimo_judge.examples_plus_skill.heldout2.jsonl \
--self-consistency runs/self_consistency.writingbench.qwen.mvp.jsonl \
--out runs/mvp_metrics.heldout2.current.summary.jsonRead AGENTS.md, CLAUDE.md, and CONTRIBUTING.md before starting work. Large features go through branch plus PR. Daily iterations should keep tests green and update repo docs when assumptions change.