Skip to content

Experiment Playbooks

This page collects concrete experiments you can run directly against the current repository.

Runnable recipes

Choose An Example By Question

Each recipe is a real command path you can run today. Start with the baseline if you want raw behavior, then move to projection or Lean table shielding when you want safety and auditability.

Recipe 1

Baseline

See what the policy does before any runtime correction.

Jump to recipe

Recipe 2

Projection Shield

Inspect where online projection changes unsafe proposals.

Jump to recipe

Recipe 3

Lean Table Shield

Run the strongest Lean-to-runtime integration path.

Jump to recipe

Recipe 4

Solar Export

Switch to a visibly different midday-surplus profile.

Jump to recipe

Recipe 5

Stress Test

Push tighter limits and watch margins shrink.

Jump to recipe

1. Baseline Without a Shield

Use this when you want to see the raw policy behavior first.

python3 -m vpp_rl.train \
  --episodes 180 \
  --algorithm double_q \
  --scenario peak_shaving \
  --shield-mode none \
  --output artifacts/peak_shaving_double_q_none_v3_trace.json

What to inspect:

  • whether the policy learns useful arbitrage at all
  • whether low SoC states become unstable
  • how much improvement exists before safety filtering

2. Projection Shield Run

Use this when you want a runtime-safe controller without requiring a pre-exported rule table.

python3 -m vpp_rl.train \
  --episodes 180 \
  --algorithm double_q \
  --scenario peak_shaving \
  --shield-mode project \
  --output artifacts/peak_shaving_double_q_project_v3_trace.json

What to inspect:

  • how often the shield intervenes
  • whether the projected action stays close to the proposed action
  • whether reward improves relative to the no-shield baseline

3. Lean Rule-Table Shield

Use this when you want the strongest Lean-to-runtime integration path.

python3 -m vpp_rl.rule_table \
  --scenario peak_shaving \
  --output artifacts/peak_shaving_shield_table.json

python3 -m vpp_rl.train \
  --episodes 180 \
  --algorithm double_q \
  --scenario peak_shaving \
  --shield-mode table_project \
  --shield-table artifacts/peak_shaving_shield_table.json \
  --output artifacts/peak_shaving_double_q_table_project_v3_trace.json

What to inspect:

  • whether the Lean-derived table changes runtime behavior
  • whether intervention reasons cluster around battery or grid limits
  • whether the resulting trace verifies cleanly in Lean

4. Solar Export Example

Use this when you want a visually different operating profile with stronger midday solar surplus.

python3 -m vpp_rl.train \
  --episodes 180 \
  --algorithm sarsa \
  --scenario solar_export \
  --shield-mode project \
  --output artifacts/solar_export_sarsa_project_trace.json

What to inspect:

  • midday charging and export behavior
  • whether the battery reserves energy for later peak pricing
  • how reward shape changes relative to peak_shaving

5. Stress Test Safety Run

Use this when you want tighter constraints and more frequent edge conditions.

python3 -m vpp_rl.train \
  --episodes 220 \
  --algorithm double_q \
  --scenario stress_test \
  --shield-mode table_project \
  --shield-table artifacts/stress_test_shield_table.json \
  --output artifacts/stress_test_double_q_table_project_trace.json

What to inspect:

  • whether interventions become more frequent
  • whether grid and SoC margins shrink sharply
  • whether the policy still beats the idle baseline under tighter limits

Suggested Reading Order

  1. Run the Runtime Demo
  2. Try recipe 1 and recipe 3 back to back
  3. Move to the English User Manual for extension points