Milestone M4
September 2026
Reconciling a California Power Bill
An engine that prices a household's own metered intervals against the published tariff sheets — and is held to its residual. Eleven real PG&E statements, worst error twenty-two cents, and a full accounting of everything still assumed.
- Statements reconciled
- 11 / 11within the ±$2.00 gate
- Worst single error
- +$0.220.43% of that statement
- Across all eleven
- 0.14%$1,085.81 modeled vs $1,084.26 billed
- Input
- 4,345kWh of the household's own hourly intervals, 328 days
gate
The gate
Most consumer energy calculators cannot tell you whether they would have reproduced your last bill, because they never try. They price a modeled load profile against a simplified rate, and the error disappears into a “typical savings” headline with no residual to inspect.
This engine takes the opposite constraint as its foundation, and enforces it as a merge gate rather than an aspiration:
Fig. 1 — modeled minus billed, by statement
scripts/reconcile_report.py · gate ±$2.00The residual is one-sided, and that is the point
Every one of the eleven errors is positive: the model runs consistently a little high. A curve-fitted model would have a residual centred on zero. This one has a small, one-directional bias with two known physical causes, both bounded and both recorded:
- Meter-read boundary. Interval sums and the utility's billed kWh differ by about 0.3 kWh per period, because a billing period ends at a meter read time, not at a calendar-day boundary. At residential rates that is roughly $0.10–$0.25 — most of the residual.
- The 2025-vintage franchise fee is still modeled as a percent-of-energy approximation; the 2026 specs use the exact per-kWh rate from Schedule E-FFS. On the two late-2025 statements it shows as a visible +$0.04–$0.05 on that one line.
The errors are physical, not tuned away.
Every statement, billed vs modeled
PG&E E-TOU-C + 3CE generation · CARE| Billing period | Billed $ | Modeled $ | Δ $ | ±$2 |
|---|---|---|---|---|
| Eleven statements | 1,084.26 | 1,085.81 | +1.55 | ✓ pass |
A second tier of evidence, and the line it does not cross
That number is PG&E. On SDG&E the engine has never reproduced a real bill, because no household has yet supplied one — a gap no amount of code closes, restated in §05.
There is a weaker class of public evidence. Every California utility publishes a rate-change notice when rates move, and SDG&E's names a schedule: 400 kWh per month on schedule TOU-DR1, split bundled against unbundled and CARE against standard. Three were filed in 2026, for three different rate vintages. Priced through the same compute_bill the reconciliation uses, the committed TOU-DR1 delivery specs agree with all three: energy rate, Base Services Charge, baseline allowance, baseline credit. It is the only outside evidence of any kind that touches TOU-DR1 — the schedule most SDG&E households are actually on.
The sharpest version of that test assumes nothing about load. Each alert publishes the figure it replaces alongside the one it introduces, so the change is published too. Differencing two vintages cancels everything the document leaves unstated, and the engine's change tracks the published change across every climate zone to within a few cents. A 0.3% error in the delivery energy rate breaks it.
What it is not, stated as plainly as what it is:
- It is not reconciliation, and none of it counts toward the ±$2 figure above. No meter, no billing period, no interval data. A published illustration is an average over an undisclosed population, not a bill that was issued to somebody.
- The document leaves two parameters unstated — how the 400 kWh spreads across time-of-use periods, and which baseline climate zone applies — so both are bounded rather than guessed. An earlier version of this work inferred the climate zone from one quarter's arithmetic; checked against the other two, the inference did not hold, and the test was replaced rather than retuned.
- The generation layer is still untested by anything, and on TOU-DR1 that is where 100% of the time-of-use price signal lives (§04). Every load-shifting and battery conclusion on this schedule rests on numbers no outside document has checked.
- PG&E's equivalent notices cannot be used at all. They quote an average residential bill across every schedule, territory and load shape, which no single tariff spec can reproduce. Recorded for provenance, asserted against never.
These fixtures sit in tests/published_impacts/, deliberately apart from the golden bills, and a test enforces the separation: nothing there may be imported by the engine, so no figure from it can reach anything a visitor reads.
Method
How the engine computes a bill
The bill is a pure function
No I/O, no network, no hidden state, deterministic:
compute_bill(interval_series, [tariff_spec_versions], billing_period) -> itemized Bill
Every utility-specific irregularity lives in a parser or a spec file, never in the billing logic. That is what makes the counterfactual honest: re-simulating the household on a different schedule is the same function with a different spec, not a second code path with its own simplifications.
Bills are computed in layers — delivery, generation, adjustment — because that is how California actually bills. The utility delivers; a Community Choice Aggregator may supply the generation. Separating the layers is what makes the finding in §3 answerable at all.
Tariff specs are data, and they carry citations
Each schedule is a versioned YAML file with an effective date and a citation
naming the exact Cal. P.U.C. sheet number and advice letter every number came from. A spec
without a citation is invalid. An unknown value is written UNVERIFIED and the
loader raises rather than defaulting — an unfinished spec cannot quietly
produce a plausible-looking dollar figure.
That rule has already paid for itself twice. Schedule TOU-DR-P is committed and deliberately does not load: it carries a RYU Event Period Adder of $1.16/kWh on event days, and the engine has no concept of an event-contingent charge. Defaulting the event count to zero would make TOU-DR-P the cheapest plan on the board — a free lunch manufactured entirely out of an unmodeled charge. And committing a genuinely incomplete spec exposed a real defect: the loader validated every file in the directory before filtering by schedule, so one half-finished SDG&E spec broke all eleven golden bills.
Rate versions split inside a billing period
A single statement routinely straddles a rate change, and delivery and generation change on different dates. In the case-study year alone: PG&E delivery changed three times, the summer season began 1 June, and the CCA's generation rate changed on 2026-02-15 — a date derived from the household's interval data by finding the split that reproduces the statement's printed peak/off-peak kWh, then confirmed against the CCA's published sheet.
Periods are N-period, day-typed, month-conditional
Time-of-use resolution is an ordered, first-match-wins rule list. SDG&E requires it: three periods, different weekday and weekend windows, and month-conditional carve-outs. Holidays price as weekend days — and the two utilities do not agree on which days those are. Both name the same eight holidays, but PG&E follows federal observance while SDG&E Electric Rule 1 states that a holiday falling on Sunday moves to Monday and no change is made for holidays falling on Saturday. So on 4 July 2026, a Saturday, a PG&E customer gets Friday the 3rd priced off-peak and an SDG&E customer gets nothing.
Two things that look like schema gaps and are not
Flattening PG&E's tiered E-1 to an average would destroy the marginal price signal the optimizer exists to compare. Instead E-1 is expressed through the existing baseline-credit machinery, which is an algebraic identity rather than an approximation:
tier1·min(u,B) + tier2·max(0,u−B) == tier2·u + (tier1−tier2)·min(u,B) 0.32561 − 0.40702 = −0.08141 ← exactly the printed baseline credit
PG&E's own unbundling confirms this is the tariff's real structure: the tier difference is a named component, the Conservation Incentive Adjustment. And SDG&E's non-bypassable charges already sit inside the published energy rate, so the netting logic carves them out of the offsettable subtotal instead of adding a line — no double count.
One household
One household, end to end
Disclosure. This is the author's own account, published deliberately and with attribution. Anonymising it would be theatre — the repository carries my name — and blurring the dates or amounts would destroy the one property that makes the reconciliation worth reading, which is that a third party can check the rates against the tariff sheets. The account is CARE-enrolled, California's low-income discount, disclosed here because the CARE modeling below is one of the more interesting results and cannot be presented honestly with the rate class hidden.
The account
2025-08-02 → 2026-06-25| Utility / schedule | PG&E E-TOU-C — peak 4–9 p.m. every day |
|---|---|
| Generation | Central Coast Community Energy (3CE), schedule MBRETCH1 — a CCA, so PG&E does delivery only |
| Rate class | CARE (low-income discount) |
| Baseline territory | T, all-electric |
| Local tax | City of Santa Cruz Utility Users' Tax, 8.5% on both layers |
| Meter | hourly intervals — this meter stores nothing finer |
| Window | 4,345 kWh over 328 days |
This is not a vanilla bundled account, and that turned out to be the point. The CCA overlay, the PCIA, the generation credit and the CARE discount that the roadmap had deferred to a later milestone were all required immediately, just to reproduce the first bill.
CARE rates are derived, then independently checked
PG&E does not print residential CARE rates in its tariff book. They were solved from the CARE rates appearing on the household's own statements — two parameters, five independent constraints, fitting to ≤1e-5:
care_energy = 0.65 × standard − 0.01038 care_baseline = 0.65 × standard
The question is whether that is a coincidence of one schedule or a real statewide rule.
Three checks say it is real. Applying it to E-1's two printed tier rates yields a CARE
tier differential of −0.05291, exactly the CARE Baseline Credit
printed on the E-TOU-C statements — a different schedule, a number never used in the fit.
SDG&E, unlike PG&E, prints its CARE tables explicitly, and they state “CARE
Discount 35%”. And on SDG&E the layer split is not merely consistent but exact:
delivery and generation re-sum to the published Total Electric Rate and Total Adjusted
CARE Rate on all twelve season × period cells. It remains an assumption on PG&E, and
--no-care produces the fully tariff-grounded ranking with no derived rate in it.
The result: switch — but not for the reason anyone expects
Re-simulating the same 4,345 kWh of real intervals under every eligible schedule, at rates in effect today, with the CCA on and off:
Annual cost by plan
scripts/rate_optimizer.py · two EV-only schedules excluded as ineligible| # | Plan | Annual $ | $ / mo | vs current | Status |
|---|---|---|---|---|---|
| 1 | E-TOU-C + PG&E | 933.50 | 86.52 | −170.72 | recommended |
| 2 | E-1 (tiered) + PG&E | 990.57 | 91.81 | −113.65 | |
| 3 | E-TOU-D + PG&E | 1,038.50 | 96.25 | −65.72 | |
| 4 | E-TOU-C + 3CE | 1,104.22 | 102.34 | — | current |
| 5 | E-1 (tiered) + 3CE | 1,160.91 | 107.60 | +56.69 | |
| 6 | E-TOU-D + 3CE | 1,208.80 | 112.04 | +104.58 |
EV-only schedules are excluded, not merely flagged, because this household has no plug-in vehicle. Flagging an ineligible plan still lets it become the headline; filtering cannot.
The $170.72 a year has almost nothing to do with the rate schedule. The recommended plan is the schedule the household is already on. The whole saving comes from leaving the CCA, and it decomposes cleanly:
Where the money actually differs
annual $ · credits in parentheses, per tariff convention| Component | Annual $ |
|---|---|
| Power Charge Indifference Adjustment — 2018 vintage | 159.88 |
| Utility Users' Tax levied on that PCIA | 13.59 |
| Franchise fee | 2.54 |
| 3CE generation being cheaper than PG&E's | (4.85) |
| Net cost of staying with the CCA | 170.72 |
The mechanism is checkable on the sheets. A 2018-vintage CCA customer pays a PCIA of 0.03679/kWh, while the 2026 bundled PCIA is −0.01011 — a credit. That ≈4.7¢/kWh gap is wider than the CCA's generation discount, so the CCA has quietly become the more expensive option for this particular vintage. Nothing about the household's usage is unusual: the finding is entirely an artefact of when service started, and generic calculators miss it because they do not model PCIA vintages at all.
What would have to be true for this to be wrong
Reporting a recommendation without its failure conditions is the thing this project exists not to do.
PG&E Rules 22.1/23.1 require six months' advance notice to elect bundled service, with Transitional Bundled Service (Schedule TBCC, short-term market prices) in between. A saving realised months later after an unpriced TBS window is not the annual figure above. This caveat prints in the report itself.
The CCA's discount product, 3Cflex, is 5.264¢/kWh cheaper than the default modeled here. Staying with the CCA and switching products is a plausible option this engine cannot yet rank.
No escalation, no pending rate cases. PCIA vintages reset annually; the gap driving this result can narrow.
E-1 overtakes E-TOU-C only if evening (4–9 p.m.) usage grew 353%, or 132% as a load-neutral shift. E-TOU-D never overtakes it on a load-neutral shift at any tested magnitude.
And the largest single driver between schedules is the baseline credit, which E-TOU-C and E-1 have and E-TOU-D does not — not the peak/off-peak spread most rate advice fixates on.
The unverified assumption, and why it cannot change the answer
The statements carry a Utility Users' Tax Adjustment: a municipal credit cancelling most of the gross 8.5% tax, whose mechanism is not derivable from the bills and could not be confirmed from public sources. Five candidate mechanisms were fitted across all eleven statements; a fixed credit of $0.21649 per billing day had the tightest fit — coefficient of variation 0.080, against 0.207 for per-kWh and 0.330 for a fraction of the net bill.
That is adopted, and still marked UNVERIFIED. What makes it safe is not the fit quality but the shape: a per-day credit is schedule-independent, so it shifts every candidate by the same constant and cannot reorder a ranking. The optimizer does not assert this — it re-ranks under all three rival mechanisms and prints all three orders, which are identical. An assumption you cannot verify should be handled by showing it does not matter, not by arguing it is probably right.
The cross-check that is still missing
The milestone asks for a comparison against the utility's own free rate-comparison tool. It
could not be completed: PG&E's Rate Plan Comparison sits behind an account login and, for
this account, returns “This account has no service agreement eligible for rate
enrollment” — most likely because generation is with a CCA. The harness is
built and waiting behind a --pge-comparison flag, and the report prints the exact
steps to close it. It is listed as open rather than quietly dropped. Four cross-checks that
are complete: every rate traces to a Cal. P.U.C. sheet and the unbundled components
sum to the printed total exactly on all four schedules; the generation credit derived from the
tariff (−0.12699 / −0.10031) reproduces the value
independently least-squares fitted from the bills
(−0.12705 / −0.10030) to 6e-05;
3CE's published sheet reproduces the bill-derived rates exactly; and the E-TOU-C leg of the
ranking is the same engine that reproduces all eleven statements within $0.22.
Public sources
Five findings a generic calculator misses
These came out of building the engine, not out of looking for marketing copy. Each is checkable from public documents.
On SDG&E, the nine-year NEM 3.0 vintage lock-in is currently worth nothing
Export rates under the Net Billing Tariff lock for nine years by application vintage, and the standard installer pitch is to sign before rates drop. SDG&E's published NBT2025, NBT2026 and current-year export tables were parsed in full and compared cell by cell: they are byte-identical for every overlapping year — 0.0 maximum absolute difference across 23,040 cells. On SDG&E today, “lock in before rates drop” has no dollars behind it. That is the opposite of the PG&E vintage story the pitch is borrowed from.
“Today” is dated, and it has a stated expiry. The tables were re-fetched 2026-09-16, after the CPUC adopted the 2026 ACC update on 2026-09-03 (D.26-09-007), and every source file is byte-identical to the one this was measured on — the adoption has not reached either utility's published export pricing yet. But each utility's own readme makes its floating (non-locked-in) table effective only through 2026-12-31, so a republish is due before then and this finding can change with it. It is re-checked by a test that fails on 2027-01-01, not cached on this page.
Export rates rise within a locked vintage, so a lock is not a flat number
The same tables show a 2026 mean full export rate of 0.0883 rising to 0.1432 by 2035. The lock fixes a 576-values-per-year schedule, not a constant. Code that models a locked vintage as one number misprices every year after the first.
On EV-TOU-5, roughly 45% of the super-off-peak delivery charge is non-bypassable
That schedule collapses its super-off-peak distribution charge — 0.04114 against 0.31711 elsewhere — but does not discount the non-bypassable charges. So in exactly the window a battery or an EV charges in, 0.02099 of the 0.04705 total is a charge no export can ever offset, against about 6.5% in the schedule's other windows. A calculator netting exports against the headline delivery rate overstates the value of shifting imports into super-off-peak by roughly 2×.
On TOU-DR1, 100% of the time-of-use price signal lives in the generation layer
The delivery total is 0.32948 for every period in both seasons — perfectly flat. The entire summer on-peak-to-super-off-peak spread of 0.30799/kWh sits in the generation component. Any tool that approximates generation from delivery — tempting, since delivery rates are easier to obtain — produces a flat price signal and understates every load-shift conclusion.
This engine made that mistake internally before the generation layer was authored, and the correction was large: a hand-entered super-off-peak price of 0.22 against a real total of 0.37660. A 71% understatement of the cost of importing in precisely the hour a battery charges — an error biased toward recommending the battery.
A perfect-foresight optimizer can settle worse than a dumb controller
The battery module always reports both a greedy time-of-use controller and a cvxpy linear program. The LP optimises a marginal-price proxy; actual dollars come from the Net Billing Tariff settlement, whose export credit caps and non-bypassable floor the LP never sees. Which one wins in settled dollars is a per-case result, not a fixed ordering. Reporting only the LP and calling it a theoretical maximum would overstate achievable savings on some configurations and understate them on others.
Approximate
Unknown
What is still assumed, approximate, or unknown
Presented in full, because a methodology writeup that lists only its strengths is marketing.
Blocked on data, not on code
Two milestones have complete, tested engines whose definitions of done cannot be demonstrated. The SDG&E reconciliation needs three real SDG&E bills; the specs are tariff-exact and the layers re-sum to SDG&E's published totals, but no household can currently supply bills or a Green Button export. The NEM 3.0 payback needs one real household with solar and export interval data. None exists.
The demonstration driver therefore runs on a clearly labelled synthetic load — and for that reason no payback figure from it appears anywhere in this document. Its rates, non-bypassable charge set and ACC export tables are all real and cited; the load is not, and a payback number computed on a fabricated load is exactly the kind of figure that survives out of context. The engine is done. The claim is not.
Assumptions carried into published numbers
PG&E CARE rates are solved from bills, not published in the tariff book.
--no-care removes them from every figure.
The Santa Cruz UUT Adjustment, handled by demonstrating ranking-invariance across all three candidate mechanisms rather than by picking one.
SDG&E's super-off-peak window went year-round effective 2026-05-01. Sourced from SDG&E's own customer-facing publications and corroborated by contemporaneous local news; the superseding advice letter and revised Cal. P.U.C. sheet were not located. Any spec vintage before that date must restore the March/April restriction.
Known approximations, all under $0.05 per bill
- The 2025-vintage franchise fee uses a percent-of-energy approximation; 2026 uses the exact per-kWh rate from Schedule E-FFS.
- The 2025 summer generation credit is fitted from a single summer sub-period; the 2026 one is tariff-exact.
- The California Climate Credit is modeled as an observed per-bill line rather than a versioned semiannual credit.
- The California Energy Commission surcharge is asserted on PG&E specs and deliberately not asserted either way on SDG&E, whose rate tables show no such column. Adding it unsourced would break the layer-identity test.
Schema limits
Rates carry only standard and care variants, so FERA and the
affordable-housing class are not modelable. Per-kWh adders have no CARE variant. There is no
tiers concept — currently correct, for the reason given in §2. And where a modeling
choice has dollar impact, it is surfaced as a decision rather than made silently: SDG&E
baseline allowances are per climate zone, so the territory must be supplied at bill time and
omitting it raises rather than defaulting, because a wrong zone silently
mis-sizes the largest credit on a California bill.
this
Reproducing this
conda env create -f environment.yml && conda activate energy-advisor pytest # over 1,700 tests PYTHONPATH=src python scripts/reconcile_report.py # the 11 line-item comparisons PYTHONPATH=src python scripts/rate_optimizer.py # ranking, why, sensitivity, assumption audit PYTHONPATH=src python scripts/nem3_report.py # NEM 3.0 engine (real rates, synthetic load)
The golden-bill reconciliation is the merge gate: nothing touching the tariff or NEM modules lands while any of the eleven statements is out of tolerance. Real interval exports and bill PDFs are gitignored and never committed; the committed fixtures carry no name, address or account number.