You are eighteen months into prequalification. Your lawyers have revised the baseline methodology twice. The grid operator has just asked, politely but firmly, for the metering documentation again. The technology works. The sites are ready. The market, it turns out, has a gate you never saw in the brochure.

Wholesale electricity balancing markets exist to keep supply and demand in real-time equilibrium. When a large generator trips offline unexpectedly, or when a cloud passes over a solar farm at the wrong moment, the grid operator calls on balancing resources: fast-acting generation or load that can respond within minutes, sometimes seconds, to close the gap. Demand-response providers are, in theory, ideal candidates. They represent industrial facilities, commercial buildings, or aggregated households that can cut consumption on command, effectively acting as a virtual power plant running in reverse. The physics work. The economics are plausible. The market structure, though, is a different problem entirely.

The gate is built from prequalification, not prejudice

The exclusion is rarely deliberate. It emerges from a set of technical prequalification criteria written, historically, with synchronous thermal generators in mind. To participate in a balancing mechanism, a resource typically must demonstrate three things: a minimum capacity threshold, a guaranteed response time, and a firm, metered baseline from which deviation can be measured and settled.

Start with the minimum capacity threshold. In Great Britain's Balancing Mechanism, the de facto floor for direct participation has sat around one megawatt. A single large factory might clear that bar. An aggregator pooling the flexible load of forty small commercial buildings probably cannot offer that capacity from a single metered connection point, and many market designs require exactly that: one grid connection, one meter, one bid. The aggregation itself is the product, but the market was designed around atoms, not portfolios.

Response time requirements cut even deeper. Many balancing products require delivery within two minutes of instruction, or in some cases within thirty seconds for the highest-value frequency services. A gas turbine receives a signal and opens a valve. A demand-response aggregator receives a signal and must dispatch instructions across a software layer to dozens of separate sites, each of which has its own control system, its own latency, and its own local decision logic. Even if the aggregate response lands within the window, proving it in advance, to a grid operator's satisfaction, is a certification process that can take months and costs money that smaller aggregators cannot front. Think of it as being asked to audition for an orchestra by submitting, in writing, a sworn affidavit that you can play the oboe.

Then there is the baseline problem. This one is genuinely thorny. Generators are settled against actual output: you said you'd produce 50 MW, you produced 47 MW, you're penalized for the 3 MW shortfall. Clean, auditable, hard to game. A demand-response provider is settled against a counterfactual: what would this site have consumed if we hadn't called it? That baseline must be estimated, and every estimation methodology embeds assumptions that favor some participants and disadvantage others. A factory running a predictable shift pattern looks good on paper. A data center whose load varies with internet traffic looks unreliable, even if it can genuinely flex. The settlement methodology, written for one kind of asset, becomes a structural veto on another.

Here is a worked example that illustrates how this friction compounds. Consider two aggregators who entered the same regional balancing market in the same year. One represented a handful of large water-treatment facilities with stable, predictable overnight loads. The other represented a mixed portfolio of cold-storage warehouses, retail units, and small manufacturers. The first cleared prequalification in a single application cycle. The second spent eighteen months revising its baseline methodology, resubmitting metering documentation, and ultimately splitting its portfolio into sub-portfolios to hit capacity minimums per connection point, tripling its administrative overhead in the process. Same underlying flexibility. Vastly different friction.

The cost nobody talks about

Compliance costs are the invisible barrier that capacity thresholds and response times merely introduce. A resource that clears every technical hurdle still faces legal fees, metering upgrades, telemetry installation, credit requirements (many markets demand collateral posted against imbalance exposure), and ongoing reporting obligations. For a utility-scale battery developer, these are line items. For a demand-response aggregator whose margin depends on thin arbitrage across many small sites, they can consume the entire business case before a single bid is submitted.

And here is the judgment that market designers are still reluctant to say aloud: the compliance architecture is not neutral. It was calibrated for capitalised incumbents, and pretending otherwise wastes everyone's time.

Grid operators are not indifferent to this. Several have introduced simplified participation routes, sandboxed trial frameworks, or aggregation-specific rule changes. But reformed rules layer onto inherited infrastructure. The telemetry protocols that a new aggregator must integrate with were written for a different era. The settlement systems run on batch cycles designed around day-ahead markets, not the sub-five-minute granularity that modern flexible load can actually offer. So ask yourself: if the rules were written fresh today, with full knowledge of what demand response can do, would they look anything like what currently exists?

They would not. Which is the point.

Balancing markets were not designed to exclude demand response. They were designed to balance the grid reliably, using the tools available at the time of their construction. Demand response arrived later, speaking a different technical dialect, and the translation is still incomplete. That gap is not a technical inevitability. It is a policy choice, repeated every time a rule revision stalls in committee, and the cost of that choice is borne entirely by the resources the grid will need most as generation becomes less predictable.