Skip to content
All insights

Insights

How to size a Microsoft Fabric capacity against real workload

Most Fabric capacities are sized by guess and corrected by pain. A two-week method for sizing from measured load, what smoothing and throttling do to the number you pick, and the three mistakes that leave an F64 doing an F16's work.

Author
Aaron Stark
Published
October 9, 2026
Length
7 min read

The number you are actually choosing

A Fabric capacity is a pool of compute, measured in capacity units (CUs), that every workload in the workspaces assigned to it draws from: semantic model refreshes and queries, Spark notebooks and jobs, pipelines, Dataflows, Eventstreams, warehouse queries, Copilot. SKUs run from F2 to F2048 and the number is the CU count, so an F64 is 64 CUs and costs 32 times an F2. Pricing is linear. There is no volume discount for going bigger, only the reservation discount for committing to a year.

Two thresholds matter beyond cost. At F64 and above, Power BI content on the capacity can be viewed by users with a free license; below F64, every viewer needs a Pro license, which is often the real reason a mid-sized company lands on F64 whether or not the workload needs it. And the Direct Lake guardrails (files, row groups and rows per table, model size, memory) scale with SKU. What happens when a model exceeds them depends on the Direct Lake variant: on SQL endpoints, queries fall back to DirectQuery and every visual starts paying for its own query; on Direct Lake on OneLake, the refresh fails and the model cannot be queried until the tables are optimized.

Smoothing changes what utilization means

Fabric does not bill on peak demand. Interactive operations, a report query, a DAX query, a warehouse query someone is waiting on, are smoothed over five minutes. Background operations, refreshes, Spark jobs, pipelines, Dataflow runs, are smoothed over twenty-four hours. A refresh that consumes a large burst of CU-seconds at 2 a.m. shows up in utilization as a thin layer spread across the whole day, not a spike.

This is why a capacity can read 40 percent utilized on the chart and still throttle. Utilization is the smoothed total; throttling is triggered by how far ahead of itself the capacity has borrowed. Bursting lets a workload consume more than the purchased CUs for a short time, and smoothing pays the overage back over the following minutes or hours. If the debt is not repaid, throttling applies in stages: after ten minutes of overage, interactive requests are delayed; after sixty minutes, interactive requests are rejected; after twenty-four hours, background jobs are rejected too. Administrators can also configure surge protection to cap background usage before it reaches the point of hurting interactive users.

So the question is never what the average utilization is. It is what the smoothed load looks like at the busiest hour of the busiest day, and how much overage debt is being carried into it.

A method that takes two weeks

  1. Install the Microsoft Fabric Capacity Metrics app and give it two full weeks that include a month-end. Shorter windows miss the refresh pile-up that month-end reporting creates.
  2. Separate background from interactive. On the overview, background is the floor and interactive is what sits on top of it during working hours. Note the hour where the two together peak.
  3. Attribute the background floor by item. The timepoint detail shows which semantic models, notebooks, and pipelines account for it. In most estates, three to five items are more than half the floor.
  4. Attribute interactive peaks by item and by user count. A single report with a poorly designed model can be the entire peak. That is a modeling fix, not a capacity fix.
  5. Size to the peak hour. The target is smoothed utilization at the peak hour comfortably under capacity, in practice 70 to 80 percent, with no sustained overage carried into it. Pick the smallest SKU that meets that, then check the two thresholds above: free viewers at F64, and the Direct Lake guardrails for your largest model.
  6. Re-measure after every structural change: a new large model, a Spark workload, a Copilot rollout. Capacity sizing is a quarterly activity, not a procurement event.

The three mistakes that produce the wrong number

Sizing for the refresh instead of moving it. Import-mode refreshes scheduled at 6 a.m. so the numbers are fresh for the morning stack into a peak that the capacity is then sized for. Spread them across the night, use incremental refresh so each run touches only changed partitions, and the floor drops. The capacity you no longer need for the 6 a.m. spike is usually one SKU size.

Fixing a model problem with capacity. A flattened wide table, a measure that iterates a fact table row by row, or a visual that issues forty queries per render will consume whatever capacity it is given. Doubling the SKU halves the pain and doubles the bill. The before-and-after on a model refactor is routinely larger than the before-and-after on a SKU change, and the refactor does not recur monthly on an invoice.

One capacity for everything. Spark and pipelines are bursty; semantic models serving people need headroom at 9 a.m. On a single capacity, a notebook that runs long at 8:45 delays every report that opens at 9. Splitting into one capacity for engineering workloads and one for semantic models costs nothing extra when the total CUs are the same, and it stops background work from throttling interactive users. The exception is a small estate where the total is under F16; there, one capacity and careful scheduling is the right answer.

Reserved or pay-as-you-go

Pay-as-you-go capacities can be paused and resumed, and scaled up or down in minutes. Reservations trade that flexibility for a lower rate. The usual sequence is pay-as-you-go for the first quarter while you measure and fix the model and scheduling issues above, then reserve at the size you have settled on. Reserving before measuring locks in the guess.

What this looks like in practice

A common pattern in mid-sized companies: an F64 bought for the free-viewer threshold, running a smoothed load that would fit an F16 or F32, with one or two large import models producing a morning spike and one report with a model problem producing the afternoon one. Fixing the two models and spreading the refreshes is days of work, not weeks. After that the capacity decision is clear, and the free-viewer question becomes a licensing calculation, Pro licenses for the actual viewer count against the F64 premium, rather than a default.

Lakewarden's capacity module, which we built for exactly this analysis, gives per-item CU attribution, hour-by-weekday load, and a SKU what-if against measured data, with years of retention rather than the metrics app's fourteen days. The method above works with the metrics app alone. The tool makes it repeatable.

If you are sizing a capacity now, or suspect yours is wrong in either direction, the free thirty-minute architecture review covers it. Bring the metrics app and we will read it together.