Skip to content
All insights

Insights

Direct Lake or Import: choosing a storage mode for a Fabric semantic model

Direct Lake removes the refresh, not the modeling work. When it beats Import, when Import is still the right answer, the two Direct Lake variants that behave differently under pressure, and the Delta table habits that decide performance.

Author
Aaron Stark
Published
October 9, 2026
Length
8 min read

The short answer

Choose Direct Lake when the data is already being prepared upstream in Fabric, the tables are large or change often, and someone owns the Delta tables underneath the model. Choose Import when the model author needs to shape data in Power Query, the volume fits comfortably in a scheduled refresh, or the content has to run without a Fabric capacity. Most estates end up with both, and the decision is made per model, not per platform.

The rest of this article is the reasoning, and the details that decide it in practice.

What actually differs

An Import model holds its own compressed copy of the data. Refresh rebuilds that copy from the source, which takes time and capacity and puts load on the source systems. In exchange, the model is self-contained: Power Query lives inside it, and it works on any Fabric or Power BI license, including Pro on shared capacity.

A Direct Lake model reads Delta tables in OneLake directly into the same in-memory engine. Its refresh, called framing, copies only metadata: it points the model at the latest version of each Delta table and completes in seconds. Data volumes that would need hours of Import refresh become available as soon as the upstream pipeline commits. The price is that all data preparation has to happen upstream, in Spark, T-SQL, Dataflows or pipelines, because there is no Power Query inside a Direct Lake model. It also requires a Fabric capacity.

That is the whole trade. Direct Lake removes the refresh; it does not remove the need for a well-designed star schema, and it moves data preparation out of the model and into the platform.

There are two Direct Lake modes, and they fail differently

Fabric now offers two variants, and choosing between them matters as much as choosing Direct Lake at all.

Direct Lake on SQL endpoints reads one Fabric item through its SQL analytics endpoint. When it cannot read a Delta table directly, it falls back to DirectQuery: for SQL views, for tables protected by SQL-based row-level security, or when the model exceeds the capacity's guardrails. Fallback keeps reports working but makes them slower, often much slower, and it happens silently unless someone is watching. The Direct Lake behavior property can disable fallback, in which case those queries fail instead.

Direct Lake on OneLake can combine tables from several Fabric items and never falls back to DirectQuery. It also supports composite models with Import tables, and offers calculated tables and a limited form of calculated columns in preview. The trade-offs: SQL-endpoint security (row, object and column level) is not enforced, so security has to be defined in the semantic model; non-materialized SQL views cannot be used as sources; and if the model exceeds a guardrail, the refresh fails and the model cannot be queried until the tables are optimized.

The practical difference is how each one degrades. On SQL endpoints, a model that outgrows its capacity gets slow. On OneLake, it stops. Neither is better in the abstract; what matters is that the team knows which one it is running and monitors for the corresponding failure.

The guardrails are per table and per SKU

Direct Lake limits are set by capacity size and checked per query, except model size, which is checked for the whole model. The ones that bite first are files and row groups per table, not rows:

SKUParquet files and row groups per tableRows per tableMax model sizeMemory before paging
F2 to F81,000300 million10 GB3 GB
F161,000300 million20 GB5 GB
F321,000300 million40 GB10 GB
F645,0001.5 billionUnlimited25 GB

A 50-million-row fact table that is appended to every fifteen minutes can hit the 1,000-file limit long before it gets near the row limit. Memory is not a hard limit: exceeding it causes columns to be evicted and reloaded, which shows up as inconsistent query times rather than errors.

Delta table habits decide performance

Direct Lake performance is mostly a property of the Delta tables, which means it is owned by whoever writes them. Five habits matter most:

  1. V-Order on. Microsoft's guidance is explicit: apply V-Order. Depending on the workspace's Spark settings it may not be on by default, so check the tables that feed the model rather than assume. It improves compression and lets the engine work on compressed data.
  2. Large, even row groups. Microsoft's guidance is roughly 1 to 16 million rows per row group. Row groups well under a million rows produce many small column segments and slow every query.
  3. Compact small files. Frequent small appends accumulate files quickly. Run OPTIMIZE on a schedule that matches the write pattern: weekly for daily loads, daily or more often for near-real-time ones. Expect the first queries after an OPTIMIZE to be slower, because the affected data has to be reloaded into memory.
  4. Avoid destructive rewrites. An overwrite, or a delete that touches most files, forces Direct Lake to reload the whole table. Partitioning by a low-cardinality column such as month (Microsoft suggests fewer than 100 to 200 distinct values) confines a daily delete-and-reload to one partition.
  5. Watch cardinality. High-cardinality columns, such as transaction IDs or timestamps with seconds, are expensive to load and rarely useful in a report. Leave them out of the model or split them.

None of this exists in an Import model, because the model owns its own storage. Moving to Direct Lake moves this work from the person who builds the model to the person who builds the pipeline, and in many companies those are not the same team.

When Import is still the right answer

Import remains the better choice in more situations than the marketing suggests:

  • Self-service models. An analyst who needs to shape data quickly, without waiting on a data engineering backlog, needs Power Query in the model.
  • Sources the model author cannot change. If the upstream tables are not yours to optimize, Direct Lake performance is not yours to control.
  • Features Direct Lake does not have. Hybrid tables, model-level partitions and user-defined aggregations exist only in Import.
  • Small, slow-changing data. A model that refreshes in four minutes overnight gains nothing from framing, and it keeps working on any license.

For an existing Import model, OneLake integration can write its tables out as Delta in OneLake without migrating the model, which makes the data available to other Fabric workloads and is a low-risk first step toward Direct Lake.

Two details that catch teams during migration

Direct Lake does not support auto date/time. A model that relied on it needs a proper marked date table before it moves, and every visual that used the automatic hierarchy needs rework. Relationship columns must also have matching data types and unique values on the one side; Import tolerated some looseness here that Direct Lake does not.

A decision sequence

  1. Is the data already prepared upstream in Fabric, or will it be? If not, start with Import.
  2. Is refresh time or data latency a problem today? If not, there is little to gain.
  3. Does someone own the Delta tables and their maintenance? If not, Direct Lake will degrade over time.
  4. If yes to all three, pick the variant: OneLake for multi-source models and when you want to rule out silent fallback; SQL endpoints when the model depends on SQL-based security or views.
  5. Check the largest fact table's file and row-group counts against the guardrails for your SKU before building, not after.

The free thirty-minute architecture review covers this decision for a specific model: what it reads, how its tables are written, and which mode fits. Bring the model and the table it struggles with.

Sources: Microsoft Learn, Direct Lake overview and Understand Direct Lake query performance, as of October 2026.