r/MicrosoftFabric 14d ago

AMA Experts + Engines | Hi! We're the Fabric Data Warehouse team - ask US anything!

30 Upvotes

Hi r/MicrosoftFabric community!

I am Tino from the Microsoft Fabric Data Warehouse and SQL Endpoint team, and I am excited to kick off r/MicrosoftFabric's new Experts + Engines AMA series with you.

At FabCon, I made a bet on stage with u/raki_rahman that if our live petabyte-scale query failed, I would eat a ghost pepper. It was a fun moment, but luckily for me Raki delivered! You can see it here: FabCon stage post.

My team is building a modern, serverless, lakehouse-native warehouse for workloads that need to run blazing fast while being very cost-effective, and keep getting better over time. We recently announced breakthrough innovations such as GPU-accelerated query execution, a new billing model, on-demand billing options, and cache cooldown. Rest assured, we are keeping close tabs on all the things the community reports that we could be doing better - such as metadata sync in the SQL Endpoint.

We lurk in here quite a bit, but would love to learn more from you all and share with you our vision and our thinking around the future.

We're here to answer your questions about:

  • Fabric Data Warehouse architecture: modern, serverless, lakehouse-native design
  • Performance, cost, and ease of use: why Fabric Data Warehouse is differentiated
  • Why Fabric Data Warehouse is a net-new product in Microsoft Fabric (and not Synapse)
  • New capabilities: GPU acceleration, billing updates, on-demand billing, cache cooldown, and so on
  • SQL Endpoint is the same thing as Data Warehouse in the Fabric experience
  • Real-world best practices for migration, workload patterns, and getting great performance
  • What we’re doing to enable AI experiences with our product

Tutorials, links and resources before the event:

AMA Schedule:

  • Start taking questions 24 hours before the event begins
  • Start answering your questions at: July 21, 2026 9:00 AM PDT / July 21, 2026 4:00 PM UTC
  • End the event after 1 hour

r/MicrosoftFabric 1d ago

Announcement Share Your Fabric Idea Links | July 21, 2026 Edition

3 Upvotes

This post is a space to highlight a Fabric Idea that you believe deserves more visibility and votes. If there’s an improvement you’re particularly interested in, feel free to share:

  • [Required] A link to the Idea
  • [Optional] A brief explanation of why it would be valuable
  • [Optional] Any context about the scenario or need it supports

If you come across an idea that you agree with, give it a vote on the Fabric Ideas site.


r/MicrosoftFabric 6h ago

Community Share Direct Lake Inernals - My findings

10 Upvotes

Hey,

I spent the past few weeks, in my leisure time, digging into Direct Lake and trying to find working patterns around eviction, ideal physical layouts, and other things. I also wanted to write a general blog about Direct Lake with some proper hard numbers and evidence, so it can act as a Direct Lake guide beyond the docs.

Here’s the blog I wrote:

https://www.vojtechsima.com/post/microsoft-fabric-direct-lake-framing-caching-performance

I would love to hear whether you have done similar digging, what numbers you got, and what kinds of edge cases you have encountered when using Direct Lake in production.


r/MicrosoftFabric 9h ago

Data Engineering New Spark SQL Query Button - Awesome But!

15 Upvotes

Really like the new query your Lakehouse with Spark SQL Feature in the Lakehouse Explorer, But!

  • Shift+Enter does not execute the query
  • F5 refreshes the browser and you lose your query
  • Run button does not run the highlighted text, but rather all that exists in the query editor

Really nice start, but little quality of life stuff would be greatly appreciated!


r/MicrosoftFabric 3h ago

Data Engineering 400MB optimal file size for SQL and PBI

3 Upvotes

Is anyone able to explain why file sizes for optimal SQL endpoint & PowerBI consumption performance is indicated to be between 400MB to 1GB?

My believe is that the difference is negligible for table sizes that would warrant small file sizes. However, is there a high level technical reason for 400MB?

It's also an odd number, unlike 64, 128, 256, 512.... etc..

https://learn.microsoft.com/en-us/fabric/fundamentals/table-maintenance-optimization

File sizes: Large (400 MB to 1 GB) for optimal SQL and Power BI performance.


r/MicrosoftFabric 11h ago

Administration & Governance Anyone tried the new OneLake lifecycle management from the June update yet?

7 Upvotes

Saw they simplified storage and lifecycle management in OneLake in the June release. Has anyone actually switched over, or still testing it out first?


r/MicrosoftFabric 9h ago

Data Engineering checkpointRetentionDuration - Lakehouse Table CDF incorrectly relying on checkpoint files

4 Upvotes

Perhaps my statement (title) is wrong, however, we are experiencing the same behaviour experienced here https://community.fabric.microsoft.com/t5/Data-Engineering/Change-Data-Feed-bug-Unable-to-reconstruct-state-despite-recent/td-p/4663530

As described here, https://www.mssqltips.com/sqlservertip/7962/microsoft-fabric-automatic-table-maintenance-checkpoint-statistics/ checkpointing is an optimization.

Official documentation across the board mentions the need for logRetentionDuration and file retention:
https://learn.microsoft.com/en-us/fabric/data-engineering/delta-lake-time-travel?tabs=sparksql

No indication of needing checkpointRetention or its impact.

Checkpointing is described as an optimization across the board in official channels, however those channels don't indicate that CDF/TT will fail if the optimization files get removed.

However, CDF fails if checkpoint files have been removed within the log retention window. Take the following scenario:

We have both the delta logs (with a long retention) and the unreferenced parquet files (not vacuumed and a long retention) for the CDF version. However, previous checkpoints have been automatically removed due to the default checkpointRetentionDuration being 2 days.

Therefore, retrieving CDF versions of records are always working if done within 2 days of the transaction, but always fail if done after 2 days even though we still maintain the delta log and the unreferenced files.

Many user forums indicate users having this same experience where the solution is to set checkpointRetentionDuration to the same duration as your logRetention.

Is this how it's supposed to work? Is this a confirmed bug? If it's by design, can Microsoft document it somewhere?

I imagine that the failure is due to reading the delta log, seeing the checkpoint, and no longer being able to reference the checkpoint file and instead of traversing the log for the cumulative changes (which it would do normally), it just fails. Considering the optimization isn't actually needed, it shouldn't fail there. It should revert to traversing the logs/parquet files for TT and CDF.


r/MicrosoftFabric 1d ago

Discussion Apps in the "Fabric" experience UI - papercut gone!

Post image
22 Upvotes

Today I noticed a nice silent update in our tenant (or I missed the announcement).

You can now add the Apps menu to your sidebar while remaining in the Fabric view. Before, this was frustrating since apps were the only option which could be added to the Power BI sidebar but remained unavailable when in the Fabric view. Now, as far as I can tell, there is no reason to leave the Fabric view if you prefer it.

Another papercut gone!


r/MicrosoftFabric 1d ago

Certification Got my dp 600 certification

7 Upvotes

Completed my dp600 certification exam and congratulations popped up on my screen

Thanks to YouTube for the videos abt the exam and what to expect from the exam.


r/MicrosoftFabric 1d ago

Security Fabric Rayfin apps security

4 Upvotes

I deployed an Fabric app(Rayfin) in a Microsoft Fabric workspace and the hosting URL appears to have no authentication in front of it anyone with the link can reach it. How do I add a proper auth layer on top?


r/MicrosoftFabric 1d ago

Community Share I built futils, an open source terminal tool for Microsoft Fabric — deploy from git, run notebooks & pipelines, refresh models, move items

Enable HLS to view with audio, or disable this notification

15 Upvotes

I spend most of my time in Microsoft Fabric — notebooks, lakehouses, semantic models — and I wanted one tool that just does the everyday chores for me instead of clicking through the portal. So I built futils, a single Go binary with an interactive menu that talks straight to the Fabric REST APIs: run notebooks and pipelines, refresh semantic models, move items between workspaces, and deploy from git. The deploy flow is the part I think is genuinely cool.

Open source: https://github.com/DanielAndreassen97/futils

Deploy from git

Deployment pipelines don't care about git, and fabric-cicd wants python and yaml when I mostly just want to see what's about to change and ship it. futils reads origin/<branch> directly — the pushed state on the remote, not your local files — so it deploys exactly what's in git, and a local change you haven't committed and pushed is never picked up by accident. It compares every item against the target workspace and sorts them into new / changed / orphan before anything is touched, and the whole diff is rendered as one HTML file you can open and share:

  • line-by-line diffs per item
  • summary cards per class
  • every GUID it will rewrite from source to target (lakehouses, workspaces, SQL endpoints, OneLake shortcuts)
  • which semantic model each report gets connected to

Fabric stores GUIDs as-is in git, so futils translates them by item name instead — no parameter.yml to maintain.

After the compare you cherry-pick exactly which items to deploy. Items that exist in the target but not in git are only deleted after their own confirm, and things that hold data, like lakehouses, get a second confirm that names each one, because deleting those deletes tables. New items land in workspace folders matching the repo structure, pipelines are published in the right order so a pipeline that calls another never lands first, and at the end it runs the notebooks/pipelines you've set up as post-deploy jobs and saves a timestamped HTML report in your repo.

The other flows

  • Run notebooks or pipelines with parameter overrides — the form pre-fills each parameter's default, so you just change what you want.
  • Refresh semantic models with a table picker.
  • Schema compare — shows drift between two environments' lakehouses, also as an HTML page.
  • Move items between workspaces.

Try it without a tenant

futils demo starts the tool against a built-in fake tenant. Quit and you're back to normal — nothing sticks.

Install

brew install DanielAndreassen97/tap/futils          # macOS / Linux

scoop bucket add futils https://github.com/DanielAndreassen97/scoop-bucket
scoop install futils                                 # Windows

Or grab a prebuilt binary from the releases page.

Feedback and issues very welcome, especially item types you deploy that I haven't covered yet.


r/MicrosoftFabric 19h ago

Service Status ⚠️ [Service Degraded] Power BI Customers in East Asia might observe latency in scheduled refreshes.

1 Upvotes

Status: Degraded | Reported: Jul 21, 2026 at 8:55 PM UTC


Power BI Customers in East Asia might observe latency in scheduled refreshes.


🤖 This post was sent from an automated and unattended service and cannot respond to questions or requests. For official updates, visit the Microsoft Fabric Service Status page.


r/MicrosoftFabric 1d ago

Data Engineering Direct Lake Best Practice: Large Gold Table Used by Multiple Semantic Models

16 Upvotes

Hi everyone,

I'm looking for some advice on a Fabric / Power BI Direct Lake architecture question.

We have a very large Gold layer fact table (e.g., accounting document postings) that is consumed by multiple semantic models. The Gold table is intended to be the central, complete business dataset, so we would prefer not to restrict or modify it for a specific reporting use case.

The challenge is that some of the semantic models only need a small subset of the data:

  • Only a few columns out of many available
  • Sometimes only a limited date range
  • Sometimes only specific document types or business areas

Since the source table is very large, we're seeing performance impacts and are looking for ways to optimize the semantic models without changing the Gold layer itself.

Our setup uses Direct Lake.

My questions:

  1. What is considered best practice in this scenario?
  2. Do you create a dedicated view layer on top of the Gold tables and connect your semantic models to those views?
  3. Can Direct Lake semantic models effectively benefit from filtering rows and columns through SQL views?
  4. Are there other recommended approaches to reduce the footprint of the semantic model while keeping the Gold table fully detailed and reusable?

I'm interested in hearing how others handle large shared fact tables in Fabric when different semantic models require significantly different subsets of the data.

Thanks!


r/MicrosoftFabric 1d ago

CI/CD Having multiple commit issues with git lately

5 Upvotes

My company started using a git repo for source control a couple months ago. So far, we've been storing Power BI reports and a few other artifacts there and it worked nicely up until last Friday.

For some reason, a member of my team committed a Power BI report into git from Service and the system prompted them that it was unable to do it with no explanation. After trying a few times, Source Control prompted us to DELETE several Power BI reports from Power BI service, which is extremely odd since the files seem to be correct in the repo and the reports open well locally.

We've raised a ticket but they seem to be clueless and all they tell is to commit an earlier branch to fix it, which solves the issue only temporarily since any further commit will replicate this issue. From my understanding, Fabric must be doing some sort of mistake and messing with the files already existing in the repo, which causes them not to recognized back into Fabric. At the same time, I'm still unable to commit any new files.

Has this (or something similar) happened to anyone around here? At this point we're just considering to outright ditch git since we seem not to be able to get a solution. Thanks.


r/MicrosoftFabric 1d ago

Community Share Power Query - Fixing Summer Time issues

Post image
3 Upvotes

Side quest from my variable library series to answer a Power Query question and post a blog I promised in a session earlier this year. Handling those summer time dates of the day before at 11pm.

https://hatfullofdata.blog/power-query-fixing-summer-time-issues/


r/MicrosoftFabric 1d ago

Discussion Lakehouse to DataFlow Gen2 to Semantic Model

3 Upvotes

Hi All,

There is a workflow that involves 6 different excel files being combined into one. I am trying to automate the workflow, and I have a few questions:

1- Is it a good idea to create dbo table after DataFlow Gen2? We don't have a database yet because we don't have enough data (probably excel file of 2 years i.e. 24 files each with approximately 25-30 rows and 15-20 columns). If we store it as a table, we can use it for future database.

2- Right now, the Power BI report comes from a semantic model that has master files, but again the master files are month-wise and there are so many master files in the folder, which is connected to semantic model. Is it a good idea to combine monthly files (from now onwards) when bringing the files in DataFlow Gen2? Or keep files separate? For example, the files are: Master File 1 Month X, Master File 2 Month Y, Master File 2 Month X, Master File 2 Month Y, and so on.

3- Sort of continuation of point (2) - Master File 1 -Month X and Master File 2- Month X are combined later to give a full picture. Is it a good idea to combine File 1 by month in DataFlow Gen2 stage, and then combine Master File 1 and Master File 2 at table stage before it becomes part of semantic model?

I hope my questions make sense. I am also working on it the first time, and I want to implement best practices.


r/MicrosoftFabric 1d ago

Data Factory AzureDatabricksWorkspace connector doesn't work in Pipeline with On-Prem Gateway?

Thumbnail
gallery
4 Upvotes

Hello fellow Fabricators.

I'm trying to create a Fabric Pipeline step that would trigger a job in Azure Databricks workspace using the built in activity for that (image #3). Our ADB workspace is accessible only through On-Prem gateway connection, querying from Databricks SQL connector works well through the gateway, but I need to be able to also trigger things when needed from Fabric.

BUT, for some reason the activity settings do not "see" the gateway connection I created beforehand (image #1) and also when I try to create a brand new connection using built-in dialog, it completely ignores any gateways when I type in ADB workspace URL (image #2).

Please advise on how to solve this 🤔 thank you!


r/MicrosoftFabric 1d ago

Data Factory SharePoint connection with Service Principal - useless and outdated

4 Upvotes

Why Sharepoint connection using Service Principal is so useless in Fabric?

According to below documentation, SP can no longer use secret to authenticate. Only client certificate authentication is possible.

Moreover, in ADF the Sharepoint List connector supports certificate authentication and it works.

https://learn.microsoft.com/en-us/sharepoint/dev/solution-guidance/security-apponly-azuread#faq

FAQ
Can I use other means besides certificates for realizing app-only access for my Azure AD app?
No, all other options are blocked by SharePoint Online and will result in an Access Denied message.


r/MicrosoftFabric 2d ago

Data Engineering New Data Engineering Docs / Reorg!

62 Upvotes

For all my fellow data engineering nerds out there, grab a coffee as I've got some more light reading for you :)

NEW DOCS:

UPDATED:

We also did a complete reorg of the left navigation pane to focus on user journeys (i.e., "Develop and author", "Transform and enrich", etc.) and cross-cutting capabilities (i.e., Apache Spark, Delta Lake, etc.)

While there's certainly more work to be done, please share any feedback to help prioritize what is added/updated next.

Cheers!


r/MicrosoftFabric 1d ago

Data Engineering Bricked MLVs ?

3 Upvotes

Hello dear community,

We have a gold layer composed of MLVs, one of the refresh seems to have ran for 576 hours before being deadlettered and cancelled. I can see the job as Failed/Deadlettered status in the monitoring tab.

Monitoring Page Details of the refresh job

But on the Manage MLVs page, Fabric still seems to believe this run is ongoing and proposing me to cancel the run which immediately fails with "Couldn't cancel refresh" the HTTP error code is 400.
If i go to the job instance page, it's still displayed as running.

Job instance page details

This is a problem as this means I can't start a new refresh.
Is there any way for me on my end to fix this without duplicating then deleting the bricked lakehouse ?

Thanks a lot for your help and time.


r/MicrosoftFabric 2d ago

Community Share Workshop covering RAG architecture, vector search, Fabric, and knowledge graphs together, thought this would be relevant here

14 Upvotes

Came across this and thought it'd be worth sharing here, most resources cover vector search, Microsoft Fabric, or knowledge graphs separately, but this one actually puts them together as parts of the same enterprise RAG architecture, which is closer to how these systems actually get built in practice.

It's a hands on session on August 8, led by Brian Bønk, a Data Platform MVP and Microsoft FastTrack Solution Architect. Goes through the full pipeline, ingestion, chunking, metadata, vector search, retrieval tuning, evaluation and governance, and then knowledge graphs and ontology as an extension pattern for stronger grounding and traceability. There's also a section on using Fabric and Power BI specifically for business adoption, which is something I don't see covered together with the RAG side very often.

You come out of it with an actual rollout plan rather than just slides, which is the part I found most useful when I looked into it.

Link if anyone wants to check it out: https://www.eventbrite.co.uk/e/design-enterprise-grade-rag-systems-with-llms-vector-search-tickets-1992561384740?aff=rn


r/MicrosoftFabric 2d ago

Community Request We brought fabric-cicd into the Fabric CLI. Where does fab deploy fit in your workflow?

15 Upvotes

Hi everyone. I work on the Fabric CLI team, and I’d like to hear from people who are automating Fabric deployments.

At FabCon earlier this year, we introduced fab deploy. It runs fabric-cicd under the hood, but packages the deployment flow as a Fabric CLI command.

The idea was to give teams that already use fab for authentication and automation a way to deploy Fabric items without writing and maintaining a Python wrapper.

With a YAML configuration, fab deploy can:

  • Publish multiple workspace items from a Git checkout or locally exported definitions
  • Resolve dependencies when logical IDs are available
  • Apply environment-specific values across dev, test, and production
  • Control which items are published or unpublished in each environment

A basic deployment looks like this:

fab deploy --config config.yml --target_env dev

I don’t see the CLI and the Python library as competing deployment engines. The CLI uses the library.

Calling fabric-cicd directly may be the better fit when you need deeper Python customization or want to compose deployment behavior programmatically. The CLI may be a better fit when you want a declarative, one-command step in GitHub Actions, Azure DevOps, or another automation environment.

I’m interested in how people make that choice:

  • Where would a CLI command fit better than calling the Python library directly?
  • What keeps you on the Python API today?
  • If you considered fab deploy but decided against it, what got in the way?
  • What would make the CLI path meaningfully more useful in your CI/CD workflow?

Positive or negative, both kinds of feedback are useful.

Docs:
[https://microsoft.github.io/fabric-cli/commands/fs/deploy/](vscode-file://vscode-app/Applications/Visual%20Studio%20Code.app/Contents/Resources/app/out/vs/code/electron-browser/workbench/workbench.html)

One note if you try it: start with a non-production workspace and review the publish and unpublish settings first. Both operations are enabled by default.[]()


r/MicrosoftFabric 2d ago

Certification Passed DP-700 on My First Attempt! Sharing My Experience

33 Upvotes

First of all let me clarify this is not an easy Exam, it requires very in depth knowledge and I judged a lot of the answers based on my experience with data engineering which I didn't learn from the learning path of fabric

Score & Prep Time:
830/1000 - Passing Marks: 700

2 Months + 3 Years of Experience in Data Engineering

Time Give 100 MINS

Total Questions: 49 (44 Normal Ones & 5 Case Study Questions)

All of my Experience:
Firstly I prepared from Aleksi's Playlist of DP-700 (This Helped a lot in fundamental understanding of each fabric component), then took the Microsoft Learn DP 700 Course on Youtube, I also took a lot of practice tess from Microsoft, I also repeated the Aleksi's Playlist again and did a small project in Fabric to practice hands on learning because at the end, the purpose is to be able to Implement data engineering workflows in fabric and provide value.

The crazy part starts from here, when I felt confident enough to give the exam and scheduled it and just before the exam their system crashed, and showed failures to verify my Identity and Machine to be qualified for the exam. I got into a lot of stress at the time, but thanks to the support of Microsoft they rescheduled it to another date. In the meantime I used Claude & ChatGPT to take my practice exams and also contacted my senior colleagues to share some tips and experience (which helped a lot) as its always better to communicate with some who has more experience. Anyways the date has come and I tested my system and again and finally it worked but then something happened and just a minute before the exam started my Laptop got shut down 🥲 I got into a lot of anxiety because the time has begun and the OnVue software takes half an hour to run the system test and verify the identity, but at the end the exam started after 1 hour of the scheduled time, the first 2 - 5 questions were really hard compared to the rest of the exam but still I managed to finish the exam in under 1 Hour.

What to Expect:
90% of the questions were scenario based and practical in terms of using fabric things.
Practice all the items and features in fabric and you should be able to answer these questions to yourself:

How to use it?
Where to use it?
Why to use it?
Pros & Cons of using it and not using it.

Practice:

Practice KQL More than SQL as Microsoft asks questions related to KQL more!

Good Luck for your Exam!!


r/MicrosoftFabric 2d ago

Data Engineering SQL Endpoint refresh

6 Upvotes

I know this topic has been posted about a lot.

We have many pipelines that we run which follow the same pattern, but the last step is a warehouse stored procedure that populates many fact and dim tables using lakehouse table as the source. Each lakehouse table is dropped and recreated with each load, it is an inherited process. We get errors all the time about the tables not existing in the lakehouse when the warehouse stored proc tries to read from them. So, I implemented a notebook to run in between the lakehouse load and the warehouse load to refresh the lakehouse SQL endpoint using the API

{FABRIC_API_BASE}/workspaces/{workspace_id}/sqlEndpoints/{sql_endpoint_id}/refreshMetadata

Our pipelines are schedule throughout the night and it seems that the first one will still get the metadata errors, but any subsequent pipelines will load successfully. I found this section in Microsoft documentation, and am wondering if my first process fails because it is waiting for the endpoint to spin up and for the background process to complete. Being that my notebook only waits for the SQL endpoint to refresh, is it possible that the background process has not yet completed before my stored proc runs? Like I mentioned, any subsequent pipelines finish, but I'm wondering if that is the case since the endpoint is already spun up.

A background process is responsible for scanning the lakehouse for changes and keeping the SQL analytics endpoint up-to-date for all the changes committed to lakehouses in a workspace. The Microsoft Fabric platform transparently manages the sync process. When a change is detected in a lakehouse, a background process updates metadata and the SQL analytics endpoint reflects the changes committed to lakehouse tables. Under normal operating conditions, the lag between a lakehouse and SQL analytics endpoint is less than one minute. The actual length of time can vary from a few seconds to minutes depending on many factors that this article discusses. The background process runs only when the SQL analytics endpoint is active and it halts after 15 minutes of inactivity.

Any recommendations? Or do I move the refresh SQL endpoint notebook to earlier in my pipeline so that it is up and running when the lakehouse tables get loaded?


r/MicrosoftFabric 2d ago

Security Workspace Identity can be Last Modified By of a pipeline

3 Upvotes

I’ve been involved in a few different discussions lately where setting Last Modified By on a pipeline came up and it has been noted that it works for SPN but not WI.

In case anyone is curious, it does actually work fine to set the last modified by of a pipeline to the WI.

When triggered via schedule, it runs as the WI security context.

I don’t suggest you do it necessarily, but if you want to run something that’s completely secretless, it should work fine.

The other options:
User accounts tokens will fail eventually, so that’s no good.

SPN token expiry mechanism isn’t documented, so last modified by SPN works but you have to periodically reset the SPN token as per many users experiences.

WI doesn’t have passwords or secrets. Thus, we can assume it won’t just stop working due to not being able to get a token.

Drawback: WI scoped to the workspace

Update from u/banner650 :‘Unfortunately, at this point in time, WI is treated like an SPN if it ends up as the Last Modified By for a Pipeline (or other item) and uses a captured refresh token that can expire. As a result, it is no more durable than an SPN. We have some work in flight right now to improve this situation, but I don't have a timeline that I know that I can share yet for it.’