Before AI, data: why we're building a Data Engineering practice at Light-it
Every healthcare roadmap I see in 2026 has AI on it. Scribes, copilots, prior auth automation, patient engagement agents.
What almost nobody puts on the roadmap is the thing all of that runs on.
Gartner predicts, that through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data. In healthcare the picture is sharper. A survey of health plan executives found that 92% identify the absence of a single, connected data source as a major barrier to scaling AI across value-based arrangements.
By the time an AI feature underperforms, the model is rarely why. The cause usually sits three systems upstream: mismatched identifiers, undefined metrics, data nobody trusts.
Why we're talking about this now
Data engineering isn't new for us at Light-it. We've been building ETL/ELT pipelines, integrations, and analytics layers inside our healthcare projects for years. What's new is that we're now offering it as a standalone practice, and we've joined the Databricks Partner Program to back it up.
In this post, I'll cover the product and business side of data engineering: why this matters and what it unlocks.
Healthcare has a data problem, not a data shortage
Healthcare organizations sit on more data than almost any other industry. None of it agrees with the rest: different systems, different formats, different definitions of the same patient.A typical digital health company pulls from an EHR, a billing or RCM platform, a scheduling tool, a CRM, product analytics, maybe wearables or lab feeds. Each system has its own identifiers, formats, and definitions of what a "patient" or an "active user" even is. Add PHI, HIPAA, and audit requirements on top, and most teams end up with one of two setups: spreadsheets stitched together by hand, or a dashboard nobody fully trusts.
Both lead to the same outcome. Decisions get made on gut feel, and AI initiatives stall before they reach production. In one
2026 industry survey, data access, quality, or preparation topped the list of reasons healthcare and life sciences AI projects stall, at 73%.
What we learned building our own products
One of the best arguments I can make for data engineering doesn't come from a survey. It comes from our own products.
When we built CompliantChatGPT, and later as it grew into ClinicFrame, we invested early in pipelines and dashboards that connected product usage, subscriptions, and customer behavior into one place we could actually analyze. Not because it was exciting, but because we had real product questions and no reliable way to answer them.
One of those questions was about our free tier. We dug into engagement across free users: daily and monthly active users (DAU and MAU), the DAU/MAU ratio as a signal of stickiness, and how usage compared to the plan's limits.
The data told a story we hadn't expected. A meaningful group of free users were relying on the product as part of their daily clinical workflow, yet few of them ever reached the point where upgrading made sense. For them, the product was already essential. For us, the plan structure wasn't reflecting that value or supporting continued investment in the features they depended on.
So we rethought the model, reshaping our entry-level offering around the usage patterns we were actually seeing. The result was a meaningful wave of conversions to paid plans and a pricing structure that finally matched how clinicians used the product.
That was one of several decisions made on top of the same data foundation, and together they had a significant, measurable impact on revenue.
The lesson for us was simple: the teams that make better product decisions aren't smarter. They just see their business more clearly, and faster. Without reliable pipelines, that free-plan insight would have stayed buried in raw event logs.
Data engineering isn't just dashboards
Dashboards are the visible part. Data engineering is everything that makes a dashboard worth looking at:
- Integration: getting EHR, RCM, product, and third-party data flowing into one place, reliably.
- Quality and trust: validating, cleaning, and normalizing data so a metric means the same thing to everyone in the room.
- Compliance by design: handling PHI with proper access controls, lineage, and auditability from day one, not bolted on later.
- Readiness for AI: structured, governed data that models and agents can actually use in production.
From a business perspective, it comes down to three outcomes: faster decisions, decisions you can trust, and AI initiatives that make it past the pilot.
A few questions worth asking
If you lead a healthcare product or operation, try these questions:
- Can you answer your three most important business questions without having to look in multiple places or asking multiple people?
- Do your product, finance, and clinical teams agree on the same numbers?
- If you launched an AI feature next quarter, is the data it needs clean, connected, and compliant today?
A 'not really' to any of these is the gap this practice exists to close.
What's next
Over the coming months, we'll go deeper on how we approach healthcare data platforms, why we chose Databricks as a partner, and what a practical path from fragmented systems to AI-ready data looks like.