Powering Growth With Unified Data

Factored unified data sources and tables into a scalable cloud platform powering trusted analytics.

Key Takeaways:

One unified platform turned fragmented data sources into a scalable foundation for analytics, operations, and AI.
Metadata-driven automation now governs 100's+ tables while sustaining over 99% pipeline reliability.
Trusted, self-service data now powers daily decisions for most of the workforce.

Putting Trusted Insights in the Hands of the our clients Workforce

A leading sustainable infrastructure provider is on a mission to transform how individuals and businesses access critical resources. As the company experienced rapid expansion, it recognized the critical need for a highly scalable data architecture capable of supporting all its long-term data, analytical, and AI ambitions.

By leveraging Microsoft Azure Platform as a Service (PaaS) components alongside Databricks, Factored successfully built a centralized, unified platform. This robust system processes millions of records daily and manages tables containing billions of records. Ultimately, this transformation has empowered a significant majority of our client's workforce with self-service analytics, provided real-time operational insights, and laid the foundation for advanced machine learning applications.

Fragmented Data Couldn’t Keep Pace

To maintain its rapid growth trajectory and operational excellence, our client needed to centralize its entire analytical, data, and AI ecosystem into a single, highly scalable architecture. Prior to this initiative, the company faced the challenge of managing a vast, fragmented, and complex array of data sources.

To achieve their goals, the system needed to seamlessly integrate:

Real-time IoT Telemetry

High-frequency equipment data streaming from numerous distributed sites.

Asset Lifecycle Data

Metrics spanning site construction, development, and active execution stages.

Customer Insights

Behavioral patterns, billing histories, and service usage data.

Third-Party Integrations

External feeds from market providers and public data sources.

To ensure the success of this monumental shift, our team established five core principles to guide the platform's development:

  • Treat Data Quality and Data Security as "first-class citizens."
  • Maintain a consistent Data Pipeline Framework.
  • Create a Common Data Platform presented exactly as needed by the business.
  • Build an analysis tool ecosystem that meets the needs of varying skill sets.
  • Implement agile but robust and iterative Data Governance.

To address these complex challenges, Factored designed and implemented a Common Data Platform framework on Azure and Databricks to ingest and process data into an expanded, four-layer Medallion Architecture. This ensures multiple layers of curated, high-quality data are ready for specific consumption needs:

  • Bronze: Hosted on Azure Data Lake Storage Gen 2, this restricted area receives and stores all raw data in its native format (typically Parquet).
  • Silver: Stored using a lakehouse, this layer standardizes the data format into Delta Lake structures. Crucially, this is where automated data quality checks are executed and Personally Identifiable Information (PII) is secured, providing advanced consumers a safe place to explore foundational data.
  • Gold: Stored in the Databricks lakehouse, this layer serves as the client’s primary data warehouse. It utilizes star schemas and connected data marts to provide standardized, easily consumable data for enterprise reporting.

Under the Hood: Balancing Flexibility, Scalability, and Trust

Our data engineering team pioneered a framework built on custom flexibility and automated scalability, managed through Azure DevOps for seamless CI/CD and agile project management.

1. Custom Queries and Transformations (Flexibility)

Using orchestrated Databricks Jobs, the engineering team automated the ingestion of data from highly diverse sources. The framework allows engineers to utilize SQL to query previous staging areas or use specialized PySpark connectors to load data into a Spark DataFrame. This ensures that regardless of the data's origin, it can be efficiently brought into the processing pipeline.

2. Parametrized Processing via YAML Configurations (Scalability)

To achieve true scalability, we introduced a metadata-driven approach. Using standard YAML configuration files, engineers determine the exact nature of each column in the Spark DataFrame, driving powerful automated transformations and governance features simultaneously:

  • Automated ERD Deployment: Generating Entity-Relationship Diagrams automatically directly from the YAML definitions.
  • Intelligent Lookups: Executing Dimensional Lookups with embedded logic for Slowly Changing Dimensions (SCD Type 1 or Type 2) and time validity.
  • Metadata Management: Maintaining all metadata required for the highly curated Gold tables in one centralized configuration.
3. "First-Class" Quality and Security
  • Data Quality with Great Expectations: Integrated directly into Databricks workflows, Great Expectations acts as automated unit testing for data. It validates information flowing through the pipelines, automatically pausing workflows or sending alerts if data quality rules are breached before data reaches the Silver or Gold layers.
  • Centralized Security & PII Protection: Driven entirely by the central YAML configuration files, the system utilizes a default hashing function to automatically secure PII data based on roles during the table maintenance process. This YAML-driven approach allows for the easy creation of new, tailored hash functions if required. Azure Active Directory handles seamless underlying authorization, while Azure Key Vault securely manages all credentials.
4. Lightweight Data Governance and AI-Driven Consumption
  • Unity Catalog for Unified Governance: With a governance model built on Databricks Unity Catalog, users can easily navigate the vast amount of accessible data. Leveraging the same foundational YAML files, Unity Catalog provides end-to-end automated data lineage, comprehensive documentation, and centralized security. Tables and columns containing PII are automatically tagged via YAML, allowing users to safely filter and search for them using the Unity Catalog search bar.
  • AI-Integrated Reporting: Serving as the principal enterprise consumption layer, our client integrated their reporting tools directly with the curated data layers. By leveraging semantic models and Databricks Genie spaces, this integration provides a highly interactive, self-service experience. This ecosystem isolates the complexity of underlying tables, empowering both technical and non-technical stakeholders to intuitively explore data and uncover insights at scale.

The Platform Turned Data Into a Daily Operating Capability

After years of continuous development, the Cloud Data Platform has fundamentally transformed how our client operates, shifting the organization to a truly data-driven culture.

Platform Scale & Adoption:

  • Dozens of Data Sources: Actively connected to their data platform with a framework to receive more connections efficiently.
  • Hundreds of Tables: Actively maintained by the automated framework.
  • Multiple Dedicated Datamarts: Established in the highly curated Gold layer.
  • Near-Perfect SLA Achievement: Maintained consistently for pipeline reliability.
  • Broad Workforce Adoption: The majority of Full-Time Employees actively utilize the data platform via integrated reporting tools and semantic models for their daily work.
Business Outcomes:
Proactive Operations

Accelerated visibility into field asset performance enables quicker insights and faster decision-making.

Maximized Revenue

Real-time operational data empowers teams to respond proactively to anomalies, maximizing resource yield and revenue across the organization.

Streamlined Reporting

Generating complex performance reports and financial statements has become significantly easier, more accurate, and highly efficient.

The Next Phase Makes the Platform Faster, Simpler, and More AI-Ready

Our client views the platform as a living product and continues to innovate with several strategic initiatives:

  • Serverless Computing for Jobs: Transitioning to serverless architecture to ensure immediate start times for time-sensitive micro-batch workflows, optimize overall compute costs, and deliver faster nightly workflows.
  • Liquid Clustering: Implementing Databricks Liquid Clustering to optimize underlying storage structures, simplify partition maintenance, and dramatically enhance query performance through enhanced statistics and data distribution.
  • Databricks AI/BI Genie & Databricks One: Further focusing on enhancing the experience for non-SQL stakeholders by expanding self-service capabilities and isolating the complexity of underlying table relationships.
  • Lakeflow Ecosystem Integration: Simplifying the connector codebase to reduce maintenance efforts for key platforms (like Salesforce and SQL Server databases).
  • Declarative Pipelines for Power Users: Enhancing the experience for SQL-proficient power users by providing a solid, SQL-based pipeline-building experience that is strongly protected and governed under the platform framework.
  • Advanced MLOps Workflow: Building a dedicated MLOps workflow to simplify, standardize, and scale the management of new machine learning models currently under development for complex scenarios.

Covering 100% of U.S. time zones, becoming a natural extension of your team

Elite engineers ready for flexibility, scalability, and measurable impact.
Build IP that belongs to you
Proven work with the Fortune 500
Start Building
Start Building

Continue Reading

AI Governance Advantage
Governance powers production AI
Agent Sprawl
100+ agents demand governance
AI Observability
Visibility builds trusted AI