Data Quality Framework Accelerator

Catch Bad Data Before It Reaches the Business.

A reusable Data Quality Framework Accelertaor built on PySpark. Define your rules in configuration, run one engine across every dataset, and see exactly which record failed and why.

employee_dataset with 5 rules enabled
Emp_ID Emp_Name Email DOB Experience Status Validation remark
E1001Priya Nairpriya.nair@corp.com1990-04-128 P Passed all rules
E1002Arjun Mehtaarjun.mehta@corp1987-11-0312 R Invalid email
E1003nulll.fernandes@corp.com1995-02-274 R Emp_Name cannot be null or empty
E1004Sara Khansara.khan@corp.com1994-13-406 R DOB must be a valid date
E1005Rohit Dasrohit.das@corp.com1992-07-19five R Experience must be integer
E1001Priya Nairpriya.nair@corp.com1990-04-128 D Duplicate record
Rows processed 0 Passed 0 Rejected 0 Duplicates 0
A sample employee dataset with deliberate issues. The framework adds the last two columns.

Every Dataset Breaks in Its Own Way

By the time a dashboard looks wrong, the bad record has already travelled through reports and business processes. These are the issues that cause it.

  • Required fields left empty
  • Duplicate records
  • Emails and phone numbers in the wrong format
  • Dates and numbers that aren't valid
  • Columns that contradict each other
  • Broken business rules
  • Values missing from reference or master data

Writing separate checks for each dataset means a lot of hard-coded logic, and three problems that grow with every new source.

Inconsistent Validation

Each dataset ends up checked a different way, so "clean" means something different from one team to the next.

Code Changes for Every Rule Change

When the business updates a rule, a developer has to find it, change it and redeploy it.

A Pass or Fail That Explains Nothing

A single result doesn't say which record failed, which rule it broke, or what needs fixing.

One Engine. Your Rules Live in Configuration.

The framework separates what to check from how checks run. Validation rules sit in a configuration file, and a single PySpark engine applies them to any dataset you point it at.

  1. Input Data

    The dataset you want to validate, loaded into a Spark DataFrame.

  2. Configuration

    Which rules apply, to which columns, in what order, with what failure message.

  3. DQ Engine

    The reusable PySpark engine that runs the enabled rules by priority.

  4. Results

    A status and a plain-language remark added to every record.

  5. Audit

    Execution details and summary metrics, stored for every run.

A new dataset needs a new configuration, not new validation code.

Watch the Walkthrough

An employee dataset taken from configuration to execution, results and audit.

Seven Rule Types, Ready to Configure

Together they cover the data quality dimensions that matter most: completeness, uniqueness, validity, conformity, consistency, integrity and text quality.

Rule What it catches Dimension
Null and empty Missing values in mandatory fields Completeness
Duplicate Records that appear more than once Uniqueness
Datatype Values that don't match the expected type, like text in a numeric field Validity
Regex Badly formatted emails, phone numbers, PAN numbers or postal codes Conformity
Expression Business rules and relationships between columns Consistency
SQL or lookup Values that don't exist in reference or master data Integrity
Language Text and character-level requirements Text quality

Rules Are Configured, Not Hard-Coded.

If Employee Name is mandatory, you say so in the configuration file. The engine does the rest. Each rule is described by the same six fields.

rule_name
Which check to run
enabled
Turn a rule on or off without deleting it
priority
The order rules run in
applies_to
The columns the rule checks
parameters
Settings the rule needs, like a pattern or expression
fail_message
The remark written on records that fail

        config.json
      

Rules Change. The Engine Doesn't

As business requirements move, you add or edit rules in configuration. The execution flow stays the same.

  1. 01

    Today

    Employee Name cannot be null. One null_values_check entry covers it.

  2. 02

    Tomorrow

    The business wants Employee Name in a specific format too. Add a format rule to the configuration and run again.

  3. 03

    Later

    A check the library doesn't support yet? Add the rule type to the framework once, then use it from configuration on any dataset.

Same Engine, Any Dataset.

The configuration also says which dataset each rule belongs to, so one framework serves every domain.

  • Employee
  • Recruitment
  • Finance
  • Sales
  • Production
  • Insurance

What Happens When You Run It

You don’t run validations one by one. The framework handles the whole lifecycle in a single execution.

  1. Load the Data

    The input dataset is read into a Spark DataFrame.

  2. Read the Configuration

    The framework finds the configuration for this dataset.

  3. Select Enabled Rules

    Only rules switched on for this dataset are picked up.

  4. Run by Priority

    Rules execute in the order their priority sets.

  5. Mark Every Record

    Each row gets a validation status and, if it failed, a remark.

  6. Record the Run

    Execution metrics are calculated and the summary is saved to Delta.

Know Which Record Failed, and Why

Instead of “this dataset has bad data,” every record carries one of three statuses, so data owners know exactly where to look.

Pass

The record passed every configured rule.

Reject

The record failed one or more rules.

Duplicate

The record was identified as a duplicate.

The Remark Tells You What to Fix.

Every rejected record gets a validation remark written from the rule's failure message. That makes the output something a business user can act on, not just a developer.

  • Invalid email
  • CTC required for Confirmed Employee
  • DOB must be a valid date
  • Experience must be integer

Every Run Leaves a Record.

Beyond record-level results, each execution gets its own ID and summary metrics, persisted to Delta. You can see how quality changes across runs and datasets over time.

Data quality isn't a one-time check. It's something you measure.
Execution summary Sample run on employee_dataset
Execution ID
dq-0924-0831
Rows processed
6
Rows passed
1
Rules applied
5
Rows flagged
5
Runtime
4.8 s

Failures by Rule

  • datatype_check2
  • null_values_check1
  • regex_check1
  • duplicate_check1
  • expression_check0

Built Once, Used Everywhere

Find bad data early, understand exactly what’s wrong, and give the people who own it what they need to fix it.

Reusable

One engine validates every dataset you configure.

Scalable

Built on PySpark to handle large datasets.

Auditable

Each execution is logged to Delta with its metrics.

Maintainable

Business rules change in configuration, not in code.

Actionable

Every failed record says what went wrong.

Standardized

The same validation approach across teams and projects.

Frequently Asked Questions

How the Data Quality Framework validates data, how rules are managed, and how results are tracked.

A data quality framework is a standard way to check data for errors before it reaches reports, dashboards or business processes. Our Data Quality Framework accelerator is a reusable, configuration-driven engine built on PySpark. It applies validation rules to any dataset and flags exactly which records fail and why.

Instead of writing separate validation code for each dataset, you define rules in a configuration file and let one PySpark engine run them. The framework loads data into a Spark DataFrame, applies the enabled rules by priority, and adds a validation status and remark to every record.

It supports seven rule types: null and empty checks, duplicate checks, datatype validation, regex validation for formats like email, phone, PAN and postal codes, expression rules for business logic, SQL or lookup checks against reference data, and language validation. Together they cover completeness, uniqueness, validity, conformity, consistency, integrity and text quality.

Yes. Rules are configured, not hard-coded. You can add, edit or turn off a rule by updating its name, priority, target columns, parameters and failure message in configuration. If you need a completely new type of check, it's added to the framework once and then reused across any dataset.

Yes. The configuration tells the engine which rules belong to which dataset, so the same framework can validate employee, recruitment, finance, sales, production, insurance or other data. Each new dataset needs its own configuration, not new validation code.

Every run gets a unique execution ID and a summary of rows processed, rows passed, rules applied, failures per rule and runtime. This summary is stored in Delta, building an audit history, so you can see how quality changes across runs and datasets.

Don't Wait for Bad Data to Reach the Business.

See the framework run on a real dataset, from configuration to audit, and talk through how it fits your data platform.