Sumeh: 14개 엔진의 데이터 품질을 위한 단일 API

작성자

카테고리:

← 피드로
DEV Community · Demetrius Albuquerque · 2026-07-21 개발(SW)
Cover image for Sumeh: one API for data quality across 14 engines

Demetrius Albuquerque

Every data quality framework eventually forces you into a choice. Great Expectations gives you a rich ecosystem, but you’re managing suites, checkpoints, and a data context — it’s a platform, not a library. Soda pushes you toward SodaCL and SodaCloud, with the open-source layer getting thinner every release. Pandera is excellent for schema and type validation, but stops there — no aggregation rules, no SQL engines, no row-level tagging.

Sumeh takes a different position: it’s a library, not a platform. It does one thing — run validation rules against data — and does it across every major engine with the same API and the same return type.

from sumeh import pandas
from sumeh.core.rules.rule_model import RuleDefinition

rules = [
    RuleDefinition(field="email", check_type="is_complete", threshold=1.0),
    RuleDefinition(field="age",   check_type="is_between", min_value=18, max_value=120),
]

report = pandas.validate(df, rules)
good_df, bad_df = report.split()
print(f"Pass rate: {report.pass_rate:.2%}")

Enter fullscreen mode Exit fullscreen mode

Same rules, same ValidationReport, whether the engine underneath is Pandas, Polars, PySpark, Dask, DuckDB, BigQuery, Snowflake, Athena, PyFlink, or Ray.

The part I’m actually proud of: single-pass bifurcation

Most validation tools check your data and hand you a report. Sumeh also splits the dataset into clean rows and quarantine rows — in the same scan that computes the metrics:

report = pandas.validate(df, rules)
good_df, bad_df = report.split()
# good_df — original columns, clean
# bad_df  — original columns + _dq_errors (which rules failed, per row)

Enter fullscreen mode Exit fullscreen mode

No double scanning, no extra joins. And on PySpark specifically, this matters more than it sounds: the bifurcation is built entirely from Column expressions evaluated lazily across the cluster. .collect() is never called, not even for sampling. That’s a real limitation in libraries like Deequ — pulling a full dataset to the driver for row-tagging is a driver OOM waiting to happen once you’re past toy data sizes.

What v3.0 changed

The big architectural break already happened in v2.0 — a full rewrite: namespace-first API (from sumeh import pandas; pandas.validate(...)), a ValidationReport object instead of a bare tuple, SQL generation rebuilt on SQLGlot’s AST instead of string concatenation, and the cuallee dependency dropped entirely.

v3.0, released this month, builds on that same foundation — added functionality, no API breaks. If you’re on v2.x, upgrading should be a non-event.

How it compares

I want this comparison to be accurate rather than flattering, so here’s what I found checking each project’s current docs rather than trusting my own assumptions:

Sumeh Great Expectations Soda Core pandera cuallee Engines / backends 14 3 execution engines (Pandas, Spark, SQLAlchemy — the latter covers ~9 SQL dialects) ~10 data sources (Postgres, Snowflake, BigQuery, Databricks, Redshift, and others) 4 native backends (pandas, pyspark, polars, ibis), plus dask/modin via the pandas backend ~7 (PySpark, Pandas, Snowpark, Polars, DuckDB, BigQuery, Daft) Row-level split of good/bad data Native, one method call (report.split()), same pass as validation Possible — unexpected_index_list/unexpected_list expose failing rows, but it’s a manual extraction step, not a built-in split Surfaces “bad” rows per check, but not a single clean/quarantine split call Has drop_invalid_rows to remove bad rows, no built-in quarantine output No built-in split SQL generation approach SQLGlot AST, compiled per dialect at call time Compiles through SQLAlchemy per dialect Compiles SodaCL checks to SQL N/A (DataFrame-native, not SQL-based) N/A (DataFrame-native, not SQL-based) Metadata/catalog export Zero-SDK payload generation for OpenMetadata — you own the HTTP calls Official first-party OpenMetadata integration via OpenMetadataValidationAction Integrates with Soda Cloud (hosted) for monitoring/lineage None built-in None built-in PySpark bifurcation avoids .collect() Yes — fail_condition Column expressions, evaluated lazily N/A (bifurcation isn’t a first-class feature) N/A N/A Uses PySpark’s Observation API, not .collect(), for metrics pip install and go, no mandatory config files Yes No — requires a Data Context Yes, YAML-based Yes Yes

A few honest caveats on top of the table: Great Expectations’ OpenMetadata support is more mature than Sumeh’s, just built differently — official SDK-backed action vs. Sumeh’s dependency-free dict generator that you POST yourself. And cuallee deserves a specific callout — its clean API is what inspired Sumeh’s design in the first place. Sumeh extends that idea to more engines, layers SQL generation via SQLGlot on top, and adds a profiler alongside it.

Where this is actually at

Sumeh is not production-ready — that’s the honest answer, and it’s the reason for this post. It’s early: a handful of engines have solid test coverage (Pandas, PySpark, DuckDB), others are newer and less battle-tested. This is exactly the stage where outside testing matters most, before assumptions calcify.

Concretely, here’s where more eyes would help:

  • Ray Data and PyFlink — these engines have thinner test coverage than the core batch engines. If you already use either in your stack, running Sumeh against real workloads there is high-value.
  • SQL dialect edge cases — every SQL engine shares a compiler that builds queries as SQLGlot AST and compiles to the target dialect at call time. Snowflake, Trino, Doris, Athena, Redshift — each has its own quirks. If you run any of these in production, generating validation SQL against your actual schema (sumeh sql rules.csv --table your.table --dialect snowflake) and checking the output would catch a lot.
  • Bifurcation at real scale — tested so far on dev-sized data, not production volume. If you can point it at something large and report back on memory behavior, timing, and correctness, that’s exactly the kind of feedback this needs.
  • Date and schema validation rules — the newest and least-covered rule categories.

Try it

pip install sumeh                # Pandas included by default
pip install sumeh[pyspark]       # or polars, duckdb, bigquery, snowflake, ...

Enter fullscreen mode Exit fullscreen mode

Repo: github.com/maltzsama/sumeh
Docs: maltzsama.github.io/sumeh

To contribute:

git clone https://github.com/maltzsama/sumeh.git
cd sumeh
git checkout develop
poetry install --with dev
poetry run pytest

Enter fullscreen mode Exit fullscreen mode

Issues, PRs, and bug reports are genuinely welcome — especially in the areas above. If something breaks in your setup, open an issue with the traceback; it’ll get read and acted on.

원문에서 계속 ↗

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다