Skip to content

DataCitizen

Democratizing enterprise data engineering.

DataCitizen Data Groups Management, listing pipelines with source and target catalogs, attribute counts and run status

DataCitizen turns complex, multi-week enterprise ETL pipelines into a guided, self-service data platform — letting business analysts and citizen engineers ingest, map, enrich and orchestrate production-grade pipelines in minutes, without writing code.

Role
Lead Product Designer & Design Systems Engineer
Platform
Web application · B2B SaaS / enterprise cloud
Stack
React 18, Vite, Material UI v5, custom CSS design tokens
Timeline
16 weeks · audit → design system → prototype → testing
Team
1 product designer (me), 1 product manager, 3 full-stack engineers
Impact
85% less pipeline onboarding time; no code for 90%+ of ETL work

Context & Problem

The enterprise data bottleneck.

In most enterprises, data is trapped between two extremes. Business analysts and citizen data users understand the domain logic — how customer revenue must be calculated or reconciled — but lack the SQL, PySpark and DevOps skills to build pipelines. Central data engineering teams, meanwhile, are buried under backlogs of routine ingestion and schema-mapping tickets, averaging three to six weeks per pipeline.

Traditional workflow

With DataCitizen

Business analyst → Jira ticket → data-engineer backlog (4–6 weeks) → PySpark/SQL scripts → manual QA bug loop

Citizen data user → visual wizard and AI copilot → instant mapping and validation → auto-generated code and a live job

How might we give non-technical domain experts the power to ingest, map and transform complex heterogeneous data independently — while giving enterprise engineering teams the transparency, schema validation and deterministic code generation they demand?

My Role

Design, system and prototype.

  • End-to-end UX/UI design

    Information architecture, persona definition, wireframing, high-fidelity UI and design validation.

  • Design system engineering

    DS Conduit — a dual-mode token architecture in MUI v5 and CSS custom properties, holding 4.5:1+ contrast across high-density data tables.

  • Front-end prototyping

    High-fidelity interactive React components: the ingestion wizard, drag-and-drop attribute mapping, data enrichment and the SwitchieSense copilot.

Constraints

What the design had to survive.

  • High information density

    Users needed 50+ table attributes, data types, confidence scores and preview samples on one screen without cognitive overload.

  • Heterogeneous formats

    Delimited (CSV/TSV), fixed-width (`.dat`), relational databases (PostgreSQL, Snowflake, Databricks Delta Lake) and cloud buckets (AWS S3, Azure Blob).

  • A dual trust barrier

    Business users distrust cryptic regex and code; data engineers distrust black-box transformations. Every visual mapping had to produce inspectable, exportable code.

  • Strict performance

    In-browser schema parsing and table diffing had to stay smooth, with no UI stutter.

Problem Framing

From code to intent.

As a revamp of an existing platform, the direction came from the current product rather than new field research — auditing where analysts got stuck in the legacy tool, and the recurring gaps across its finance, retail and healthcare use cases.

Domain analysts don’t lack logic — they lack syntax. They think in tabular transformations (“trim the postal code”, “look up country by ISO code”), not in AST parse trees or PySpark joins. Forced into code-heavy tools, they fall back to spreadsheet chaos.

Four pain points ran through the legacy workflow:

  • The ingestion cliff

    Connecting cloud buckets required IAM keys, VPC endpoints and JSON schemas — intimidating most non-engineers before they even began.

  • Schema-mapping blindness

    Mapping 100+ source columns in legacy tools meant endless repetitive dropdowns and frequent key-column mismatches.

  • Formula anxiety

    Regex and nested CASE WHEN statements for cleansing led to syntax errors and abandoned workflows.

  • Validation disconnect

    Errors surfaced only after batch jobs failed at midnight, forcing full pipeline rewrites.

Definition

Who it serves, and what it has to do.

  • Priya Sharma · Citizen Data Analyst

    • Goal — merge CRM and billing data for BI.
    • Frustration — waiting four weeks for simple attribute transformations.
    • Behaviour — thinks in spreadsheet formulas and business logic.
  • Marcus Vance · Senior Data Architect

    • Goal — maintain governance and avoid bad ETL.
    • Frustration — writing repetitive glue code for business teams.
    • Requirement — audit trails, clean schemas, reproducible SQL and Spark pipelines.

Jobs to be done

  • New raw files arrive

    Ingest them into the catalog automatically, so nobody configures database connection strings by hand.

  • Reconciling source and target

    Offer automated mapping recommendations with confidence scores, so hundreds of fields can be mapped in seconds.

  • Cleansing messy data

    Let transformation rules be written in plain English, so there is no SQL or regex to debug.

End-to-end workflow

  1. 01

    Ingestion wizard

    S3, Azure, database or CSV.

  2. 02

    Catalog & entity group

    Choose where the data lands.

  3. 03

    Attribute mapping

    AI confidence matching across source and target.

  4. 04

    Data enrichment

    LLM-written rules with a live debugger.

  5. 05

    Validation & profiling

    Report before anything is scheduled.

  6. 06

    Scheduler & code export

    Run the job, or take the generated code.

  • The ingestion manager for a customer data migration, listing catalogs with their database type, description and actions

    Steps 01–02 — the ingestion manager, holding catalogs and their connection types

Ideation

The messy middle.

  1. 01

    Infinite node canvas → guided structured canvas

    Explored
    A free-form infinite node canvas connecting nodes with bezier curves.
    Failed because
    Once schemas passed ~30 columns it devolved into an unreadable spaghetti bowl that overwhelmed analysts.
    Solved with
    A dual-pane, high-density table with drag-and-drop mapping and auto-match confidence badges. Unmapped fields filter instantly, and match confidence reads at a glance.
  2. 02

    15-field monolithic form → 7-step progressive wizard

    Explored
    One long modal holding connector config, parsing options, schema paths and target catalogs.
    Failed because
    64% drop-off from cognitive overload and early validation errors.
    Solved with
    A seven-step ingestion wizard — select source, configure, select files, select catalog, file info, refinery, ingestion settings — with auto-detection of delimiters, headers and data types.
  3. 03

    Formula box → natural-language rule builder

    Explored
    A standard formula input with autocomplete.
    Failed because
    It still demanded exact syntax, which was the barrier the platform existed to remove.
    Solved with
    A multi-modal rule creator: categorized instant functions (string, aggregate, analytical, numeric, date/time, conversion) paired with an LLM rule fetcher and a real-time live-preview debugger.

The Solution

Four moves that removed the code.

  • Guided multi-source ingestion wizard

    Removes the technical barrier to connecting heterogeneous sources.

    • Connector presets for AWS S3, Azure Blob, PostgreSQL and Databricks Delta SQL.
    • Auto-schema inference with dynamic preview cards.
    • In-place profiling — row counts, null rates and type distributions before saving.
  • High-density entity & attribute mapping

    Eliminates column-by-column mapping fatigue.

    • Confidence-scored match engine highlighting direct and fuzzy matches.
    • Visual chips marking primary keys, direct mappings and calculated fields.
    • Collapsible source/target explorer for complex one-to-many and joined schemas.
  • Enrichment & natural-language rule builder

    Enables complex transformations without coding syntax.

    • Categorized function palette — 40+ operations across string, date/time, coalesce, lookup and hashes.
    • Plain-English prompting generates exact SQL and functional transformation rules.
    • Integrated live debugger showing before-and-after data states per row.
  • SwitchieSense conversational copilot

    Reduces onboarding friction and context-switching to docs.

    • Persistent, dockable copilot aware of the active workspace context.
    • Proactive next-step suggestions, such as offering to auto-map remaining attributes.
  • The Choose Source step of the ingestion wizard, showing connector tiles for Amazon S3, Azure Blob, GCP Storage, PostgreSQL, MySQL, Azure SQL, SQL Server and Snowflake

    Ingestion wizard, step 01 — connector presets instead of IAM keys and JSON schemas

  • The entity mapping screen, showing a draft target entity with its selected sources, source table, filter and join conditions, column counts and data profile

    Entity mapping — sources, joins and column counts against one target

  • The attribute mapping table, pairing target attributes with source attributes by transformation type, key column and a confidence score per row

    Attribute mapping — every suggested match carries its confidence score

Design System

DS Conduit.

To hold enterprise-grade rigour across 25+ views, DS Conduit was tuned for high-density, dual-mode data platforms.

Brand primary
#5c9707 DataSwitch green · hover #4d7f06 · tint #a8cc60
Dual surface
Light #eff1f3 / #ffffff ⇄ dark #1d1d1d / #262626
Data table head
Light #E8E8E8 ⇄ dark #3A3A3A / #505050, with a 2px bottom accent
Typography
Montserrat for UI and actions, Fira Code for data and SQL
Semantic pills
Text blue, number emerald, date purple, diff alert amber
  • High-density accessibility

    A standardized --table-header-bg variable across every MUI table held WCAG AA contrast (4.5:1+) in both themes.

  • Monospace data legibility

    Fira Code for column names, SQL expressions and metrics prevented optical misalignment across numerical rows.

  • The Data Validation Studio, comparing source and target columns in a dense table with match-status pills and per-column validation rules

    Data Validation Studio — the density and the status pills the tokens were tuned for

Outcome & Impact

Three weeks to twenty minutes.

Pipeline creation time
Reduced from ~3 weeks to under 20 minutes.
Data mapping accuracy
98.4% first-time mapping success rate.
Engineering backlog
62% drop in routine data-ingestion tickets.
Usability (SUS)
Up from 54 on the legacy CLI to 86.5.
Zero-code adoption
91% of standard pipelines built with no coding.

Reflection

What it taught me.

  • Explainability over magic

    Automated mapping must show its work — confidence scores, underlying logic, generated code. When users understand why a suggestion was made, trust follows.

  • Density is a feature

    Consumer-app minimalism frustrates data analysts. High density with clear hierarchy, keyboard navigation and monospace data typography is far more efficient.

Future roadmap

An interactive, column-level lineage graph across multi-stage ETL jobs, and Figma-style real-time commenting on schema mappings and validation discrepancies.