Senior Data Engineer

Switch Intelligence Inc REMOTE 2026-09-11

We are looking for a strong Data Engineer who is also an excellent Python developer to think through already existing architecture, improve, modify or redesign, and build the data backbone of the Switch Intelligence Platform — a multi-tenant, source-agnostic data streaming and intelligence system. You will build scalable pipelines that connect and integrate with multiple data sources, APIs, databases, and external services, which will feed a Neo4j knowledge graph, a postgres, lakehouse, and an AI-driven signal loop. The ideal candidate has 5+ years of professional experience, strong software engineering fundamentals, and hands-on experience building production-grade, multi-tenant data systems.

Key Responsibilities

  • Design, build, and maintain scalable, reliable, production-grade data pipelines for batch and real-time (streaming) processing.
  • Develop integrations with multiple data sources, APIs, databases, SaaS platforms, and third-party services.
  • Build reusable, tenant-aware ingestion and transformation frameworks that support per-tenant object and field customization without sacrificing data integrity.
  • Implement Change Data Capture (CDC) and reliable write-back flows to and from source systems, with idempotency and exactly-once/at-least-once semantics as appropriate.
  • Develop robust ETL/ELT pipelines with validation, cleansing, normalization, transformation, and enrichment workflows.
  • Build API-driven data services and integrations using Python.
  • Contribute to schema evolution tooling: declarative (YAML-based) schema definitions, validation, diffing, and code generation.
  • Containerize and deploy data services using Docker and Kubernetes (GKE).
  • Design systems for handling large volumes of structured and unstructured data in the pipelines.
  • Implement reliable error handling, retries, logging, monitoring, and data-quality checks.
  • Work with relational databases and optimize data storage and retrieval.
  • Build asynchronous and event-driven data processing workflows.
  • Build automated testing, CI/CD, deployment, and monitoring for data applications.
  • Collaborate closely with AI/ML, Backend, and Engineering teams.
  • Research and evaluate new data technologies, APIs, frameworks, and integration patterns.
  • Troubleshoot complex production data and integration issues.
  • Collaborate on graph translation in the data pipeline: ingest data from source systems and translate it into a governed knowledge graph model.
  • Build and operate metering and usage-aggregation workloads with windowed aggregation, late-arriving data handling, and billing-grade correctness.

Ideal Candidate

The ideal candidate combines strong Data Engineering expertise with excellent software engineering and Python skills. You should be comfortable taking an ambiguous data integration requirement through to the pipeline architecture, writing the Python code, connecting the external systems, building the pipeline, testing it, and deploying it.

You should be able to reason about real architectural trade-offs — for example, when to use typed columns versus JSONB versus extension objects for per-tenant flexibility, and why a metering engine or money chain cannot be built on a fully generic EAV model. We value honest technical pushback grounded in correctness, integrity, and operability.

You should be a hands-on engineer and strong coder, rather than someone focused primarily on data analysis or low-code ETL tools.

Nice to Have

  • Experience with Apache Airflow, Dagster, Prefect, or similar orchestration frameworks.
  • Experience with Redis or similar caching/NoSQL technologies.
  • Experience with Neo4j or another graph database in production.
  • Experience with metering, usage aggregation, or other systems with strict correctness guarantees.
  • Experience with schema evolution/migration tooling or code generation from declarative schema definitions (e.g., YAML/DSL-driven codegen).
  • Experience with Spark, PySpark, or other distributed data processing technologies.
  • Experience with Google Cloud Platform (GCP) — our target platform — or AWS/Azure.
  • Experience with Google Cloud Storage, S3, or similar object storage systems.
  • Experience with data provenance, lineage, or field-level assertion/audit models.
  • Experience with Ansible, SonarQube, Artifactory, or similar DevOps tooling.
  • Experience with data observability and monitoring platforms.
,[5+ years of professional Data Engineering experience., Strong Python development skills, including excellent knowledge of OOP, clean architecture, testing, and writing maintainable production code., Strong experience designing and implementing ETL/ELT data pipelines., Event streaming / messaging experience with Kafka, Google Pub/Sub, RabbitMQ, or similar — streaming is core to this role, not optional., Hands-on experience integrating with REST APIs, third-party services, databases, and external data sources., Experience building multi-tenant data platforms or SaaS integrations, including tenant isolation and per-tenant customization., Experience with PostgreSQL and strong SQL skills, including working with JSONB for semi-structured/sparse fields., Hands-on experience ingesting from Salesforce, HubSpot, Dynamics, or a comparable commercial source system., Hands-on experience with at least one modern warehouse or lakehouse, such as Snowflake, Databricks, or BigQuery., Experience with Change Data Capture (Debezium or similar) and reliable write-back patterns to source systems., Experience with data processing frameworks and/or distributed data processing., Experience with Docker and Kubernetes., Strong understanding of Git and software development workflows., Experience with CI/CD pipelines (GitHub Actions, Jenkins, Bitbucket Pipelines, or similar), automated testing, deployment, and monitoring for data applications., Experience building and maintaining production-grade data applications., Strong understanding of data modeling, data quality, validation, schema evolution, and pipeline reliability., Ability to independently troubleshoot and resolve complex technical problems.] Requirements: Python, ETL Tools: GitHub, SharePoint, GitLab.