Share this job
Senior Software Engineer - API & Data Connectors - MarTech/AdTech
San Francisco, California, United States
Apply for this job

PLEASE CLICK HERE TO SEE *ALL* OF OUR JOB OPENINGS!


Senior Software Engineer - API & Data Connectors


This role sits in Data Products, which leverages the company's go-to-market, operations, and SRE functions and adheres to the company's engineering standards and tooling, while running its own product-focused, sprint-based development cycle. It’s the right fit for engineers who understand hyperscaling startups: rapid iteration, resourcefulness, and comfort with shifting priorities.


As a Senior Software Engineer on API & Data Connectors, you will own and extend core data ingestion framework and SDK abstractions. You will design and maintain the custom connector abstraction layers that sit at the core of our data pipeline, alongside a high-throughput Flask runtime service. This role is a hybrid of framework/SDK design and high-performance data engineering—you’ll build native high-throughput connectors (DuckDB, Arrow, Iceberg, S3) while maintaining extensible adapters (PyAirbyte) to bridge hundreds of enterprise data sources into our Causal AI engine.


What You’ll Do:

  • Architect Connector Abstractions: Design, extend, and maintain core Abstract Base Classes (Connector and WatcherConnector ABCs) that establish clean, typed SDK-like interfaces for all data ingestion
  • Build High-Performance Native Connectors: Develop zero-copy, memory-efficient native connectors for local filesystems, Amazon S3, S3 file watchers, and Apache Iceberg utilizing DuckDB relations, Apache Arrow, and Flowbee DataRef primitives.
  • Own the Flask Runtime Service: Build, maintain, and optimize the Flask connector service that exposes standardized runtime operational contracts: catalog, check, discover, and read.
  • Integrate Broad Source Catalogs: Maintain and scale our optional PyAirbyte integration (airbyte>=0.20.0), using PyAirbyte as an optional adapter backend to instantly unlock 300+ external Airbyte data sources.
  • Optimize Data Transfer Semantics: Ensure low latency and high memory efficiency when materializing and passing data references (DataRef) through pipeline execution layers.
  • Developer Experience & Extensibility: Ensure internal and external developers can easily author, test, and deploy custom connectors against our framework with minimal friction and strong typing guarantees.


What Will Help You Succeed:

Engineering Fundamentals & Frameworks

  • 5+ years writing production software, with deep mastery of Python 3.10+ (async/await, strict type hints, ABCs, Pydantic, and clean OOP abstraction patterns).
  • Experience building and operating microservices/APIs using Flask or similar lightweight web frameworks, with a focus on high-throughput JSON/stream serialization.
  • Proven track record of building extensible SDKs, abstract interfaces, or pluggable framework architectures used by other engineers.


Data Ecosystem & Ingestion Depth

  • Deep experience with modern embedded analytical engines and columnar formats: DuckDB relations, Apache Arrow, and Apache Iceberg.
  • Hands-on experience with cloud storage ingestion patterns (Amazon S3, event-driven S3 file watchers, local storage abstractions).
  • Familiarity with standard ELT/ETL lifecycle concepts (catalog, schema discovery, health check, read/stream).
  • Experience integrating or extending data connector frameworks like PyAirbyte (airbyte>=0.20.0), Singer, or Meltano.


Development & Deployment

  • Experience handling complex Python dependency management and packaging (Poetry, pip-tools) for modular framework plugins.
  • Daily comfort with Linux, Docker, Git, automated testing (pytest), and CI/CD pipelines.


Nice to Have:

  • Familiarity with custom event-driven file watching and streaming architectures.
  • Performance optimization experience dropping into Rust or C++ for bottleneck operations.
  • Experience working with low-overhead data reference passing engines (e.g., Flowbee DataRef).


The role is right for you if:

  • You appreciate elegant code abstractions just as much as raw data-throughput performance.
  • You enjoy building the foundational tools, framework ABCs, and SDKs that empower an entire data platform to talk to any data source on earth.


 

Job-3645228

*LI-MY1

#LI-Onsite

Apply for this job
Powered by