À propos de ce poste Data Platform Engineer chez Hubble Network
Hubble Network was founded with the intention of delivering on the promise of what Internet-of-Things (IoT) was supposed to be. We're building a global Bluetooth® network dedicated to machine-to-machine connectivity. We differentiate ourselves as the first modem-less and gateway-less, direct-to-satellite network from off-the-shelf Bluetooth® Low Energy chips. Hubble is ideal for applications in logistics, AgTech, and maritime where economies of scale for volume consumer and enterprise asset tracking is a priority. Our goal is to be the first billion-endpoint-connected network in the world.
Hubble is an early-stage, venture-backed startup supported by some of the best investors in the world. In their previous lives, the founding team has been successful in raising $100s of millions in venture funding, developing the Amazon Sidewalk network, launching billions of dollars of space assets, and leading their teams to successful exits, both through acquisition and IPO. We are now looking to bring on talented team members who are the best at what they do to help us make Hubble a reality for the world.
About the Role
This role will be based in San Francisco, CA
This is the founding data hire. You will build Hubble's data platform from the warehouse layer up, make decisions on architecture, determine processes for managing schema and data models, and work with many stakeholders.
We know what we want: a layered lakehouse with Landing, Staging, Warehouse, Mart stages. Each layer has an owner, a purpose, and a quality bar. You’ll work with the Infrastructure team to manage the pipes: Fivetran, Redshift, orchestration, monitoring. You own everything downstream of raw: the transformations, the models, the metric definitions, the quality gates, and the performance bar.
The guiding principle is data democratization. Every person at Hubble should be able to answer their own questions. Success is not how many queries you run for other people, it's how few they need you to run.
You'll be a peer to Platform Engineering and to the team building our internal operations platform. You'll be an embedded consultant to Finance, Product, and Customer Success. You’ll think about how to present a strong story from data.
We expect you to use AI agents as part of how you work, and data work rewards this more than most: schema exploration, model scaffolding, test generation, and documentation are all places where agentic tooling earns its keep if you review the output critically.
Key Responsibilities
- Own the transformation layer end-to-end: Defined models across Staging, Warehouse, and Mart. Star-schema design, SCD2, and data contracts enforced by tests at every layer.
- Build the data layer our internal tools run on: mart tables designed for specific operational workflows, sitting behind a generated internal metrics API. You own the mart and the API's performance characteristics; Engineering owns the framework and the frontend. Hold the line at p95 under 5 seconds on mart queries and under 500ms on metric endpoints.
- Make freshness a commitment, not an aspiration: core data and critical operational metrics streamed in real time with Apache Iceberg, with monitoring that catches breakage before a stakeholder does.
- Instrument the funnels: Design event collection across our different product surfaces, unified into a customer 360 spanning authentication, usage, and billing. Acquisition, activation, engagement, retention, revenue.
Data in space: We collect large amounts of telemetry from our satellites which is critical to mission operations - help our Mission Operations group land this telemetry in dashboards and monitoring systems so they can take action on data immediately.
- Automate revenue reconciliation: Daily comparison of Stripe Platform billing against Hubble usage, with discrepancies flagged before month-end and drill-down to the device and packet level. Turn a multi-day manual close into a few hours.
- Define metrics once: Active devices, packet volume, MRR, contract utilization. Every number has a canonical definition, an owner, and visible lineage from source to dashboard. “Which revenue number is correct?” should stop being a question anyone asks.
- Implement classification in the pipeline: Security defines the tiers and the guardrails. You tag datasets by sensitivity, enforce retention, masking, and anonymization, and keep lineage and access logs auditable. Hubble does not store PII, and the pipeline is where that commitment either holds or quietly fails.
- Build for self-service: Semantic layers that let business users work with customers and revenue rather than join keys, curated datasets, a data catalog, and documentation good enough that people stop asking you.
- Choose boring technology: Fivetran, dbt, Redshift, Metabase are all options - we lean towards buy over build. Bring a strong opinion about which is which, and be able to defend it.
Technical Requirements
Required
- 7+ years building production data systems, with 2+ years at senior or staff scope owning architecture rather than executing someone else's.
- Deep SQL and dimensional modeling: star schemas, slowly changing dimensions, grain discipline. You can explain fact tables and grains.
- dbt (or similar) in production: model organization, macros, incremental strategies, tests as contracts, and CI that blocks bad merges.
- Cloud data warehouse depth, ideally Redshift (or similar): distribution and sort keys, vacuum and analyze behavior, workload management, and the ability to take a slow query apart and make it fast.
- Python or Scala for the specialized pipelines: custom connectors against proprietary APIs, backfills, and reconciliation jobs.
- Event and product analytics instrumentation: you've defined a tracking plan, argued about event taxonomy, and built funnel tables that survive contact with a changing product.
- Data quality as engineering: anomaly detection, freshness monitoring, and incident response with real escalation paths. You've been paged for a broken pipeline and you've fixed the class of bug, not the instance.
- Fluency with AI agents in your own workflow: you already use agentic coding tools to ship real work, and you have a point of view on where they help and where they quietly introduce errors that only show up three models downstream.
Architecture & Scale
- Experience with high-volume machine-generated data, telemetry, events, or IoT at billions of rows, where your partitioning and incremental strategy created a cost-efficient and performant data product.
- Semantic layer and metric definition work: you've built the abstraction that keeps one number meaning one thing across every tool.
- Streaming or near-real-time ingestion (Kinesis, Kafka, Redpanda) with Apache Iceberg or Pinot and a clear view of when it's warranted and when a scheduled batch is the honest answer.
- Ownership of a BI tool as a platform: permissions, curated collections, and the discipline that keeps a Metabase instance from becoming an unmaintained graveyard.
Preferred
- Interest or experience in IoT, satellite systems, telemetry, or geospatial data.
- Billing, usage-based pricing, or revenue reconciliation and recognition workflows.
- SOC 2 or similar compliance work: classification frameworks, access reviews, audit evidence.
- Open source contributions or public technical writing.
What We Look For
Critical Thinking
You interrogate a metric request before building it, because the question behind it is often different from the question asked. You spot the difference between a data quality problem and a source system problem, and you fix upstream when you can. You make decisions with incomplete information and revise them as you learn.
Partnership
You work as a peer to Platform and Engineering, with clear ownership boundaries and no throwing work over the wall. You sit with Finance and Product long enough to understand what they're actually trying to decide. You disagree constructively and commit fully once a decision is made.
Ownership & Pragmatism
You've shipped a data model you later had to tear out, and you learned something specific from it. You explain lineage and trade-offs to people who don't write SQL. You know when to ship the imperfect table that unblocks a decision this week and when to slow down and get the grain right.
Compensation & Benefits
- Salary: $153000-$250,000 (commensurate with experience)
- Comprehensive benefits: Health, Dental, Vision, HSA/FSA options
- Unlimited PTO
- Commuter benefits (if working from HQ)
- Learning & Development allowance
- Health & Wellness stipend
- Sabbatical program
- Work on state-of-the-art satellite systems
The posted compensation range reflects multiple levels within Hubble Network's Talent architecture. Final level placement and corresponding compensation will be determined through the interview and assessment process, taking into consideration factors such as the candidate's relevant experience, qualifications, skills, and geographical location.
ITAR REQUIREMENTS
Hubble is required by the U.S. Government to comply with various space technology export regulations including the International Traffic in Arms Regulations (ITAR). All applicants must be U.S. citizens or lawful permanent residents (“green card holders”) as defined by ITAR (22 CFR §120.15). More information on ITAR can be found here.
Hubble is committed to creating a diverse environment and is proud to be an equal-opportunity employer. Each individual has the right to work in a professional environment that promotes equal employment opportunity and prohibits discriminatory practices, including harassment. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status.
Seattle Washington