🚀 Join Our Data Products and Machine Learning Development Remote Startup! 🚀
Mutt Data is a dynamic startup committed to crafting innovative systems using cutting-edge Big Data and Machine Learning technologies.
We’re looking for a Data Engineer Senior to help take our expertise to the next level. If you consider yourself a data nerd like us, we’d love to connect! 🐶🚀
🚀 What We Do
Leveraging our expertise, we build modern Machine Learning systems for demand planning and budget forecasting.
Developing scalable data infrastructures, we enhance high-level decision-making, tailored to each client.
Offering comprehensive Data Engineering and custom AI solutions, we optimize cloud-based systems.
Using Generative AI, we help e-commerce platforms and retailers create higher-quality ads, faster.
Building deep learning models, we enhance visual recognition and automation for various industries, improving product categorization, quality control, and information retrieval.
Developing recommendation models, we personalize user experiences in e-commerce, streaming, and digital platforms, driving engagement and conversions.
🌟 Our Partnerships
Amazon Web Services
Astronomer
Databricks
🌟 Our Values
📊 We are Data Nerds
🤗 We are Open Team Players
🚀 We Take Ownership
🌟 We Have a Positive Mindset
🔍 Curious about what we’re up to? Check out
our case studies and dive into our
blog post to learn more about our culture and the exciting projects we’re working on! 🚀
Responsibilities 🤓
Own end-to-end pipeline reliability across Bronze, Silver, and Gold layers (PySpark + Delta Lake on Microsoft Fabric).
Build and maintain ingestion notebooks for new retailers and syndicated data partners as we scale.
Harden existing pipelines - idempotent replaceWhere patterns, partition strategies, ZORDER optimization, schema enforcement, and dedup logic.
Run and improve our daily and weekly orchestration through Fabric Data Pipelines and scheduled notebook runs.
Diagnose and resolve runtime issues across the medallion stack - including the parts of Fabric that don’t behave the way the docs say they do.
Manage lakehouse shortcuts, Azure Blob Storage accounts (for non-HNS sources), and data source authentication via Entra / Key Vault.
Partner with the AI engineering team to keep the Gold layer clean and queryable for our RAG, reporting, and attribution use cases.
Contribute to data modeling decisions across our unified retailer schemas, keeping naming conventions consistent across the platform.
Document patterns so the next engineer can move at our pace.
Required Skills 💻
4+ years building data pipelines in production.
Deep PySpark and Delta Lake experience - you’ve shipped real medallion architectures, not just read about them.
Hands-on Microsoft Fabric experience: lakehouses, notebooks, Data Pipelines, OneLake shortcuts. If you’ve worked through Fabric’s quirks (cells not reliably sharing Python variables, shortcut type limitations on non-HNS storage, etc.), that’s exactly the experience we want.
Strong SQL, including window functions, CTEs, and analytical patterns. Comfortable reasoning about partition pruning and predicate pushdown.
Comfortable owning end-to-end pipeline reliability - not just writing the happy path. You think about reruns, backfills, late-arriving data, and what happens at 4am when something breaks.
Azure ecosystem familiarity: Blob Storage, Entra ID, Key Vault, Static Web Apps.
Declarative, sparse code style. You prefer fixing the schema over patching the symptom.
Strong written and spoken English; able to work effectively in a distributed team with overlap to North American business hours.
🎁 Perks
Remote-first culture – work from anywhere! 🌍
AWS, DBT, Google Cloud, Azure & Databricks certifications fully covered
In-Company English Lessons.
Birthday off + an extra vacation week (Mutt Week! 🏖️)
Referral bonuses – help us grow the team & get rewarded!
Maslow: Monthly credits to spend in our benefits marketplace.
✈️🏝️ Annual Mutters' Trip – an unforgettable getaway with the team!