About this Data Engineer, AWS Glue role at NTT DATA, Europe & LATAM, Branch in USA, Inc.
NTT DATA is a team of more than 190,000 diverse professionals, operating in more than 50 countries throughout the world. The sectors where we have activities include: telecommunications, finance, industry, utilities, energy, public administration and health.
Our mission? Offer technological solutions, business, strategy, development and maintenance of applications, while being a benchmark in consulting. All thanks to the collaboration between teams, the human quality of our people and the fact that we do not conform to what is established, we always seek innovation that brings us closer to the future.
Our essence has led us to the forefront of technology, breaking paradigms and providing solutions that truly respond to the needs of each client. Our talent has led us to be one of the top 6 technology companies in the world.
Because #Greattech, needs #GreatPeople, like you
NTT DATA is looking for high-achieving team players that are quickly adaptable to new challenges and entrepreneurial ventures. We are looking for a Data Engineer, AWS Glue to work with our global client in the U.S. for a remote opportunity in LATAM (including Mexico, Brazil, Chile, Peru).
Overview:
We are seeking a Data Engineer specialized in moving data from heterogeneous source systems into Amazon S3 and processing it across the lake, with AWS Glue as the core engine. The primary focus of the role is ingestion: connectivity to sources, full and incremental extraction, landing zone design, and reliable, repeatable loads at scale. On top of that foundation, capabilities across every layer of the medallion architecture, from raw landing through cleansed and gold layer. Accountable for pipeline reliability, data quality and processing cost efficiency in enterprise or highly regulated environments.
Responsibilities:
- Design, build and maintain ingestion pipelines that move data from relational, file, API and streaming sources into Amazon S3, with full, incremental and CDC based loads.
- Define and maintain the landing and raw zone layout: partitioning, file formats, naming conventions, compression, retention and immutability of source data.
- Onboard new source systems end to end, covering connectivity and network configuration, credential management, extraction strategy, volume and window analysis, and coordination with source owners.
- Guarantee load completeness and integrity through reconciliation against the source, control tables, reprocessing procedures and clear handling of failed or partial runs.
- Implement processing across all layers of the medallion architecture, applying cleansing, standardization, deduplication, conformance and business rules at the appropriate stage.
- Register and maintain datasets in the Glue Data Catalog, keeping schemas, partitions and metadata aligned with Athena, Redshift and Lake Formation.
- Tune job performance and cost, monitoring DPU consumption, execution times and file layout, and refactoring jobs that exceed agreed targets.
- Embed data quality controls into pipelines, define rulesets, handle rejected records and publish quality metrics to stakeholders.
- Orchestrate multi step pipelines with Glue Workflows, Step Functions and EventBridge, including dependency management, alerting and recovery.
- Provide L2 and L3 support for production pipelines: incident analysis, root cause investigation, backfills and continuous improvement backlog.
- Apply and document development standards, naming conventions and promotion procedures across Dev, QA and Production environments, and collaborate with governance and security teams on access policies, lineage capture and metadata quality.
Requirements:
- 5+ years of data engineering experience.
- Must have working proficiency in English and Spanish, both written and spoken.
- Hands-on experience building ETL pipelines with AWS Glue and PySpark.
- Experience moving data from databases and files into Amazon S3.
- Strong knowledge of Python, SQL, Spark, and data transformation.
- Experience designing reliable, scalable, and restartable data pipelines.
- Understanding of S3 data lakes and layered data architectures.
- Experience with AWS workflow, monitoring, and data-quality tools.
- Experience supporting production pipelines and troubleshooting failures.
Nice-to-Have:
- Experience with additional AWS data services such as DMS, Athena, Redshift, or Lake Formation.
- Familiarity with Informatica, Denodo, Terraform, or CloudFormation.
- Experience in enterprise or highly regulated environments.
- Relevant AWS, data engineering, or database certification.
Why NTT DATA?
Empowerment and rewards are the cornerstone of our career development model. We are a young, fast-growing company, with a highly innovative and entrepreneurial spirit, because of this professional experience and growth will be unmatched. Our talent and positive attitude allows us to transform our goals into achievements, and projects into realities.
NTT DATA is committed to hiring and retaining a diverse workforce. We are proud to be an Equal Opportunity/Affirmative Action-Employer, making decisions without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other protected class. NTT DATA is an Equal Opportunity Employer Male/Female/Disabled/Veteran and a VEVRAA Federal Contractor.