Sobre esta vaga de Materials Science Analyst na Gramian Consulting Group
About Gramian
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
Role Overview
We are looking for a materials science professional to contribute to an advanced AI research project focused on building rigorous STEM coding datasets. The role combines materials science expertise, Python-based scientific computing, and AI evaluation, with a focus on creating well-structured scientific problems, implementing verified solutions, and developing tests that accurately distinguish correct from incorrect model outputs.
Responsibilities
Design scientific coding tasks with one main problem and at least 3 logically connected sub-problems.
Implement verified Python solutions with complete unit test coverage.
Create discriminative test cases that distinguish correct from incorrect AI-generated outputs.
Perform quality control checks using the Central Task Platform (CTP), including Tier 1 structure checks and Tier 2 quality rubrics.
Revise tasks and solutions based on QC feedback.
Optimize tasks against Pass@K evaluation criteria across multiple LLM judges, including GPT, Gemini, and Nemotron.
Validate scientific correctness, determinism, well-posedness, and expected outputs.
Maintain a low rework rate and high first-submission quality.
Participate in project reviews, feedback sessions, and standups during required overlap hours.
CONTRACT: Freelance / Contractor
COMMITMENT: Full-time commitment; overlap requirements to be confirmed
LOCATIONS: Remote; eligible locations to be confirmed
PROCESS: Not specified
Requirements
Design scientific coding tasks with one main problem and at least 3 logically connected sub-problems.
Implement verified Python solutions with complete unit test coverage.
Create discriminative test cases that distinguish correct from incorrect AI-generated outputs.
Perform quality control checks using the Central Task Platform (CTP), including Tier 1 structure checks and Tier 2 quality rubrics.
Revise tasks and solutions based on QC feedback.
Optimize tasks against Pass@K evaluation criteria across multiple LLM judges, including GPT, Gemini, and Nemotron.
Validate scientific correctness, determinism, well-posedness, and expected outputs.
Maintain a low rework rate and high first-submission quality.
Participate in project reviews, feedback sessions, and standups during required overlap hours.