Data Scientist, AI/ML
Job Overview
Data Scientist, AI/ML Today’s complex, fast-paced systems have become a minefield of reliability risks, any of which could cause an outage that costs millions and destroys customer confidence.
That’s why high-availability teams use Gremlin to find and fix reliability risks before they become incidents.
Job Description
**If you don’t think you meet all of the criteria above but still are interested in the job, please apply. Nobody checks every box, we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.
Key Responsibilities
- Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide.
- As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world’s largest organizations where high availability is non-negotiable.
- About the Role of the Data Scientist, AI/ML As a Data Scientist, AI/ML at Gremlin, you will have the opportunity to improve the reliability of the internet at large by turning millions of chaos engineering experiments into automated failure analysis and remediation.
- You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations).
- You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.
- In this role, you’ll get to: Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior Develop scalable data pipelines and feature stores to process, enrich, and serve large volumes of experiment data for both model training and real-time inference Collaborate closely with platform engineers and SREs to integrate AI-driven failure analysis and remediation capabilities directly into Gremlin’s core product Apply advanced techniques, including causal inference, graph ML, time-series modeling, and reinforcement learning, to continuously improve the accuracy and actionability of automated failure analysis Translate insights from millions of chaos experiments into AI-powered features that help customers automatically understand blast radius, pinpoint root causes, and accelerate recovery Research and productionize novel ML approaches, including causal AI and agentic systems, that turn raw chaos experiment data into automated, reliable remediation strategies We’ll expect you to have: Experience as a self-driven and collaborative problem solver with strong communication skills 5+ years professional experience building and productionizing machine learning, ideally for distributed systems, infrastructure, or DevOps and SRE use cases with more overall years of experience in software development.
- Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning Experience building data pipelines and feature stores that support both offline training and real-time inference Experience with agile development environments and practices Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices Comfort partnering with platform engineers and SREs to turn research into shipped product features Strong at breaking down ambiguous problems into concrete actions and milestones Bonus Experience: Experience with chaos engineering, site reliability engineering, or distributed systems Background in agentic AI systems or large-scale causal inference in production Experience standing up MLOps tooling such as model serving, monitoring, or feature store infrastructure Working in Remote first environments Has been on-call and participated in an incident management program *The role does not offer sponsorship employment benefits.
- We set the standard for reliability and equip leading organizations with the mindset and expertise needed to drive reliability improvements that move the world forward.
Required Skills and Qualifications
- You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations).
- You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.
- In this role, you’ll get to: Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior Develop scalable data pipelines and feature stores to process, enrich, and serve large volumes of experiment data for both model training and real-time inference Collaborate closely with platform engineers and SREs to integrate AI-driven failure analysis and remediation capabilities directly into Gremlin’s core product Apply advanced techniques, including causal inference, graph ML, time-series modeling, and reinforcement learning, to continuously improve the accuracy and actionability of automated failure analysis Translate insights from millions of chaos experiments into AI-powered features that help customers automatically understand blast radius, pinpoint root causes, and accelerate recovery Research and productionize novel ML approaches, including causal AI and agentic systems, that turn raw chaos experiment data into automated, reliable remediation strategies We’ll expect you to have: Experience as a self-driven and collaborative problem solver with strong communication skills 5+ years professional experience building and productionizing machine learning, ideally for distributed systems, infrastructure, or DevOps and SRE use cases with more overall years of experience in software development.
- Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning Experience building data pipelines and feature stores that support both offline training and real-time inference Experience with agile development environments and practices Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices Comfort partnering with platform engineers and SREs to turn research into shipped product features Strong at breaking down ambiguous problems into concrete actions and milestones Bonus Experience: Experience with chaos engineering, site reliability engineering, or distributed systems Background in agentic AI systems or large-scale causal inference in production Experience standing up MLOps tooling such as model serving, monitoring, or feature store infrastructure Working in Remote first environments Has been on-call and participated in an incident management program *The role does not offer sponsorship employment benefits.
- We recognize that salary varies from person to person depending on level of experience and we welcome direct conversations about it.
- The final offer will vary based on assessment of a candidate's skills and ability and our budget and market data.
Benefits and Perks
- Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning Experience building data pipelines and feature stores that support both offline training and real-time inference Experience with agile development environments and practices Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices Comfort partnering with platform engineers and SREs to turn research into shipped product features Strong at breaking down ambiguous problems into concrete actions and milestones Bonus Experience: Experience with chaos engineering, site reliability engineering, or distributed systems Background in agentic AI systems or large-scale causal inference in production Experience standing up MLOps tooling such as model serving, monitoring, or feature store infrastructure Working in Remote first environments Has been on-call and participated in an incident management program *The role does not offer sponsorship employment benefits.
- Gremlin offers competitive total compensation packages including 401k Matching, Equity and other benefits such as flexible time off and paid company holidays.
USA Jobs Today role summary
Role Summary
Data Scientist, AI/ML at Gremlin is a remote United States position and is listed as Full Time.
Data Scientist, AI/ML Today’s complex, fast-paced systems have become a minefield of reliability risks, any of which could cause an outage that costs millions and destroys customer confidence.
Review the employer's original description and official application page before applying; USA Jobs Today organizes source-supported details for readability without changing stated requirements.
Application Checklist
- Review the stated qualifications, including: You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations). You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.
- Confirm that the official application page still lists the role as open.
- Verify the work location and any United States eligibility or work-authorization requirements.
- Tailor your resume to the responsibilities and qualifications stated by Gremlin.
- Apply only through the official link shown on this page and never pay USA Jobs Today to submit an application.
Original job information is provided by the hiring organization or listed source. USA Jobs Today organizes source-supported details for readability without inventing salary, benefits, qualifications, sponsorship, or remote eligibility.
Work Location and Schedule
This role is listed as Remote USA with location information shown as Remote USA. The employment type is Full Time.
About the Company
Gremlin is the organization connected with this listing. USA Jobs Today displays this opportunity for job discovery only, so applicants should verify company details, application instructions, and eligibility on the official employer website.
Application Notes
This job was reviewed for USA-only relevance. Always apply through the official employer website, review the full job details carefully, and avoid sharing sensitive personal or payment information outside a trusted application process.
Report this job if it looks expired, suspicious, inaccurate, or unsafe.More USA job search options
Related resources for this job seeker
Use these links to browse more USA jobs, compare related categories, prepare your resume, and read USA job search guidance before applying.
