2 Candidate Submittal Slots, New High Level PolicyBill Rate - MSP Owner: Rob FintonLocation: White Plains, NY or Fort Lauderdale, FL - Position can be Onsite or RemoteDuration: 6 monthsGBaMS ReqID: 10914533Competencies: 10+ years experience requiredDigital : Machine LearningDigital : DevOpsQuick JD:Senior DevOps Engineer with deep expertise in designing, automating, and operating cloud-based AI/ML platforms.The ideal candidate will have hands-on experience building scalable, secure, and production-grade machine learning environments, with a strong focus on AWS SageMaker, MLOps practices, and modern cloud infrastructure.Experience with Generative AI platforms and services, including Amazon Bedrock, Azure OpenAI, vector databases, RAG architectures, and LLM deployment patterns.Familiarity with ML frameworks such as TensorFlow, PyTorch, MLflow, Kubeflow, or similar technologies.Experience supporting GPU-based workloads and optimizing infrastructure for AI model training and inference.This role will be instrumental in establishing and evolving the organization's AI/ML platform capabilities by delivering robust, automated, and secure cloud infrastructure that accelerates innovation while ensuring operational excellence.Role Summary : Senior DevOps Engineer - AI/ML Platform Engineering (AWS/Azure)This role will be instrumental in establishing and evolving the organization's AI/ML platform capabilities by delivering robust, automated, and secure cloud infrastructure that accelerates innovation while ensuring operational excellence.Senior DevOps Engineer with deep expertise in designing, automating, and operating cloud-based AI/ML platforms. The ideal candidate will have hands-on experience building scalable, secure, and production-grade machine learning environments, with a strong focus on AWS SageMaker, MLOps practices, and modern cloud infrastructure.Key Responsibilities & Qualifications• Extensive hands-on experience designing, implementing, and managing AI/ML infrastructure and MLOps platforms in AWS and/or Azure.• Strong expertise with AWS SageMaker, ML lifecycle management, model training and deployment pipelines, feature stores, model monitoring, and platform automation.• Proven experience building and supporting enterprise-scale MLOps ecosystems, including CI/CD pipelines, Infrastructure as Code (Terraform/CloudFormation/Bicep), containerization, and cloud-native architectures.• Experience integrating AI/ML platforms with modern data ecosystems, including technologies such as Snowflake, data lakes, streaming services, and analytics platforms.• Deep knowledge of cloud services including AWS ECS, EKS/Kubernetes, networking, security, IAM, observability, and high-availability architectures.• Responsible for enabling secure, scalable, resilient, and production-ready AI/ML platforms that support Data Science, Generative AI, and advanced analytics initiatives.• Serve as a trusted technical advisor to engineering, data science, and platform teams, providing architectural guidance, operational best practices, and real-time troubleshooting support.• Demonstrated ability to rapidly assess platform, infrastructure, and deployment challenges and recommend scalable, cost-effective, and secure solutions.• Strong understanding of DevSecOps principles, cloud governance, compliance requirements, and automation strategies for enterprise AI workloads.Excellent communication and collaboration skills, with the ability to bridge the gap between Data Science, Engineering, Operations, and Cloud Infrastructure teams.Preferred Experience• Experience with Generative AI platforms and services, including Amazon Bedrock, Azure OpenAI, vector databases, RAG architectures, and LLM deployment patterns.• Familiarity with ML frameworks such as TensorFlow, PyTorch, MLflow, Kubeflow, or similar technologies.• Experience supporting GPU-based workloads and optimizing infrastructure for AI model training and inference.