Sundus

MLOps Engineer

Sundus

United Arab Emirates

Accepting Applications Full-time On-site LinkedIn
Posted 1 month, 1 week ago 5 views 0 applications
Job Description

Job Code: 6315

Job Title: DevOps / MLOps Engineer (AI \& LLM Platforms)

Location: Abu Dhabi

Contract: 1 year and renewable

Experience: 7+ years

Role Purpose The DevOps / MLOps Engineer is responsible for the setup, automation, and maintenance of infrastructure and deployment pipelines for AI/ML and microservices-based applications. This role focuses on enabling efficient development, testing, and deployment of AI solutions, including LLM workloads, while ensuring system reliability, scalability, and performance.

Key Responsibilities

  • Infrastructure Support \& Environment Management
  • Set up and maintain compute infrastructure, including GPU-enabled environments.
  • Configure and manage Linux-based systems for development and production environments.
  • Provisioning and configuration of cloud and on-prem infrastructure.
  • Monitor system resources and assist in performance tuning and optimization.
  • Containerization \& Deployment
  • Build and manage containerized applications using Docker.
  • Deploy and manage applications on Kubernetes clusters under guidance from senior engineers.
  • Creating deployment configurations, Helm charts, and environment setups.
  • Support scaling and orchestration of microservices and AI workloads.
  • CI/CD Pipeline Implementation
  • Develop and maintain CI/CD pipelines for application and AI model deployment.
  • Automate build, test, and deployment processes using tools like Azure DevOps, GitHub Actions, or Jenkins.
  • Ensure smooth promotion of code and models across environments (dev, test, prod).
  • Troubleshoot pipeline failures and deployment issues.
  • MLOps \& AI Deployment Support
  • Deploying machine learning models and LLM-based services.
  • Integration of AI components into production systems.
  • Contribute to model versioning, monitoring, and lifecycle management.
  • Work with AI engineers to operationalize RAG pipelines and inference services.
  • Monitoring, Logging \& Issue Resolution
  • Implement and maintain monitoring and logging solutions (e.g., Prometheus, Grafana, ELK).
  • Track application performance, system health, and availability.
  • Respond to incidents, troubleshoot issues, and escalate when required.
  • Assist in root cause analysis and continuous improvement.
  • Automation \& Scripting
  • Write scripts (Python, Bash) to automate repetitive operational tasks.
  • Support Infrastructure as Code (IaC) initiatives using tools like Terraform or ARM templates.
  • Improve operational efficiency through automation and tooling.
  • Collaboration \& Support
  • Work closely with Senior DevOps/MLOps Engineers, AI Engineers, and Development teams.
  • Support developers in environment setup, debugging, and deployment processes.
  • Follow DevOps and MLOps best practices and continuously improve operational workflows.

Required Skills \& Qualifications

  • Bachelor s degree in Computer Science, Engineering, or related field.
  • 7+ years of experience in DevOps or platform engineering roles.
  • Basic to intermediate experience with Linux system administration.
  • Hands-on experience with Docker and containerization.
  • Familiarity with Kubernetes (deployment and basic management).
  • Experience with CI/CD tools (Azure DevOps, GitHub Actions, Jenkins, etc.).
  • Basic understanding of cloud platforms (Azure, AWS, or GCP).
  • Scripting skills in Python, Bash, or similar.
  • Understanding of version control systems (Git).

Preferred Skills

  • Exposure to AI/ML model deployment and MLOps practices.
  • Familiarity with LLM deployment concepts and tools.
  • Basic knowledge of GPU environments and high-performance computing.
  • Experience with monitoring and logging tools (Prometheus, Grafana, ELK).
  • Knowledge of Infrastructure as Code (Terraform, ARM templates).
  • Understanding of microservices architecture.

Key Performance Indicators (KPIs)

  • Deployment success rate and pipeline stability.
  • System uptime and availability.
  • Resolution time for incidents and issues.
  • Efficiency of CI/CD processes.
  • Infrastructure utilization and basic cost optimization.
  • Support effectiveness for development and AI teams.

Stakeholders \& Reporting

  • Reports to: Senior DevOps / MLOps Engineer / Platform Lead
  • Key Stakeholders:

+ AI Engineers \& Data Scientists + Backend \& Frontend Developers + DevOps / Platform Team + QA \& Release Management Teams

Max 3 MB. JPEG or PNG recommended.

Choose which emails you want from Jobaro. You can change this later in Settings.

About Company
Share this job