Skip to content
// SYSTEM OPERATIONAL · AIOPS & MLOPS ARCHITECT

Hi, I'm Yazan.

I'm an AIOps Engineer. I specialize in fusing AI Agent Engineering with MLOps to transform infrastructure, turning manual operations into intelligent, self-healing systems powered by SRE and DevOps best practices.

Over the past 6+ years, I've architected and scaled cloud platforms across high-profile government institutions, Middle Eastern retail conglomerates, and fast-paced tech startups.

Beyond code and cloud platforms, I am an active contributor to open-source software, spending the last 4+ years sharing deep technical knowledge as a writer for the Fedora Project.

Experience 6+ Years
Core Focus AIOps / SRE
Open Source Fedora Writer
Security Air-Gapped
// CAREER TELEMETRY

Professional Background

[ ACTIVE // CURRENT ROLE ] Feb 2026 — Present
LOC // SAUDI ARABIA · ON-SITE

Senior DevOps Engineer

Governata

Architecting and owning the cloud infrastructure for an enterprise data governance SaaS platform, delivering secure, highly available deployments aligned with SDAIA & NDMO framework specifications to high-profile government and institutional clients across Saudi Arabia.

>_
Cloud Infrastructure & Governance: Architected and provisioned production-grade AWS infrastructure using Terraform, ensuring scalable, repeatable, and auditable environments aligned with SDAIA (Saudi Data & AI Authority) and NDMO (National Data Management Office) regulations.
>_
Air-Gapped & Government Deployments: Designed and deployed fully isolated, air-gapped environments for tier-1 government clients, including the Royal Court of Saudi Arabia, the Ministry of Health (MOH) , the General Entertainment Authority (GEA) , and integrations serving the Unified National Platform (my.gov.sa) , meeting strict national security and data sovereignty requirements.
>_
Automation & Operations: Implemented Jenkins as a central automation server and engineered modular Ansible playbooks, eliminating manual ops toil and standardizing multi-client configuration management.
Stack // Terraform AWS Ansible Jenkins Air-Gapped SDAIA & NDMO Compliance IaC
Contract Sep 2024 — Dec 2025
LOC // AMMAN, JORDAN · REMOTE

Senior DevOps Engineer

Majid Al Futtaim

Contracted to modernize DevOps practices at one of the Middle East's largest retail conglomerates, owning mobile release pipelines and cloud infrastructure automation end-to-end.

>_
Mobile CI/CD Pipelines: Designed and implemented CI/CD pipelines for Android and iOS engineering teams using Fastlane, automating full build, test, and release workflows to App Store and Google Play — significantly reducing release friction.
>_
Infrastructure as Code: Managed and maintained multi-environment cloud infrastructure using Terraform and Ansible, guaranteeing consistency across development, staging, and production environments.
>_
Kubernetes Deployments: Automated containerized microservice deployments to AWS EKS clusters using Helm charts, standardizing rollouts and eliminating manual intervention in production.
Stack // Terraform Ansible AWS EKS Helm Kubernetes CI/CD Mobile DevOps
Full-Time May 2022 — Jul 2024
LOC // AMMAN, JORDAN · ON-SITE

Senior DevOps Engineer

Baaz, Inc.

Owned the complete DevOps lifecycle at a fast-scaling social media platform, building the infrastructure foundation to deploy code from commit to production with high velocity and zero downtime.

>_
CI/CD Architecture: Designed end-to-end CI/CD pipelines using Jenkins, reducing manual deployment overhead and enabling low-risk microservice releases.
>_
Automation at Scale: Built automated operations scripts using Python, Bash, and Ansible that accelerated deployment cycles and eliminated human error.
>_
Observability Stack: Deployed a unified telemetry setup combining Zabbix for infrastructure metrics and ELK Stack for MongoDB cluster log analysis for proactive incident detection.
>_
ChatOps & SRE On-Call: Engineered real-time alerting using AWS Chatbot, SNS, and CloudWatch directly to Slack. Acted as primary on-call SRE lead for high-severity production incidents.
Stack // Jenkins Ansible Python AWS ELK Stack Zabbix
Full-Time Oct 2021 — May 2022
LOC // AMMAN, JORDAN · HYBRID

Site Reliability Engineer

Vardot

>_
Backup & Restore Automation: Built automated Jenkins backup and disaster recovery jobs from scratch to safeguard high-traffic enterprise digital platforms.
>_
Multi-Cloud Operations: Managed and monitored infrastructure across AWS, OVH, DigitalOcean, and Platform.sh, optimizing for uptime and performance.
>_
Incident Response: On-call SRE handling production triage, rapid resolution, and blameless post-mortem writing.
Stack // Jenkins AWS OVH DigitalOcean Linux
Full-Time Nov 2020 — Jun 2021
LOC // AMMAN, JORDAN · ON-SITE

DevOps Engineer

NADSOFT

>_
Configured CI/CD automation and utilized Ansible for system & service configuration management.
>_
Implemented server monitoring using Nagios across DigitalOcean and AWS infrastructure.
Stack // Ansible Nagios DigitalOcean AWS
// COMPLETE CAREER VERIFICATION & RECOMMENDATIONS AVAILABLE ON LINKEDIN linkedin.com/in/yazanmonshed
// OPEN SOURCE & COMMUNITY

Extracurricular Activities

Contributor Nov 2020 — Present

Technical Writer

Fedora Project (Writer)

Authoring and publishing deep-dive technical articles on container runtime engines (Podman), Linux system administration, and CLI workflows for the global Fedora developer community.

Podman Linux Containers Fedora Magazine System Admin
// OPERATIONAL PHILOSOPHY

Working with Me

A brief blueprint of how I operate, communicate, and solve infrastructure engineering challenges.

Doc-First & Async Communication

I believe in writing things down. I document architecture decisions (ADRs), SRE processes, and runbooks to reduce context-switching and enable team members to move forward autonomously.

// PRINCIPLE 01

Toil Reduction & Automation

If a manual task has to be performed twice, it belongs in an automation script or declarative IaC configuration. Automation is key to eliminating operational friction.

// PRINCIPLE 02

Actionable Alerting Over Noise

Alerts should point to real customer impact and have clear runbooks. I strive to design high-fidelity, actionable monitoring setups to prevent pager fatigue.

// PRINCIPLE 03

Blameless Post-Mortems

Outages are learning opportunities. I advocate for blameless post-mortems that investigate systemic gaps in architecture or processes, not individual mistakes.

// PRINCIPLE 04
// CONNECT

Elsewhere

Let's work together

Have a project in mind? Let's discuss how I can help with AIOps, AI Agent Engineering, MLOps, or cloud infrastructure.

Reach me on GitHub, X, or LinkedIn.

A.I.D.A. SYSTEM MONITOR

Uplink status: SECURE
Agent role: DevOps Assistant