Senior SysOps/DevOps Engineer

CYBERIOUS El Menzah, Tunis
Nouveau Keejob.com
Contrat
CDI
Salaire
2000 - 3000 TND
Secteur
consulting / étude / stratégie
Date
16 septembre 2026

Offre externe provenant de Keejob.com

Cette offre a été publiée sur Keejob.com. Cliquez sur le bouton pour postuler directement sur leur site.

Postuler

Description du poste


Job Description

Senior MLOps / DevOps Engineer

Location

This position is remote

Job Type

Full Time

Relocation Provided

No

Compensation

Commensurate with experience

The Role

We're seeking a Senior SysOps/DevOps Engineer to own platform operations for our AWS infrastructure. You'll be the technical owner of the operations domain: the compute, networking, Kubernetes, and serverless foundations under an AI platform that generates and maintains research on 10,000+ companies, serves institutional clients (including top-tier asset managers), and runs a growing agentic AI workload.

This is a senior, hands-on ownership role on a three-person SysOps/DevOps team — you'll architect and operate the platform alongside teammates focused on CI/CD & security automation and on IAM & internal networking. You'll work directly with the CTO and the AI/full-stack engineering teams, and your work is a direct input to our SOC 2 Type II program and enterprise-client security reviews.

What You'll Work On

Scalable Compute for AI Workloads

  • Architect and operate auto-scaling execution infrastructure for batch AI generation pipelines (queue-based workers, scale-up/scale-down policies, cost-aware scheduling)

  • Own our Kubernetes footprint (EKS), including the vector-database layer (Qdrant) powering retrieval at large scale: capacity planning, upgrades, encryption, performance tuning

  • Own our AWS Lambda architecture for pipeline glue and internal APIs, with cost per invocation tracked and optimized

  • Support the rollout of managed agent runtimes (Amazon Bedrock AgentCore) — networking, isolation, and observability for agentic AI services

Networking & Security Posture

  • Own external and internal network architecture: VPC design, NACLs, security groups, private connectivity, VPN/Tailscale access paths

  • Drive infrastructure hardening for SOC 2 Type II and enterprise security reviews: encryption at rest/in transit, private subnets for data stores, documented access paths

  • Contribute infrastructure evidence to our compliance tooling and to client/investor technical diligence (architecture docs, security posture)

Reliability, DR & Operations

  • Own capacity planning and load validation for platform-scale events (library generation at 10k entities plus concurrent client read load), publishing capacity numbers

  • Design and exercise disaster-recovery: backup strategy, DR tabletops, documented RPO/RTO

  • Own observability for the infrastructure layer (CloudWatch, dashboards, alerting) and act as senior escalation for infrastructure incidents

  • Run OS patching as a standing practice: quarterly baseline cadence with Critical ≤7-day / High ≤14-day remediation SLAs, compliance reported monthly

Cost Engineering

  • Treat infrastructure cost as a first-class metric: instance selection (including Graviton/CPU-optimized fleets for ML inference), storage tiering, right-sizing, and scale-down discipline

  • Partner with the AI team on the unit economics of generation and retrieval — infrastructure choices that cut cost-per-entity without sacrificing reliability

Requirements

Essential

  • 4+ years of SysOps/DevOps/SRE experience, with 2+ years operating production AWS at platform-ownership level

  • Deep AWS expertise: EC2, EKS, Lambda, RDS, S3, VPC networking, IAM fundamentals, CloudWatch

  • Production Kubernetes experience: cluster operations, upgrades, capacity planning, stateful workloads

  • Strong networking fundamentals: VPC design, security groups/NACLs, VPN/private connectivity, TLS

  • Proven auto-scaling and queue-based architecture experience for batch or high-throughput workloads

  • Security-first operational mindset: patching discipline, encryption, least-privilege, audit-ready documentation

  • Scripting/automation proficiency (Python and/or Bash)

  • Fluent English communication skills (written and verbal)

Highly Valued

  • Experience supporting SOC 2 (Type I/II), ISO 27001, or similar compliance programs — evidence collection, control implementation, auditor interaction

  • Vector database operations (Qdrant, Weaviate, Milvus) or other stateful data infrastructure on Kubernetes

  • Experience running infrastructure for ML/AI workloads (GPU/CPU inference fleets, batch pipelines, model serving, Bedrock/SageMaker)

  • Distributed PostgreSQL operations (replication, pgEdge/Aurora, migration off public subnets, zero-downtime changes)

  • Disaster-recovery design and testing (RPO/RTO definition, tabletop exercises)

  • Cost-optimization track record with measurable results (FinOps practices, Graviton adoption, savings plans)

  • Exposure to compliance tooling and security scanning pipelines

  • Experience in fintech, or other environments with enterprise security review processes

Technical Stack You'll Use

  • Cloud: AWS (EC2, EKS, Lambda, RDS/PostgreSQL, S3, CloudWatch, SQS, Bedrock)

  • Containers & Orchestration: Kubernetes (EKS), Docker

  • Data Infrastructure: Qdrant (vector DB on EKS), PostgreSQL/pgEdge, Neo4j

  • Networking & Access: VPC, Tailscale/VPN, Auth0 (SSO/SAML integration points)

  • Compliance & Security: compliance tooling, dependency/secret scanning pipelines, CVE remediation workflows

  • Automation: Python, Bash, CI/CD (GitHub Actions / AWS CodeBuild)

  • AI Platform (what you'll support): LangGraph agentic workflows, Bedrock AgentCore, ONNX/CPU inference fleets

What We Offer

  • High Impact: Own the operations domain outright — your architecture decisions carry the platform's two hardest goals: scale and unit cost

  • Autonomy: Ownership of infrastructure from design to production operation, working directly with the CTO

  • Modern Stack: Infrastructure for cutting-edge agentic AI at enterprise-grade security standards

  • Remote First: Work from anywhere with strong English communication

  • Learning: Rapid exposure to AI-platform operations, compliance engineering, and enterprise fintech infrastructure

Personal Attributes

  • Self-Directed: Thrive with minimal supervision, define your own milestones and deliver

  • Pragmatic: Balance perfection with shipping; incremental infrastructure improvement over big-bang re-architecture

  • Calm Under Fire: Systematic incident response; root causes over quick patches; blameless postmortems

  • Security-Minded: Treat auditability and least-privilege as defaults, not chores

  • Collaborative: Work effectively across AI, full-stack, and QA engineering in a distributed, multi-time-zone team

  • Efficient: Small team means high productivity and smart tradeoffs are essential

To Apply

Please include:

  • Resume highlighting relevant platform/infrastructure ownership experience

  • Brief description of your most impactful infrastructure project (what you built or scaled, the constraint you were solving for — cost, scale, or compliance — and the measurable result)

  • Link to GitHub/portfolio if available (optional but helpful)

Informations complémentaires

Entreprise
CYBERIOUS
Localisation
El Menzah, Tunis
Type de contrat
CDI
Salaire
2000 - 3000 TND
Secteur
consulting / étude / stratégie
Site de l'entreprise
Voir le profil

Pour postuler à cette offre, visitez le site de Keejob.com

Voir l'offre sur Keejob.com