IVAN MALIANOV

SENIOR DATA PLATFORM RELIABILITY ENGINEER / PLATFORM SRE

email: malyanoviy@gmail.com • Telegram: @malyan_off • LinkedIn: malyanoff • github: MalianovIY • Almaty, Kazakhstan

PROFESSIONAL SUMMARY


Data platform reliability, DevOps, and AIOps engineer with experience in business-critical systems across telecom, fintech, and data-intensive environments. Strong in OpenShift/Kubernetes operations, Kafka/CDC, PostgreSQL, ScyllaDB, observability, incident response, RCA, release support, security hardening, and operational automation. Builds Python automation and applied AIOps tools for production support, with careful permission boundaries and manual review where automated output affects operational decisions.

CORE SKILLS


Platform Reliability: production readiness, L3 support, incident and problem management, RCA, high availability, backup/restore, post-change validation, runbooks

Containers & DevOps: OpenShift, Kubernetes, Docker, Docker Compose, Linux, GitLab CI/CD, Jenkins, Bitbucket, Ansible, Bash, pipeline-as-code

Data Platforms & Integrations: Kafka, Debezium/CDC, PostgreSQL, ScyllaDB, Spark, HDFS, SFTP, Avro, Parquet, Iceberg, data pipelines, Data Quality

Observability: VictoriaMetrics, Prometheus, Grafana, Alertmanager, Netcool, ELK/EFK, metrics, dashboards, alerts, logging, SIEM log forwarding

AI & Automation: Python, LangGraph, MCP, on-premises open-source LLMs, Jira, Confluence, PostgreSQL, Elasticsearch, Jenkins, SQLite, prompt design, audit trail

Security & Identity: HashiCorp Vault, Ranger, Kerberos, LDAP, Red Hat IdM/IPA, service accounts, SSH keys, secrets management, access review

PROFESSIONAL EXPERIENCE


Beeline KZ — Data Platform Reliability Engineer / DevOps & AIOps Engineer Sep 2024 — Present

Work on reliability, observability, releases, technical acceptance, and L3 diagnostics for a CDC/data middleware platform used in a large BSS transformation.

  • Supported production stabilization across OpenShift, Kafka/CDC, PostgreSQL, ScyllaDB, integration jobs, orchestration flows, and network interactions.
  • Investigated memory and disk exhaustion, network unavailability, configuration defects, post-restart failures, connection exhaustion, OOM events, prepared statement errors, and data pipeline failures.
  • Reviewed architecture diagrams, operational documentation, troubleshooting procedures, change plans, recovery steps, and UAT results before production rollout.
  • Built and migrated monitoring for infrastructure, databases, containers, connectors, certificates, disks, and scheduled jobs using VictoriaMetrics, Prometheus, Grafana, Alertmanager, and Netcool.
  • Validated backup/restore and resilience scenarios for PostgreSQL and ScyllaDB, including failover, switchover, snapshot-based recovery, incremental restore, and full node recovery.
  • Implemented highly available PostgreSQL connectivity using Patroni, etcd, and HAProxy, then verified related data pipeline behavior after the change.
  • Rebalanced OpenShift workload resources, reducing recurring OOM/restart issues and correcting overprovisioned requests/limits.
  • Hardened platform access by helping remove shared and obsolete accounts, identifying exposed secrets, moving technical users to SSH keys, and coordinating access changes with information security.
  • Automated operational checks, configuration corrections, storage cleanup, deployment log collection, and parts of the CI/CD support flow using Python, Bash, Jenkins, and GitLab CI/CD.
  • Localized vendor defects, prepared technical evidence, requested RCA where fixes addressed only symptoms, and coordinated DBA, DevOps, network, security, development, and vendor teams until service recovery was confirmed.
AIOps initiatives
  • Designed and implemented DQ Test-to-Incident Agent, a Python service that compares failed Data Quality checks with Jira history and helps operators find cases where a business incident has not yet been created.
  • Integrated the DQ agent with Jira, Confluence, PostgreSQL, and Jenkins; kept writes limited to a service table and deterministic columns, with manual specialist review before incident creation.
  • Proposed the architecture for Log Pre-Analysis Agent, a Python/LangGraph workflow that turns an incident description into an analysis plan, searches Elasticsearch and Jenkins, compacts logs, and returns likely causes with next checks.
  • Built the log analysis workflow around a custom MCP server, regex and filter tools, SQLite state, and on-premises open-source LLMs; the tool runs read-only and does not perform remediation.
  • Provided technical leadership for two developers on the log analysis MVP through task definition, architecture decisions, result review, and mentoring.
Aitas KZ — DevOps Engineer Feb 2023 — Sep 2024
  • Developed DevOps practices for a large agricultural holding: GitLab CI/CD, Docker/Docker Compose deployments, quality gates, automated tests, security scans, monitoring, and operational automation.
  • Migrated critical on-premises infrastructure, including Kafka, PostgreSQL, GitLab, and Vault, to Yandex Cloud using Hystax, with no data loss recorded in the source CV.
  • Containerized Pentaho Community Edition, simplified dependency management, and supported further updates.
  • Installed and supported GitLab CI/CD, Ansible, HashiCorp Vault, Prometheus, Grafana, and data processing tools.
  • Automated deployment, backups, code quality checks, key management, reverse proxy configuration, and incident support routines.
SberTech / Sber — Python Developer Nov 2021 — Jan 2023
  • Developed a Python web service for second-line support of Hadoop environments, including configuration collection, service status checks, analytics, reports, Grafana data, and automated support actions.
  • Built and integrated Python modules for Ambari and Ranger APIs.
  • Participated in architecture design, legacy refactoring, MVP-to-production delivery, and internal product presentations.
  • Developed cluster health metrics and monitoring; source CV records reduced manual L2 support effort and incident volume after rollout.
  • Mentored and coordinated two interns who completed the program and received team offers.
Divent / Fleetcor — Software Developer / Technical Support Manager Dec 2018 — Nov 2021
  • Developed backend services and automation with Python, Django, PostgreSQL, Celery, RabbitMQ, C++, OpenCV, and Boost.
  • Built software for media processing and a stereo-vision system.
  • Administered and developed Salesforce applications using APEX, jQuery, and Lightning Apps.
  • Managed technical support processes, maintained digital interactive installations, and tested new functionality before rollout.

EDUCATION


Moscow Power Engineering Institute — Master’s, Robotics and Mechatronics · 2016–2018
School 21 / École 42 — C, Python, Linux, Docker & networking projects · Nov 2018–Jan 2022

LANGUAGES


English — C1 · Russian — Native