Farrukh Islamov
Data Engineer
Summary
Data engineer with 6 years building high-volume ingestion and transformation pipelines on Spark, Airflow and Snowflake. Cut pipeline failure rate from 6% to 0.4% while scaling daily ingestion to 400M events across a logistics data platform.
Experience
- Built an Airflow pipeline ingesting 400M events/day into Snowflake, cutting pipeline failure rate from 6% to 0.4% via dbt-based data quality tests.
- Migrated batch ETL jobs to Spark structured streaming, reducing data latency for downstream dashboards from 4 hours to 12 minutes.
- Partitioned and converted 8 core tables to Parquet on S3, cutting monthly warehouse compute costs 34%.
- Built a Kafka-based CDC pipeline replicating 15 production tables into Redshift, replacing a nightly batch job that took 6 hours.
- Wrote dbt models and tests standardizing shipment data across 3 regional systems, cutting reporting discrepancies to near zero.
- Automated schema-change detection with Terraform-managed infrastructure, reducing pipeline breakage incidents by 60%.
