Pranav K. Sudhir AI Engineer

Cloud data pipelines

Batch and streaming pipelines on AWS at 2 TB+ a day, plus a zero-loss migration off legacy ETL.

Where
HCL Technologies
When
2021 – 2023
Context
Fortune 500 client
Type
Data platform
2 TB+processed daily
−45%job runtimes
99.5%on-time SLA
0data lost in migration

How it works

  1. 1SourcesBatch and streaming
  2. 2IngestKinesis, Lambda, S3
  3. 3TransformGlue and PySpark
  4. 4WarehouseRedshift, tuned
  5. 5Deliver99.5% on time

The problem

The client ran a large daily data load on legacy on-premises SQL Server jobs, with slow runtimes and rising compute costs.

What I built

  • Built batch and streaming ETL/ELT on S3, Glue, Lambda, Kinesis and Redshift with PySpark.
  • Migrated 40+ legacy SQL Server jobs to AWS with parallel-run validation.
  • Tuned partitioning, file compaction and Redshift sort and distribution keys.
  • Orchestrated 50+ production pipelines in Airflow with quality checks and alerting.

Stack

S3GlueLambdaKinesisRedshiftPySparkAirflowJenkins