Cloud data pipelines
Batch and streaming pipelines on AWS at 2 TB+ a day, plus a zero-loss migration off legacy ETL.
- Where
- HCL Technologies
- When
- 2021 – 2023
- Context
- Fortune 500 client
- Type
- Data platform
2 TB+processed daily
−45%job runtimes
99.5%on-time SLA
0data lost in migration
How it works
- 1SourcesBatch and streaming
- 2IngestKinesis, Lambda, S3
- 3TransformGlue and PySpark
- 4WarehouseRedshift, tuned
- 5Deliver99.5% on time
The problem
The client ran a large daily data load on legacy on-premises SQL Server jobs, with slow runtimes and rising compute costs.
What I built
- Built batch and streaming ETL/ELT on S3, Glue, Lambda, Kinesis and Redshift with PySpark.
- Migrated 40+ legacy SQL Server jobs to AWS with parallel-run validation.
- Tuned partitioning, file compaction and Redshift sort and distribution keys.
- Orchestrated 50+ production pipelines in Airflow with quality checks and alerting.
Stack
S3GlueLambdaKinesisRedshiftPySparkAirflowJenkins