Hi there ๐
I am a Senior Data Engineer and Data Architect with 12+ years of experience building large-scale data platforms, distributed pipelines, and analytical systems.
My core stack includes Python, PySpark, Apache Spark, AWS, Databricks, ClickHouse, PostgreSQL, and SQL. I have worked on platforms processing billions of records, with a focus on data architecture, performance, reliability, financial analytics, and AI-enabled data applications.
I am also a Databricks Certified Data Engineer Associate.
This GitHub mostly contains projects, experiments, and code supporting my technical articles on Medium.
- Lead data platform modernization and large-scale financial data engineering work at MobSquad / CoreStack.
- Designed and benchmarked a sharded ClickHouse analytics platform supporting billion-row workloads.
- Built and optimized Spark and PySpark pipelines processing billions of records.
- Reduced large-scale Spark processing time by 30% and infrastructure cost by 15%.
- Built anomaly detection systems that processed billions of records and identified a cost spike that helped Zoom save approximately $600K.
- Built forecasting, cost attribution, recommendations, and multi-cloud analytics platforms.
- Designed automated validation and reconciliation pipelines for financial data.
- Work with Claude Code, MCP, ReACT, and RAG for AI-assisted engineering and data applications.
- Led and mentored distributed engineering teams.
I publish hands-on articles about:
- Apache Spark and PySpark
- Data architecture and performance
- ClickHouse
- dbt
- Data warehousing
- MCP and AI-enabled engineering
Medium: @suffyan.asad1
- LinkedIn: Suffyan Asad
- Medium: @suffyan.asad1
- MS in Business Analytics | The George Washington University
- BS in Computer Science | FAST-NU

