Skip to content
View SA01's full-sized avatar

Block or report SA01

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please donโ€™t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this userโ€™s behavior. Learn more about reporting abuse.

Report abuse
SA01/README.md

About Me - Suffyan Asad

Hi there ๐Ÿ‘‹

I am a Senior Data Engineer and Data Architect with 12+ years of experience building large-scale data platforms, distributed pipelines, and analytical systems.

My core stack includes Python, PySpark, Apache Spark, AWS, Databricks, ClickHouse, PostgreSQL, and SQL. I have worked on platforms processing billions of records, with a focus on data architecture, performance, reliability, financial analytics, and AI-enabled data applications.

I am also a Databricks Certified Data Engineer Associate.

This GitHub mostly contains projects, experiments, and code supporting my technical articles on Medium.


๐ŸŒŸ Career Highlights

  • Lead data platform modernization and large-scale financial data engineering work at MobSquad / CoreStack.
  • Designed and benchmarked a sharded ClickHouse analytics platform supporting billion-row workloads.
  • Built and optimized Spark and PySpark pipelines processing billions of records.
  • Reduced large-scale Spark processing time by 30% and infrastructure cost by 15%.
  • Built anomaly detection systems that processed billions of records and identified a cost spike that helped Zoom save approximately $600K.
  • Built forecasting, cost attribution, recommendations, and multi-cloud analytics platforms.
  • Designed automated validation and reconciliation pipelines for financial data.
  • Work with Claude Code, MCP, ReACT, and RAG for AI-assisted engineering and data applications.
  • Led and mentored distributed engineering teams.

โœ๏ธ Technical Writing

I publish hands-on articles about:

  • Apache Spark and PySpark
  • Data architecture and performance
  • ClickHouse
  • dbt
  • Data warehousing
  • MCP and AI-enabled engineering

Medium: @suffyan.asad1


๐Ÿ’ฌ Connect


๐Ÿ“š Education

  • MS in Business Analytics | The George Washington University
  • BS in Computer Science | FAST-NU

Pinned Loading

  1. spark-read-jdbc-tutorial spark-read-jdbc-tutorial Public

    This repository contains the code and examples for my article on Medium, which explains how to parallelize reading data from JDBC sources in Apache Spark.

    Python 7 5

  2. dbt-tutorial dbt-tutorial Public

    This repository contains the code and project files for my article on Medium, which serves as an introduction to DBT (Data Build Tool) and how to build data transformations with it.

    Dockerfile 7 3

  3. clickhouse-docker-compose-cluster clickhouse-docker-compose-cluster Public

    An example three node one replica Clickhouse cluster

    Dockerfile 3 2

  4. spark-custom-datasource-tutorial spark-custom-datasource-tutorial Public

    Contains the code and examples for my article on Medium, which explains how to create a custom JDBC read-only data source in Apache Spark 3

    Scala 4 1

  5. spark-window-functions spark-window-functions Public

    This repository contains the code and examples for my article on Medium, which provides an introduction to Window Functions in Apache Spark.

    Python 2

  6. docker-spark-cluster docker-spark-cluster Public

    Forked from mvillarrealb/docker-spark-cluster

    A simple spark standalone cluster for your testing environment purposses

    Dockerfile 7 7