Skip to content
View jiangxt2's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report jiangxt2

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jiangxt2/README.md
Thunderkeg / jiangxt2 — building distributed data and AI infrastructure with Spark, Ray, Daft, Gravitino, and OLAP systems.

About me

I'm Thunderkeg. I build and contribute to open-source software for distributed data and AI infrastructure.

My work spans SQL analytics, data connectivity, and machine learning workflows—from distributed training to model delivery and inference. I develop tools and contribute improvements across the Spark, Ray, Daft, and Gravitino ecosystems, with a focus on OLAP databases and lakehouse systems.

Focus & stack

Area What I work on
Distributed analytics Spark SQL, Ray Data, Daft, and reliable OLAP delivery
Data connectivity & lakehouse Doris, ClickHouse, Hive, and governed data access
Model delivery & inference Ray Train, Ray Serve, model bundles, and batch / online inference

Languages: Python · Scala · Java

I'm interested in how multimodal lakehouse systems bring data governance, distributed analytics, and AI workflows together.

Open-source contributions

Apache Spark

My Spark work focuses on SQL functions and expression correctness, especially bitmap and set operations, type conversion, and consistent behavior across SQL, Scala, PySpark, and Spark Connect.

Apache Gravitino

My Gravitino work centers on Doris, ClickHouse, and Lance catalog integrations, with a focus on type compatibility, metadata consistency, and index and partition support. My ongoing work also covers governed Doris reads and integration testing with processing engines.

Ray

My Ray work focuses on Ray Data’s data-source and file-reading APIs, including ORC and Hive reads, Parquet schema inference and source-path metadata, and interoperability with Daft.

Daft

My Daft work focuses on data ingestion and scan efficiency, including ORC reading, Iceberg file statistics, SQL expression translation, and Gravitino catalog compatibility.

Let's connect

I welcome conversations about data connectors, distributed analytics, and model delivery.

WeChat official account QR code. Scan with WeChat to follow.
WeChat Official Account · Scan with WeChat to follow.

Pinned Loading

  1. Pista Pista Public

    Spark SQL runtime for parameterized batch SQL, Catalyst functions, processors, and target-specific OLAP delivery.

    Scala

  2. Tributo Tributo Public

    Ray-native ML SDK for distributed data, training, verified model bundles, and batch/online inference

    Python 4 1

  3. daft-doris daft-doris Public

    Community-maintained Apache Doris datasource for Daft

    Python

  4. ray-doris ray-doris Public

    Community-maintained Apache Doris datasource for Ray Data

    Python

  5. ray-clickhouse ray-clickhouse Public

    Community-maintained ClickHouse datasource for Ray Data

    Python 1 1

  6. Yamazaki Yamazaki Public

    Incubating a governed, evidence-driven AIOps control plane for ClickHouse and Apache Doris.

    Python