m
mayankgupta_160

Mayank G

@mayankgupta_160

Data Engineer for Databricks PySpark and ETL Pipelines

Índia
Inglês, Hindi
Algumas informações são exibidas no idioma inglês.
Sobre mim
I help businesses build reliable data pipelines and turn raw data into useful reporting. I specialize in Databricks, PySpark, Spark SQL and Delta Lake, from CDC ingestion to analytics-ready datasets. My experience includes pipelines processing about 22M rows per day, reducing latency from 6+ hours to under 10 minutes, and cutting compute costs by about 18%. I can help with pipeline development, performance reviews, data modeling and Tableau reporting. Message me with your data sources, current challenge and desired outcome so we can define the right scope.... Saiba mais

Habilidades

m
mayankgupta_160
Mayank G
offline • 
Tempo médio de resposta: 1 hora

Conheça meus serviços

Consultoria em Engenharia de Dados
I will make ai understand your data
Aplicações Web Full Stack
I will build you a ai first website

Portfólio

Experiência profissional

Reliance_Retail

Software Developer

Reliance Retail • Período integral

Jul 2025 - Present1 yr 2 mos

Led data pipelines for e-commerce inventory, returns and customer analytics, from CDC ingestion to Tableau reporting. • Reduced pipeline latency from 6+ hours to under 10 minutes. • Processed about 22M rows per day using Databricks, PySpark, Spark SQL and Delta Lake. • Built Oracle GoldenGate and Kafka ingestion with data-quality checks across Bronze and Silver layers. • Optimized incremental loads, reducing compute costs by about 18% while sustaining a 99%+ daily job-success rate. • Developed Gold-layer KPI datasets and improved documentation and governance across 400+ catalogued tables. • Enabled self-service analytics through Unity Catalog metadata and natural-language querying via Databricks MCP. Core tools: Databricks, PySpark, Spark SQL, Delta Lake, Kafka, Oracle GoldenGate, Unity Catalog, Tableau and AWS S3.

Celebal_Technologies

Data Science

Celebal Technologies • Período integral

Oct 2024 - Aug 202510 mos

Built customer analytics and applied AI solutions for business teams. • Defined churn logic and behavioral features for TE Connectivity using Databricks and PySpark. • Conducted exploratory analysis and feature engineering, and trained Random Forest models on AWS SageMaker. • Built a Customer Data Platform with FastAPI REST APIs and Neo4j. • Deployed a knowledge-graph chatbot using Amazon Bedrock to answer business questions. Core tools: Databricks, PySpark, AWS SageMaker, FastAPI, Neo4j and Amazon Bedrock.