All posts 13

Nora: An AI Executive Assistant That Never Sends as You

An email-native AI executive assistant built on one hard guarantee: she drafts, you send, and every action she takes is cryptographically gated behind DKIM sender verification on the way in and explicit human confirmation on the way out.

AIllmagenticpythonfastapisecurityproduct engineeringcloudflare
Vitals: A Governed Medallion Lakehouse for Healthcare Data

A health-data lakehouse that turns messy FHIR, claims, wearable, and notes data into three trusted outputs — analytics marts, an ML feature store, and a vector index — with de-identification, OMOP standardization, and cross-engine parity testing that caught a real de-id gap.

data engineeringdatabrickslakehousedbtairflowsparkdata qualityhealthcaremlops
Building a Self-Managing MBTA On-Time-Performance Lakehouse (Databricks + GCP)

A live transit-data lakehouse that ingests MBTA real-time feeds, computes on-time performance, and runs, heals, and improves itself — with an agentic layer that writes its own insights and opens pull requests.

data engineeringdatabricksgcplakehousesparkstructured streamingterraformci-cdagentic
Building a Customer Analytics Pipeline with Airflow, dbt and Spark

A production-grade ELT pipeline that automates daily identification of high-value customers using Apache Airflow, dbt-spark, and Apache Iceberg.

data engineeringairflowdbtsparkdockericeberganalyticspipeline
Analytics Engineering with dbt: From Raw Data to Business Intelligence

A comprehensive guide to implementing modern analytics engineering practices using dbt, BigQuery, and Looker Studio. Learn to build scalable data transformation pipelines with dimensional modeling, testing, and deployment strategies.

analytics engineeringdbtbigquerydata transformationdimensional modelinglooker studioELTdata pipeline
Building a Data Pipeline with BigQuery: From Storage to Analytics

A comprehensive guide to implementing a scalable, cost-effective data pipeline and warehouse using Google BigQuery, featuring external tables, partitioning, clustering, and performance optimization with NYC taxi data.

data engineeringbigquerydata warehousecloudanalyticspartitioningclusteringnyc taxi
Orchestrating Data Pipelines with Apache Airflow: A Comprehensive Guide

A practical guide to building and orchestrating robust data pipelines using Apache Airflow, covering local, cloud (GCP), and Kubernetes deployments.

data engineeringairfloworchestrationtutorialdockergcpkubernetesbigquerydata warehouse
Data Pipeline Orchestration using Kestra

Workflow Orchestration using Kestra tool

data engineeringbeginnerstutorialdockerkestraorchestrationbigquerydata warehouse
Simple Data Pipeline

A step by step guide of a simple data pipeline

data engineeringbeginnerstutorialdockerterraformgcppythonpostgresql
The Role of AI in Modern Data Architectures

How artificial intelligence is transforming data systems and workflows

AIdata architectureinnovation
Machine Learning Pipeline Design

Best practices for creating efficient and scalable ML pipelines

machine learningpipelinesMLOps
Getting Started with Data Engineering

An introduction to the fundamentals of data engineering

data engineeringbeginnerstutorial
Welcome to My Professional Website

Introduction and overview of what you can find on this site

introductionabout