Software Engineer — AI & GenAI Systems
Akhil Uthappa
I build production software and GenAI applications — retrieval-augmented generation, agentic AI, and the data infrastructure underneath them. MS in Computer Science from Boston University.
01 — Experience
Where I've built things
Analytics engineering, applied ML, and GenAI systems — working across data, product, and engineering teams to turn ambiguous problems into shipped solutions.
Manager, Analytics Engineering · American Express
New York, NY
- Build scalable data processing and analytics solutions using Python, SQL, APIs, and cloud platforms, translating complex business requirements into production-ready technical solutions.
- Build automated data workflows and analytical pipelines integrating multiple data sources to improve data quality, reliability, and downstream application performance.
- Design and implement ML/AI solutions using predictive modeling, Generative AI, LLMs, RAG, embeddings, vector databases, and agentic AI frameworks to automate analytical and business workflows.
- Collaborate with Product Managers, Software Engineers, and Data Scientists across the Agile development lifecycle, translating requirements into technical designs and prioritized engineering initiatives.
Software Engineer · Civera
Boston, MA
- Improved software architecture using Spark and AWS, reducing processing times by 40% and operational costs by 30%, increasing system adoption across 9 states.
- Built a computer vision application that identifies PII on voting ballots, while fine-tuning the underlying model on Lightning AI.
- Improved ETL processes and analytics solutions by building custom Python packages and optimizing data structures for election data.
Generative AI R&D Engineer · Cleveland Clinic
Boston, MA
- Developed and optimized a Med-PaLM model with Retrieval-Augmented Generation (RAG) to answer queries about academic research papers, using Hugging Face Transformers and a FAISS vector database for similarity search, with Elasticsearch for data retrieval.
- Trained the model with parallel training on Boston University's Shared Computing Cluster, using A100-80G GPUs with CUDA to accelerate LLM training and development.
Software Engineering Intern, Machine Learning · Staples Inc.
Boston, MA
- Built a text-to-SQL LLM proof of concept on Azure using the Spider dataset, Streamlit, and Snowpark, cutting query-generation time by 40%.
- Engineered an analytics-driven logging system that cut in-store device debugging time by 50%.
Machine Learning Research Assistant · Massachusetts General Hospital
Boston, MA
- Collaborated with stakeholders to design, develop, and test data pipelines using Python and SQL, researching to identify patterns.
- Implemented scalable classifiers for radiation-treated brain cancers with TensorFlow and optimized pipelines with Apache Spark, cutting processing time by 20% and improving report precision.
Earlier
02 — GenAI
Applied GenAI work
Grounded in production work, not demos — retrieval, model integration, and the infrastructure that makes LLM applications reliable.
Retrieval-Augmented Generation
Developed and optimized a Med-PaLM model with RAG to answer queries about academic research papers, using Hugging Face Transformers, a FAISS vector database, and Elasticsearch for retrieval.
Cleveland Clinic — Generative AI R&D Engineer
Agentic AI & Automation
Designs and implements ML/AI solutions using LLMs, RAG, embeddings, vector databases, and agentic AI frameworks to automate analytical and business workflows.
American Express — Analytics Engineering
Text-to-SQL / NL-to-Code
Shipped a proof of concept translating natural language into SQL over the Spider benchmark, deployed on Azure with Streamlit for the interface and Snowpark for query execution.
Staples Inc. — ML Engineering Intern
Vector Search & Embeddings
Built similarity search with FAISS inside a production RAG pipeline, and used Elasticsearch to keep retrieval latency low at query time.
Cleveland Clinic — Generative AI R&D Engineer
LLM Fine-Tuning & Evaluation
Trained and evaluated LLMs with parallel training on GPU clusters (A100-80G, CUDA), and works hands-on with LangChain and prompt engineering for applied AI systems.
Cleveland Clinic · American Express
Applied ML Infrastructure
Trains and serves models with TensorFlow and PyTorch across AWS SageMaker and Spark, including a distributed GAN trained with Horovod across multiple workers.
Civera · Distributed DCGAN
03 — Projects
Selected work
A curated set, not everything — each one links to real code.
04 — Stack
Tools I reach for
Languages
- Python
- TypeScript / JavaScript
- Java
- SQL
- C++
- Go
Frontend
- React
- Next.js
- Tailwind CSS
Backend
- Node.js
- FastAPI · Flask · Django
- REST APIs
- Microservices
- Apache Kafka
AI / ML
- LLMs, RAG & Agentic AI
- LangChain
- Hugging Face Transformers · FAISS
- TensorFlow · PyTorch
Cloud & Data Infrastructure
- AWS · Azure · GCP
- Apache Spark · Databricks · Airflow
- PostgreSQL · MongoDB
- Docker · Kubernetes
Tooling
- Git
- GitHub Actions (CI/CD)
- Elasticsearch
- Grafana · Sentry
05 — Contact
Let's talk
Open to software engineering and AI-focused roles, collaborations, or just a good technical conversation.