Skip to content

Rohan Sitaniya

Senior applied AI engineer · Enterprise systems & agent infrastructure

Finding signal in noise and building reliable intelligent systems.

I build AI agents that automate workflows in regulated production environments, with clear execution boundaries, controlled access, and independent evaluation.

My current focus is ensuring that agents neither grade their own work nor manipulate the tests and rewards used to evaluate them.

Determinism where necessary. Probabilistic models where required.

Selected work

AI agents in enterprise workflows, and the systems that make their work verifiable.

More independent work

Ratchet

An agent-assisted delivery loop for data integrations, with scoped changes, independent evaluation, and human review. A schema score of 0.91 still hid an empty output field; per-field yield made the failure visible.

Vendor vs Valor

A research engine for build-or-buy decisions. Parallel research, source-backed verification, and a challenger pass make the comparison reviewable; a person makes the final decision.

Skills, demonstrated

Architecture, delivery, and evaluation, with the work behind each.

AI architecture and enterprise integration

  • Jitterbit

    Owned the iPaaS Planner-Executor architecture and rollout; built APIM automation for the API lifecycle.

  • American Express

    Built a GenAI governance platform connecting quantitative and qualitative checks to production review and automated model gating.

Retrieval and context engineering

  • Jitterbit

    Workflow-step RAG, semantic chunking, and hybrid retrieval. Context precision improved 60%+; retrieval latency fell 30%. Bounded context and memory reduced LLM spend ~40% while holding quality.

Evaluation and reliability

  • Jitterbit

    Evaluated context precision, context recall, and answer faithfulness/groundedness, using DeepEval checks, Langfuse traces, and human review.

  • American Express

    20+ quantitative and 30+ qualitative governance checks reduced review time by ~60%.

  • Assay

    Protected grading evidence and investigated why passing visible tests can still miss task failure.

  • Ratchet

    Added per-field yield after an improving schema score hid unusable output.

Production engineering

  • Jitterbit

    Checkpointed execution, API integrations, request tracing, and CI/CD-backed Kubernetes services. APIM handled 10K+ daily interactions.

  • American Express

    Shipped transaction categorization for underwriting, reducing bank-statement review from days to hours.

Technical leadership

  • American Express

    Scoped the GenAI governance framework and led a five-member team through delivery.

  • Jitterbit

    Scoped iPaaS with business teams and owned architecture through production rollout. Client interactions and monitoring logs informed engineering decisions.

Journey

From research and financial systems to enterprise agents and independent infrastructure.

May 2026–present

Self-directed

Independent work

Building Assay, Ratchet, and Vendor vs Valor: agent execution, independent evaluation, and evidence-backed decision workflows. Each project has a case study documenting the engineering decisions and their limits.

Oct 2024–Apr 2026

Jitterbit

Senior Applied AI Engineer

Scoped iPaaS with business teams and owned its architecture through production rollout. The Planner and Executor combine workflow-step retrieval, validated handoffs, and checkpointed execution.

Direct client interactions were limited; those conversations and production logs informed retrieval, context management, cost, and latency decisions. Also shipped APIM for natural-language API lifecycle management, handling 10K+ daily interactions.

Aug 2020–Sep 2024

American Express

AI Engineer

Scoped a GenAI governance platform and led a five-member team. Automated gating combined 20+ quantitative and 30+ qualitative checks, reducing governance time by ~60%.

Worked with the commercial new-accounts team to scope bank-statement underwriting. Transaction categorization reached ~85% accuracy and reduced the review cycle from days to hours.

Also built risk and complaint intelligence systems and developed Transformer-based fraud representations.

May–Jul 2019

Algonomy

ML Engineer Intern

Fine-tuned BERT and GPT-2 for personalized query auto-completion over 1.5 million business analytics queries.

May–Jul 2018

National Taiwan University

Research Intern

Analyzed axolotl RNA-Seq data to identify regenerative epigenomic markers and mapped them to the human genome.

2015–2020

IIT Kharagpur

Integrated Dual Degree · B.Tech + M.Tech

Thesis: non-invasive detection of human cancers from multi-genomic TCGA data using deep learning.

Writing

Notes on the choices behind applied AI: retrieval, model adaptation, and production performance.

Dec 12, 2024 · 9 min read

Fine-Tune, Prompt, or RAG? A Decision Guide

Fine-tuning is the most over-reached-for tool in the LLM toolbox. A grounded guide to deciding between prompting, retrieval, and training, before you spend the weeks.

All writing and project findings

Get in touch

Interested in Forward Deployed Engineer and Lead/Senior applied AI roles. If your team is building agents, automating enterprise workflows, or working on reliable AI systems, I’d like to talk.

Technical questions about the work are welcome, too.

© 2026 Rohan Sitaniya