Portfolio of Abhinav Tiwary: AI/ML Engineer and Full-Stack Developer. Featuring open-source projects in autonomous AI agents, leakage-free ML pipelines, and scalable enterprise architectures.
Open Source Projects — Abhinav Tiwary
DecisionForge (Full-stack)
GST Input Tax Credit audit intelligence platform combining BigQuery analytical views with Gemini explanations protected by a numeric-grounding verification guard.
Corporate tax teams face severe financial exposure (18% annual interest and penalties under Section 16(2)(aa) of CGST Act) when reconciling purchase registers against supplier GSTR-2B filings, while standard LLMs hallucinate financial amounts. DecisionForge pairs BigQuery analytical views with a custom numeric-grounding verification guard that audits every generated figure against source invoice data. Non-grounded numbers are automatically discarded in favor of deterministic templates, eliminating financial liability from model hallucinations.
Impact: Audits Section 16(2)(aa) 18% annual interest exposure; classifies invoices into 7 categories with ₹100 tolerance; segments risk across 4 financial tiers (Critical >₹50k, High >₹25k, Medium, Low); dual Python/Node test suites.
GitHub Repository: https://github.com/abhinavtiwary15/Decisionforge
AgentGuard (GenAI / Agentic)
Autonomous AI security operations center orchestrating 5 specialized Azure OpenAI agents for real-time threat triage, investigation, and automated containment.
Security operations teams face alert fatigue and slow manual triage across high-volume enterprise telemetry, requiring automated multi-agent coordination. AgentGuard orchestrates five specialized autonomous AI agents (Sentinel, Oracle, Nexus, Striker, and Herald) to ingest logs, execute RAG investigations over MITRE ATT&CK, route high-impact incidents, and automate containment like IP blocking. Built in pure async Python with typed classes rather than bloated agent wrappers, the system includes resilient local in-memory fallbacks enabling complete offline operation.
Impact: Orchestrates 5 autonomous AI agents; includes 3 attack simulation scenarios (SQL injection, credential stuffing, insider exfiltration); zero paid cloud dependency requirement via local fallbacks.
GitHub Repository: https://github.com/abhinavtiwary15/AgentGuard
Credit Card Fraud Detection (Machine Learning)
Calibrated, leakage-free fraud classification pipeline on 284K+ transactions with SMOTE in-pipeline oversampling, tuning operating thresholds to boost precision to 87.5% at 78.6% recall.
Detecting fraud under severe class imbalance (0.17% fraud rate across 284,807 transactions) is undermined when naive models report 99.8% baseline accuracy while catching zero fraud, and default 0.50 thresholds produce unworkable false alarm rates. This leakage-free pipeline strictly encapsulates RobustScaler and SMOTE oversampling within cross-validation folds and optimizes PR-AUC over deceptive ROC-AUC metrics. By calibrating the decision threshold along the precision-recall curve to 0.9793, XGBoost precision was elevated from 36.1% to 87.5% at 78.6% recall.
Impact: Evaluated on 284,807 transactions with 0.17% fraud prevalence (578:1 imbalance); XGBoost achieved PR-AUC 0.8506 and ROC-AUC 0.9837; decision threshold calibrated to 0.9793, raising precision from 36.1% to 87.5% at 78.6% recall; 4 automated tests.
GitHub Repository: https://github.com/abhinavtiwary15/Credit-Card-Fraud-Detection-ML
VerifyHire (Full-stack)
Full-stack hiring-fraud detection platform scoring candidate authenticity across 5 real signals using deterministic NLP, Gemini Vision, and Fastify microservices.
HR tech platforms frequently present deceptive mockups and simulated features rather than honest, compliant candidate fraud verification pipelines. VerifyHire computes an auditable Candidate Authenticity Score across five weighted verification signals using stylometric resume NLP, timeline plausibility via Gemini 1.5 Flash, periodic interview webcam monitoring, and cross-tenant HMAC-SHA256 fraud hashing. The platform replaced simulated UI stubs with a production-grade Fastify, Redis/BullMQ, and Next.js 14 architecture.
Impact: Evaluates Candidate Authenticity Score across 5 weighted signals (25% Resume, 20% Work, 20% Online, 25% Vision, 10% Cross-tenant); captures video frames every 8s; HMAC-SHA256 salt security; pytest/jest test coverage.
GitHub Repository: https://github.com/abhinavtiwary15/VerifyHire
TrustCall (Full-stack)
Full-stack real-time call fraud detection platform with FastAPI, React, Chrome extension, and Celery alerting, audited to eliminate arbitrary code execution and credential risks.
Real-time voice and video call fraud tools often present polished interfaces while masking severe security vulnerabilities and uncalled alerting hooks. TrustCall connects a Manifest V3 Chrome extension to an async FastAPI backend and Celery workers to compute heuristic DSP audio risk metrics and stream live alerts over WebSockets. A rigorous security audit remediated arbitrary code execution vulnerabilities (pickle deserialization), default database credentials, and broken alerting wires before release.
Impact: Audited and eliminated 4 critical security vulnerabilities (pickle RCE, hardcoded secret, default DB password, missing RBAC); 13 automated tests across async endpoints and alerting pipelines.
GitHub Repository: https://github.com/abhinavtiwary15/TrustCall
ResearchMind (GenAI / Agentic)
Iterative multi-agent research pipeline in LangGraph catching hallucinated sources by capping critic scores and routing revisions back through web search.
Unconstrained multi-agent feedback loops incentivize LLMs to hallucinate plausible-looking peer-reviewed citations when pressured by an automated critic that cannot verify external truth. ResearchMind uncovers and resolves a 66% fake citation rate in naive loops by enforcing hard programmatic score caps (<7) until external citation verification passes cleanly. Failed claims are routed back through live web search with Tavily and BeautifulSoup, preventing fluent prose from masking unverified claims.
Impact: Exposed and remediated 66% fake citation rate in naive feedback loops; enforced hard programmatic score caps (<7) until verification passes; 4 collaborative agents; 9 automated tests.
GitHub Repository: https://github.com/abhinavtiwary15/ResearchMind
Mitra AI (GenAI / Agentic)
Full-stack loneliness intervention platform combining deterministic NLP scoring and UCLA loneliness profiling with swappable LLM coaching and a database-enforced action gate.
Conversational AI tools often trap lonely users in passive infinite-chat loops rather than encouraging real-world social reconnection. Mitra AI couples deterministic NLP relationship profiling calibrated against the UCLA Loneliness Scale with swappable LLM coaching and a database-backed action gate. The FastAPI/SQLAlchemy action gate actively halts conversational chat until users complete and report real-world reconnection actions.
Impact: Evaluates contacts against UCLA Loneliness Scale; outputs top 1-3 prioritized contacts; verified with 34 automated unit and integration tests.
GitHub Repository: https://github.com/abhinavtiwary15/Mitra-Ai
Vakil AI (GenAI / Agentic)
Statute-grounded legal Q&A assistant for Indian MSMEs featuring ChromaDB RAG, section-aware chunking, and an empirically calibrated two-layer hallucination defense.
Delivering legal information to Indian MSMEs requires strict factual reliability, as hallucinated provisions or out-of-scope legal advice carry severe financial liabilities. Vakil AI indexes official legislative statutes directly from the indiacode.gov.in REST API with immutable manifests, utilizing section-aware chunking in ChromaDB. A two-layer defense enforces an empirically calibrated vector distance cutoff (0.52–0.90 in-scope vs 1.11–1.62 out-of-scope) before prompting Gemini to require explicit statutory section citations.
Impact: Covers 4 core commercial Acts with verifiable API UUIDs; embedding distances empirically calibrated for in-scope (0.52–0.90) vs out-of-scope (1.11–1.62); 8 automated tests including two-layer defense verification.
GitHub Repository: https://github.com/abhinavtiwary15/Vakil-AI
Customer Churn Prediction (Machine Learning)
Cost-optimized customer churn prediction system framing classification thresholds around asymmetric business risk, lifting recall from 56.1% to 95.7% and saving $45K on holdout.
Standard churn models optimize abstract statistical metrics like F1 or accuracy at default 0.50 thresholds, neglecting the asymmetric business cost where losing a customer ($1,000 LTV) far exceeds an outreach intervention ($50). Benchmarking eight candidate configurations revealed near-identical ROC-AUC scores (0.845–0.847), confirming that algorithmic choice offered minimal leverage. By deriving a cost-optimal classification threshold of tau* = 0.09 from business unit economics, churner recall surged from 56.1% to 95.7%, saving ~$45,000 on holdout records.
Impact: Evaluated on 7,043 customer records (26.5% churn); 8 candidate configurations yielded 0.845–0.847 ROC-AUC; cost-minimizing threshold tau* = 0.09 boosted recall from 56.1% to 95.7% (capturing 358 of 374 churners), reducing holdout cost by ~$45,000; 4 automated tests.
GitHub Repository: https://github.com/abhinavtiwary15/Customer-Churn-Prediction-ML
Coronary Heart Disease Risk Assessment (Machine Learning)
Clinically-framed cardiac risk assessment pipeline using Platt scaling calibration and SHAP TreeExplainer, prioritizing 89.16% recall over misleading unscaled accuracy baselines.
Clinical prediction pipelines often report misleadingly high accuracy by masking clinical missing-value artifacts (such as zero cholesterol) and producing uncalibrated binary verdicts. This pipeline resolved an 18.7% missing-data artifact and trained a Random Forest optimized for Recall (89.16%) and F2-score (0.8829). Predictions are calibrated via Platt scaling (Brier score 0.105) and explained through SHAP TreeExplainer feature attributions within a Streamlit portal carrying physician equipment disclosures.
Impact: Trained on 918 patient records; resolved 18.7% (172 records) missing cholesterol artifacts; champion Random Forest achieved 89.16% Recall, 0.8829 F2-score, and 0.105 Brier calibration score; 4 automated tests.
GitHub Repository: https://github.com/abhinavtiwary15/Heart-Disease-Prediction-ML
Bengaluru House Price Prediction (Machine Learning)
Leakage-free real estate regression pipeline identifying and resolving target-informed row filtering to achieve converged CV (0.543) and test (0.561) R² scores with price-quartile diagnostics.
Filtering real estate outliers across an entire dataset using target-derived metrics (price_per_sqft) prior to train/test splitting artificially sanitizes test sets, causing test R² to unrealistically beat cross-validation scores. By detecting and fixing this subtle data leakage, Gradient Boosting CV (0.543) and test (0.561) scores properly converged, while log transforms compressed target skewness from 8.06 to 0.86. Residual analysis diagnosed a systematic ~53 Lakh luxury tier under-prediction.
Impact: Reduced target skewness from 8.06 to 0.86; converged Gradient Boosting to CV R² = 0.543 and test R² = 0.561 (remediating invalid pre-fix test R² 0.832 vs CV R² 0.696 divergence); identified ~53 Lakh luxury tier residual gap; 5 automated tests.
GitHub Repository: https://github.com/abhinavtiwary15/House-Price-Prediction-ML
Used-Car Price Estimator (Machine Learning)
End-to-end regression ML pipeline estimating used car prices with leakage-free ColumnTransformer preprocessing, frequency-thresholded cardinality encoding, and 5-model benchmarking.
Estimating used car valuations from real-world listings is corrupted by extreme typo outliers, dirty unit strings, and high-cardinality categorical features. By diagnosing an extreme divergence between cross-validation and test metrics, the pipeline traced and purged a single $26.3M listing typo and pooled 1,590 models into ~216 categories using frequency-thresholded encoding. A leakage-free ColumnTransformer pipeline benchmarks five regression models, deploying a champion Random Forest model (R² = 0.78, MAE ≈ $4,237).
Impact: Achieved Random Forest test R² = 0.78 and MAE ≈ $4,237; pooled 1,590 unique models down to ~216 frequent categories; purged a single $26.3M listing outlier; validated by 8 automated tests.
GitHub Repository: https://github.com/abhinavtiwary15/Car-Price-Prediction-ML