- Shadowing the technical team across project planning, implementation, and demos to see how AI initiatives move from concept to deployment.
- Building AI agent prototypes and automations under supervision, and maintaining internal documentation and knowledge resources.
- Researching emerging AI tools and models and presenting findings on their capabilities, limitations, and security considerations.
I BUILD USEFUL AI SYSTEMS.
RAG pipelines, secure LLM tooling, and machine learning systems that turn messy information into fast, reliable software.
About Me
I care whether AI systems hold up outside the demo.
CS major and Math minor in the FSU Honors Program. Right now I'm building a production RAG pipeline at Access to Arabia and researching differentially private graph neural networks at FSU. Most of my projects exist because I wanted to know whether something actually worked: how well an LLM resists 200+ injection attacks, whether a cheaper model can serve the same prompt, what a retrieval stack scores when you measure it honestly.
Teams beat in the programming contest held by FSU ACM in Spring '26.
Passing tests backing Lodestone, my hybrid search engine.
Portfolio projects across AI, security, and data.
Experience
Work Experience
- Designed a RAG pipeline for document ingestion, embeddings, vector search, and natural-language answers over client knowledge bases.
- Implemented REST API architecture for businesses to query internal documents through conversational AI.
- Hardened the AI workflow with RBAC, secure API-key handling, document permissions, and input validation.
- Implemented differentially private graph neural network methods for privacy-aware graph learning.
- Evaluated accuracy-privacy tradeoffs on Cora and PubMed benchmarks using PyTorch and PyTorch Geometric.
- Implemented and debugged features across internal systems and digital services.
- Improved usability and reliability through testing, iteration, and collaborative engineering workflows.
Portfolio
Personal Portfolio
InstaPlanner
AI academic planner that reads deadlines from uploaded files and plain-English notes, then generates conflict-free schedules with one-click calendar export.
Autonomous Short-Video Channel System
Unattended video pipeline that cuts, captions, quality-gates and posts Shorts to 28 channels on four platforms through their logged-in web composers, reaching 1,000,000+ views in the first month.
Job Application Agent
Job-hunting agent that polls 125+ job sources, filters new postings, tailors a per-posting resume and fills the application form with the cover letter I write, then submits once I approve, 100+ applications to date.
Agentic OS
Local control plane that schedules agent jobs through macOS launchd and re-invokes runs under a wall-clock cap with verification scripts between turns, so no run can hang or fake completion, logging 27,500+ graded run records.
Agent Turbo
Installable Claude Code plugin that makes any AI subagent fleet 2.5–3× more token-efficient and 30–50% faster, applying four dispatch rules with zero config, benchmarked on production agent sessions.
Fetchladder
Web-fetching library for AI agents that starts with plain HTTP, escalating to Chromium when a detector proves the page lied, cutting repeat visits from 5.4 to 2.5 seconds via a learned cache.
Deadbolt
Security auditor that sweeps an entire site for leaked keys, unprotected routes, and prompt injection, running 8 phases and 20 injection payloads before it will call anything clean.
SafeDeps
Claude Code skill that grounds AI coding agents in live OSV/CVE data, flagging known-vulnerable dependency versions at 100% precision and recall before they ship.
ToolProof
Tool-calling eval harness that grades LLM tool calls across 40 adversarial cases and 23 tool schemas, using a deterministic JSON Schema validator and zero LLM judges for reproducible failure scoring.
MCP Sentinel
Security scanner for AI tool servers that audits configs and tool manifests for tool poisoning, rug pulls, and supply-chain risk, catching 100% of malicious cases (26 of 26) at 86.7% precision.
Lodestone
Hybrid search engine that combines dense-vector retrieval for semantic matches with BM25 for exact terms, then fuses both result lists with Reciprocal Rank Fusion and re-scores the merged candidates with a cross-encoder.
TrustGate
Claude Code skill that scores every AI change from 0 to 100 by running your tests, lints, and risk-weighted diff analysis, so you know whether to accept it before it touches your files.
ChunkLab
Retrieval benchmarking lab that measures 8 chunking strategies against 800 SQuAD questions with BM25 retrieval, finding a 5.875-point Recall@10 spread between the best strategy (93.125%) and semantic chunking (87.25%).
Equipoise
Fairness audit bench that trains a classifier without sex or race in its inputs, then shows it still fails the EEOC four-fifths rule with selection-rate ratios of 0.3205 for Female/Male and 0.6022 for Non-White/White.
TrueOdds
Probability-calibration lab that fits temperature scaling and isotonic regression in the browser, showing when recalibration makes a model's confidence match its observed accuracy and cutting expected calibration error by 91.5% on the worst-calibrated classifier.
DriftWatch
Drift-detection dashboard that replays a logistic regression over 6,587 time-ordered sessions, catching an injected accuracy collapse from 0.8993 to 0.4400 in the exact batch it starts, and flagging 229 drift alerts on the untouched real stream alone.
PromptArmor
Red-team tool that fires 200+ prompt-injection and jailbreak attacks at a chosen LLM, computes Attack Success Rate from deterministic canary or compliance-token leakage, and ranks models by Resistance Score on a live leaderboard.
ModelRoute
LLM router that scores each prompt's difficulty with a local heuristic classifier, then routes to the cheapest capable model and reports real cost and latency against always calling the top model.
Sidetrack
Claude Code skill that diffs what a coding agent changed against what was asked, scoring 1.000 precision and 1.000 recall across a 41-scenario scope-drift benchmark while cutting what a reviewer has to read by a median of 34.3%.
Syllabus Bot
RAG app that turns any course syllabus into a question-answering bot, pulling policies, deadlines, and grading rules straight from the document.
PulseBoard
Reproducible statistics dashboard that runs real OLS regression on public data, reporting p-values and 95% confidence intervals while filters recompute the inference live in the browser.
Tool Kit
-
AI & Retrieval Systems RAG · embeddings · vector search · Azure OpenAI · Azure AI Search5
-
Machine Learning PyTorch · GNNs · NLP · scikit-learn · Pandas · NumPy6
-
Secure AI Engineering prompt injection testing · RBAC · API-key security · input validation4
-
Backend & Cloud APIs Python · FastAPI · REST APIs · Azure · Supabase · Stripe6
-
Languages & Algorithms Python · C++ · data structures · computer organization · discrete math · competitive programming · GitHub7
Have an idea?
2026 © Created & Designed by Mohammad Jeneidi & Yaser Nazzal. All Rights Reserved.