Manny Castillo

Senior Software Engineer in Test

Backend & API Test Automation · Data Pipelines · Python · Distributed Systems · CI/CD

· linkedin.com/in/mfcastillo

Summary

Senior SDET with 14+ years testing backend services, data pipelines, and large-scale distributed systems. API and integration automation in Python, Airflow pipeline validation with nightly runs, and performance work against the services underneath — plus the web and mobile E2E layer on top of it in Playwright and Maestro. Builds the test frameworks and internal tooling other engineers depend on, across enterprise SaaS, gaming infrastructure, storage, and cybersecurity.

Skills

Experience

Lead SDET, Medical Team — Backend, Pipeline & UI Automation — Sorcero Inc.

Remote · 2020 — 2026

  • Owned test strategy and end-to-end automation for the Medical Team, and mentored engineers across teams.
  • Automation reaching past the browser into the REST and GraphQL services, so failures are attributed to the layer that actually broke.
  • Tests asserting that Airflow DAGs running in Google Cloud Composer transform and land the data the product depends on — not just that the pipeline ran, but that what came out the other side was correct.
  • Scheduled nightly validation against the data pipelines, so breakages introduced during the day are caught overnight and triaged in the morning instead of being discovered in the product.
  • Owned validation of model and LLM outputs for the MLOps team — non-deterministic results need a different testing approach than an assert on a fixed value. Also integrated LLM-assisted tooling into day-to-day test engineering.
  • Moved suite execution onto Kubernetes so runs scale horizontally instead of serialising on one machine.
  • Led the migration to GitLab CI with secure workflows, and wired the suites to execute as part of deployment — on a dedicated dev/QA environment during development, then again on each promotion to staging and to production. The suites gated those promotions in practice: a release waited on a clean run, enforced as a process step rather than as an automated block in the pipeline.
  • JMeter and Locust tooling in Python identifying front-end and back-end bottlenecks, feeding directly into optimization work.
  • Dashboards plus automated triage workflows, so a red run arrives with a first guess at why rather than a wall of logs.
  • End-to-end and UI smoke coverage in TypeScript across two major customer-facing platforms, built alongside the QA team rather than in isolation, replacing a manual regression pass that consumed hours every release cycle.
  • Extended automated coverage to the native mobile app, writing E2E flows for iOS in Maestro so mobile regressions surface in the same cycle as web.
  • The full smoke suite runs against core platform functionality on each release, covering the paths a user hits first.

Software Engineer in Test — Sony PlayStation

Austin, TX · 2015 — 2020

  • Built and enhanced Daruma with reusable helper libraries and AWS tooling; it became shared infrastructure for automated testing beyond the immediate team.
  • Wrote pytest-based tests and the test infrastructure around them for microservices spanning the PlayStation Now service platform, including datacenter streaming validation.
  • Boto3 automation to provision and validate live game streams against new European datacenters ahead of launch.
  • Led QA initiatives supporting SRE through large-scale datacenter migration testing for the expansion of PlayStation Now infrastructure, validating capacity ahead of launch to 700,000+ concurrent users.
  • Designed the backend architecture for an in-house dashboard test runner — the system QA teams used to launch suites and read results.
  • Extended Prometheus exporters in Go so that quality and reliability signals were measurable rather than anecdotal.

Software QA Engineer (Infrastructure Validation) — Symantec

Boston, MA · 2014 — 2015

  • Test plans and Python automation cover enterprise security products.
  • Configured and managed the virtualization layer the appliance testing depended on.
  • Found and helped close vulnerabilities, improving product resilience.
  • Localization testing passes across supported locales.

Software Engineer — EMC

Hopkinton, MA · 2011 — 2013

  • Hardware validation and regression testing across SLICs (I/O modules), DAEs (disk array enclosures), DPEs (disk processor enclosures) and storage processors, driven by Python automation frameworks.
  • Environmental and thermal dwell testing on enterprise storage enclosures — validating hardware behaviour across temperature extremes, not just software behaviour.
  • Modified Jenkins plugins in Python to make the build system support the validation workflows the lab needed.
  • Linux builds managed and reproducible.

Projects

Cur8: playlists for a moment, grounded in hand-curated crates and measured by an eval harness

Personal project · TypeScript, Next.js, LLMs, RAG, LLM Evaluation, MusicBrainz, Spotify API, Vercel · cur8music.app/p/highway-coastline-vibes-33ba65

  • The offline tier runs on every change in about five seconds with no AI calls: keyword routing plus rule invariants (a crate never bans its own songs; nothing reached through the taste graph breaks a rule). The model tier makes the real calls and measures crate choice, invented or missing songs, and rule leaks per prompt.
  • Each case lists the crates that would be right, the crates that must never fire, artists the rules must block or keep, and what the model should read from the request (vocals, fame). A no-crate case asserts that nothing fires at all.
  • Every song the model suggests is resolved against the streaming catalog. A suggestion that can't be found is counted as invented or missing, so hallucination is a number per run rather than an impression.
  • The model can propose a banned artist; the code removes it before the listener sees it. The harness reports how often that happened, so a prompt change that makes the model lean on the guardrail shows up even though users never see the leak.
  • The engine is a flag on the harness, so model choice is a comparison on the same cases, not a guess. 13 saved runs so far, timestamped for diffing.
  • The harness reports how often the tagger's vocal/instrumental label agrees with what each crate implies, and whether MusicBrainz performer credits back it up. It also counts duplicate compositions per crate, so one piece doesn't fill a playlist under five recordings.
  • The rule graph, Spotify matching, song facets, YouTube link refresh and Apple Music key handling each have their own tests, run with Node's test runner.

Ryōri Quest: a phone-first game for learning Japanese by cooking

Personal project · JavaScript, Vite, Playwright, Supabase, WebGL, Kokoro TTS, Vercel, PWA · beta.ryoriquest.com/?invite=portfolio

  • Playwright plays the built game in Chrome at phone and desktop sizes: chapters and unlocking, a whole dish, stars, XP, the stamp and the victory screen. Mini-games are finished through a test hook that only exists on localhost, and a playthrough fails on any page error or sideways scroll.
  • Unit tests cover a typed code, a one-tap invite link, a wrong code, and exactly which paths may skip the gate. After each deploy, a smoke check confirms from outside that APIs refuse visitors without a pass and a wrong code is rejected.
  • Display names are tidied and length-checked, and reserved or unfriendly names are refused through a profanity filter, each with its own unit test.
  • Generated lines ship with the game. A recorded line replaces the generated one on every device once an admin publishes it. Tests pin down the permissions: only voice contributors may record, and only admins may publish.
  • Analytics events from the game, players per invite code, a live strip of who is playing right now, and each dish's step spread, so a mini-game that loses players shows up as data.

Job Thing: a Spotify Car Thing, jailbroken into a job board for QA roles

Personal project · Python, SQLite, JavaScript, Playwright, GitHub Actions, Chromium DevTools, Embedded Linux, Vercel · jobthing.vercel.app

  • Chromium 69 in kiosk mode on an 800x480 panel, talking to the server over USB ethernet. Every screen works with turn, press, hold and back, and the device itself never touches the internet.
  • Profiled on the device over the Chrome DevTools Protocol: every detent was rebuilding every card. The rail now rebuilds only when the list changes, and a UI test asserts the cards survive a detent so the regression cannot come back quietly.
  • Greenhouse, Lever, Ashby and SmartRecruiters public APIs plus two aggregators, in parallel, pruning postings only after a fetch succeeds. The server is Python's standard library and SQLite.
  • Titles are matched with include and exclude patterns, because the obvious keywords mislead. Every pattern change starts with a test built from a real posting that fooled an earlier version.
  • A 0 to 100 score from named rules: title fit, resume skills, seniority, location and freshness. No model output to explain away.
  • 125 Python unit tests and 59 Playwright tests across three projects: the HTTP contract, the kiosk UI driven like hardware at 800x480, and the web demo. Every run produces a report with results, traces, videos and both Python and JavaScript coverage.
  • A shim answers the kiosk's API calls from a snapshot rebuilt every morning, and each visitor's presets and saved jobs stay in their own browser. Saved jobs export as a printable checklist.
  • The tailored resume is checked in code against the master resume, so no employer, date, number or tool can be invented. The form is filled in Chrome and left for me to review; it never signs in or answers self-identification questions.

Ask the Library — self-hosted AI reading platform over 50,000 books

Personal project · Python, FastAPI, Elasticsearch, kNN Vector Search, RAG, Ollama, React, TypeScript, Playwright, Docker, Stable Diffusion

  • Eleven specs covering browse → search → detail → reader, made deterministic by intercepting at the network layer and serving fixtures. No infrastructure required, so it never flakes on a cold GPU or a slow index.
  • A second layer pointed at real infrastructure — vector search, LLM generation, the media pipeline — with environment-switchable targets for post-deploy verification. It caught GPU contention starving the search embeddings, which the offline suite by design could not.
  • Elasticsearch kNN over locally-computed embeddings — the retrieval layer the answers are grounded in.
  • Streaming, citation-grounded Q&A served by Ollama across 14B/3B/1B tiers on a two-node home lab. Nothing is sent to a third-party API.
  • One of several infrastructure faults diagnosed end to end, alongside a CUDA/cuDNN conflict silently breaking GPU inference and request starvation from single-GPU contention.
  • Public-domain acquisition from the Internet Archive filtered by OCR word-ratio heuristics, perceptual-hash detection of scanner boilerplate, and an LLM sniff test — with on-demand AI cleanup of the OCR text that survives.
  • Covers scored by Laplacian variance and a vision model, then regenerated through Stable Diffusion on a local GPU — with a before/after debug page that recomputes the scores client-side so every automated keep-or-replace decision can be audited. It exposed two classifier blind spots that became fixes.
  • A study-focused reader with themes, a two-page spread, a chapter-grounded chat panel and a vocabulary builder; plus a discovery UI with semantic search, token auth and saved lists.

Fare — a rideshare platform whose economics are actually fair

Fare Technologies · Swift, iOS, TypeScript, Firebase, Vercel, Dispatch

  • Native iOS driver app built in Swift.
  • A dispatcher-facing web app alongside the driver app, so rides can be assigned and tracked from a desk.
  • A middleware service in TypeScript rather than letting clients talk to Firebase directly — the seam where pricing and dispatch rules live.
  • Engaged through SLVRLeaf rather than built as a side project — a real client with a real deadline. Listed as in-progress on purpose: it is a live build, not a shipped product.

Tiếng Việt Learning Tools — a Vietnamese learning app that takes dialect seriously

Personal project · Swift, SwiftUI, Speech Framework, XCTest, iOS

  • North, Central (Huế) and South, each voiced by a different native speaker, with a dialect selector wired through the app so a learner can choose the accent they actually need.
  • Speaking practice captures the learner saying a word and checks it against on-device speech recognition — feedback on production, not just recognition.
  • Game logic lives in its own view model, which is what makes it unit-testable rather than tangled into the view.
  • Practice covers words, phrases and full sentences.
  • Bundled offline, so lookups work without a network round trip.
  • Pulling the matching-game state into its own view model is what made it testable at all, and it has real XCTest coverage — selection toggling, ordering, and match resolution — rather than the scaffolding Xcode gives you for free.

Sorceror — Chrome extension for regression testing

Internal tool · Sorcero · Chrome Extension, JavaScript, REST, Caching, Internal Tooling

  • Removed the manual token-fetching step that every quality engineer paid on every session.
  • Injected an overlay onto the product page exposing what testers actually needed to see — record and article counts, stakeholder counts — none of which the product surfaced on its own.
  • Surfaced the ontology backing a project alongside it, so testers could see the structure their test data was shaped by.
  • Sorting, management and basic CRUD over the test projects themselves — the setup work happened in the same place as the testing.
  • Environment-aware, so the same extension served dev, QA, staging and production without reconfiguration.
  • Caching kept the overlay fast enough that using it never cost more time than it saved.

Tesseract — a visual Playwright runner for disaster recovery

Internal tool · Sorcero · Three.js, D3.js, Playwright, TypeScript, Microservices, Disaster Recovery

  • Built with Three.js and D3.js so the running system could be seen rather than inferred from logs.
  • Bound the visual topology to the actual Playwright repo, so coverage of a given service was a thing you could look at.
  • Selecting services selected their specs — orchestration by architecture rather than by remembered file paths.
  • During DR exercises, two regions could be run and read against each other directly, which is the question a DR drill is actually asking.
  • Persisted the output of each run so drills accumulated into a record that could be compared over time.

Education

Manny Castillo
ProjectsField Notes

report/index.html

Manny Castillo

Senior Software Engineer in Test

Download résumé (PDF)LinkedIn

0

passed

0

failed

14

specs

14

years

Senior SDET with 14+ years testing backend services, data pipelines, and large-scale distributed systems. API and integration automation in Python, Airflow pipeline validation with nightly runs, and performance work against the services underneath — plus the web and mobile E2E layer on top of it in Playwright and Maestro. Builds the test frameworks and internal tooling other engineers depend on, across enterprise SaaS, gaming infrastructure, storage, and cybersecurity.

Coverage

Languages97.4%

Python · TypeScript · JavaScript · Go · Bash

Backend & API Testing99.1%

pytest · API Testing · FastAPI · REST · GraphQL · Integration Testing · Microservices Testing · Postman · Test Plans & Strategy

Data & Pipelines96.8%

Airflow · Google Cloud Composer · Data Pipeline Validation · Nightly Regression Runs · Kafka · PostgreSQL · AlloyDB · Elasticsearch

ML & LLM Validation93.6%

ML/LLM Output Validation · MLOps Support · Non-deterministic Testing · RAG & Retrieval Grounding · Vector Search (kNN) · Self-hosted LLMs (Ollama) · LLM-Assisted Test Engineering

Cloud & Infrastructure95.2%

AWS (EC2, S3, Lambda, boto3) · Google Cloud · Kubernetes · Docker · Linux · Distributed Systems

Performance & Reliability95.9%

JMeter · Locust · Grafana · Prometheus · Benchmarking · Metrics Analysis · Automated Failure Triage

CI/CD & DevOps97.3%

GitLab CI/CD · GitHub Actions · Jenkins · Git · Secure Workflows · Distributed Test Execution

Web & Mobile UI Testing100.0%

Playwright (TypeScript) · Maestro (iOS) · XCTest · XCUITest · Selenium · End-to-End (E2E) Testing · Mobile E2E Testing · UI Smoke Testing · Cross-Browser Testing · Regression Testing

Web Technologies93.2%

HTML · CSS · JavaScript · REST · GraphQL · OAuth2

Internal Tooling92.5%

Chrome Extensions · Three.js · D3.js · Data Visualization · Developer Experience

Hardware & Systems Validation99.2%

Enterprise Storage (SLIC, DAE, DPE, SP) · Environmental Dwell Testing · VMware ESXi · Debugging · Root Cause Analysis · Reliability Testing

Experience

Projects

Client and personal work first; internal tooling built at Sorcero is marked.

Education

B.S. Computer ScienceUniversity of Massachusetts, Lowell, MA2012
Study Abroad, Computer ScienceAmerican University of Sharjah, Sharjah, U.A.E.2008

1 test is still failing.

expects this candidate to be off the market. It is currently receiving "available immediately".

linkedin.com/in/mfcastillo