Available immediately for full-time roles

Software Engineer · Prediction Markets & Sports Trading

Immanuel
Anaborne

I am a software engineer focused on designing and building functional infrastructure. My work includes order execution, cross-venue matching, and studies that decide whether a strategy deserves capital.
Penn CS + English '26, based in New York.

120M → 159,635
Candidate pairs scored, cross-venue matcher
250 of 250 planted pairs recovered, from a clean clone
2.8×
Executor wake_recv p50, Python → C++
one machine, one run, an upper bound
395
Offline tests, public repo
mypy --strict · ruff clean
60,000+
Filings processed, Wharton RA
async ingestion, Pydantic-validated

Selected work

Public repositories, verified by running

Each card carries a 22-second demo. The demos are AI-generated: written and rendered with /brag and HyperFrames, from the text on this page. Every figure in them is a figure from the card beside it. The repositories, the numbers and the write-ups are mine.

10 of 10 projects

No project matches that filter.

Kalshi / Polymarket Trading Infrastructure
Python · asyncio · uvloop · SQLite · Unix domain sockets
mypy --strict ruff clean 395 tests public extraction
AI-generated

A two-process trading system for Kalshi and Polymarket: a poller that decides what to do, and a pre-warmed executor that owns how (order count, time-in-force, self-trade prevention), connected over a Unix domain socket so unbounded decision work never contends with the event loop that actually submits an order.

Evidence architecture, 4 findings, 1 caveat
poller unix socket 4-byte framed executor pre-warmed exchange API POST /orders kill switch risk gate send_wake read dispatch clear WakeAck, before dispatch is awaited rejected row telemetry.sqlite (WAL) engaged / refused fire and forget
Every refusal writes a row. A halted system's inaction lands in the audit trail, so nobody has to infer it from silence.
  • Chose uvloop over stock asyncio on a measured benchmark. Latency is a property of the machine that runs the code, so every run is committed to benchmarks/history.csv with its platform tag and a clone prints its own numbers. On the most recent committed run, wake_recv p50 falls from 0.0059ms to 0.0045ms under uvloop.
  • Designed an IDF-weighted inverted index that prices a full two-venue cross by cutting a ~1B-pair search space to ~63K scored candidates.
  • Root-caused a position-cap defect that had rejected 4,233 valid orders over 23 hours against zero open positions; replaced the counter-based estimate with live exchange-position reconciliation at no added API cost.
  • Runs strategy searches as pre-registered experiments, with criterion, sample size, and stopping rule fixed before data, and kills each on its own evidence.
What the benchmark measures Both benchmarks in the repo generate their own data and need no credentials and no network: RSA-PSS signing timed against an ephemeral keypair, and a real poller → socket → executor round trip against a fake venue client. The numbers above the fold come out of those benchmarks on a clean clone. End-to-end detect→fire on the live system needs a signed round trip to a demo account, so it is not published here.
View source & benchmarks → github.com/anaborne/prediction-market-infra
Executor Hot Path in C++
C++20 · CMake · Catch2 · OpenSSL 3 · SQLite · Unix domain sockets
CI: Linux + macOS ASan/UBSan 92 tests byte-identical to orjson
AI-generated

The executor process from the trading system above, rewritten in C++20 against the same wire protocol, so the unmodified Python poller drives either one and the two are timed against each other in the same run on the same machine.

Evidence 5 findings, 1 caveat
  • The Python poller drove the C++ executor across 4,400 frames with poller_client.py untouched: every frame accepted, no telemetry row dropped.
  • Executor-side wake_recv fell by 2.4x at p99 and 2.8x at p50 under uvloop, 4.5us to 1.6us at p50. The Python baseline runs the poller and the executor on one event loop and the C++ configuration spawns a second process, so every ratio is an upper bound on what the language change bought.
  • The encoder is byte identical to orjson: 10 golden frames, 1,000 doubles and 424 rounding cases were produced by running the Python and are asserted in the C++ suite.
  • The pre-registered expectation was wrong in both halves, and the README says so in its third paragraph. The signer row, 2.8x in the port's favour, compares two OpenSSL builds rather than two languages and stays in the table with its meaning corrected.
  • Telemetry leaves the read loop through an 8192-row ring and one writer thread. A full ring drops the row and counts it, so a database write never delays an ack.
What the comparison measures Two implementations of one wire protocol, timed on one machine in one run, with the raw rows committed on both sides. It is not a production A/B, and no number here is presented as one. The Python side is a published reference implementation of a system that was shut down on 2026-08-29.
View source & results → github.com/anaborne/executor-hotpath-cpp
Settlement Reconciliation
dbt · DuckDB · Parquet · Python 3.12 · sqlfluff · GitHub Actions
82 data tests 15 of 15 controls falsified enforced output contracts dbt docs on Pages
AI-generated

A dbt project that joins a forward test's trade decision ledger to the Kalshi settlement feed that should explain it, classifies every decision the feed has not settled, and fails the build when the unexplained population grows past the size the repository records.

Evidence 6 findings, 1 caveat
  • 281 decisions from 5 of the 12 runs in the run log, against a 263-row settlement feed, a 93.59% match rate. The 18 that do not settle are split into 11 unexplained, 6 still inside the feed's own lag window, and 1 market the exchange finalised without a binary result.
  • The settlement window is computed from the feed rather than set as a constant. The newest load is the watermark and the largest close-to-load lag across the 263 settled rows, 20.2 hours, is how far back a decision can still be pending. The 11 unexplained rows sit more than five times that lag before the window opens.
  • Every control is proven to fail when its fault is injected. verify_controls.py copies the fixture, breaks one thing, rebuilds the models, and asserts that one control catches it. 15 of 15, as its own CI job on every push.
  • A sixteenth control is published as inert rather than counted. A relationships test on the join only sees rows where the join matched, so it cannot fail, and a rowcount guard is kept as the falsifiable check in its place.
  • The one correlation that survives is confounded and says so. All 11 unexplained rows fall in the first 21 decisions under the early rulesets, and ruleset moves with collection time across those rows, so no cause is asserted.
  • 31 KB of parquet is committed, so a clone runs the whole project, the tests, and the mutation harness with no warehouse account and no access to the private source.
What is not diagnosed The 11 unexplained rows are not explained. The extract does not contain the answer and none is asserted. The reverse break, a settlement with no decision, is 0 and cannot be anything else on this data, because the source enforces that key. The models run on demand and nothing here is a live pipeline.
View source & docs → github.com/anaborne/settlement-reconciliation
Kalshi Prop Calibration
Python · numpy · scipy · pandas
pre-registered 100 tests corrections published
AI-generated

A pre-registered test of whether a forecast built from free public data beats the Kalshi mid-quote on NFL and NBA player props. The headline is a null. The finding worth keeping is structural: the midpoint convention stops being a probability when one side of the book is absent.

Evidence 5 findings
  • Hypotheses, features, model form, split, fee treatment and kill criteria were written down before any data was pulled. Later corrections are appended and dated, with the wrong text left standing.
  • The midpoint convention (yes_bid + yes_ask)/2 stops tracking the strike when one side of the book is absent, which happens on 16.0% of 102,860 scored observations. In the 87,053-observation training partition those one-sided rows are 17.3%, and they alone invert the sign of the headline Brier skill score.
  • Caught a scoring-hour boundary bug that would have priced 88,796 of 144,176 observations (61.6%) with an hour of market information the model was denied, in the market's favour, before any score was recorded.
  • Free public data reproduces the exchange's own settlement on 61,978 of 61,984 resolvable markets (99.990%), with all six disagreements named by ticker.
  • Every S1 statistic reproduces from a clean clone in about a minute with no network, off 4.7 MB of committed extracts.
View source & data → github.com/anaborne/kalshi-prop-calibration
NFL Pricing Model
Python · pandas · walk-forward Elo
pre-registered 272 games publishes its loss
AI-generated

A walk-forward Elo model over the 2025 NFL season, graded against the de-vigged market moneyline. The constants were fixed before any result was read, and the places the model loses to the market are in the repository.

Evidence 3 findings
  • Walk-forward Elo over 272 games using only pre-game information, constants fixed before any result.
  • Brier 0.224 against the market's 0.212, reported with a full calibration diagram including where it lost.
  • Where model and market disagree by more than 0.10 (n=57), the model's Brier is 0.2614 against the market's 0.2040 and its directional hit rate falls to 50.9%. The gap widens with disagreement, which is the opposite of a tradeable edge.
View source & data → github.com/anaborne/nfl-pricing-model
Kalshi Weather Market Calibration Study
Python · pandas · pre-registered backtest
pre-registered kill criterion triggered $0 data cost
AI-generated

A pre-registered test of whether Kalshi's daily temperature markets are mispriced. They are not. The kill criterion was written down before any backtest data was pulled, it triggered in every stratum, and the strategy was killed on that evidence.

Evidence 3 findings
  • Over 235,145 evaluated bucket-hours, the market mid-quote scored Brier 0.0772 against the model's 0.1739.
  • Pulled 60,906 settled markets and 6.8M hourly observations at $0 cost, from unauthenticated public endpoints.
  • Kill criterion fixed in writing before any backtest data, triggered in every stratum, strategy killed on that evidence.
View source & data → github.com/anaborne/kalshi-temperature-calibration
Cross-Venue Settlement Identity Study
Python · standard library · pre-registered gate
pre-registered killed at gate A $0 data cost
AI-generated

Two venues listing the same game are not automatically listing the same contract. Before pricing anything across Kalshi and Polymarket US, this asks whether a matched pair settles on the same source, at the same time, under the same edge-case rules. On the sample drawn, none of them do.

Evidence 4 findings, 1 caveat
  • 74 game-moneyline pairs joined on league, date and teams from live open markets on both venues. 0 are settlement-identical: the settlement source matches on 0 of 74, the settlement time on 0 of 74, and the cancellation handling on 74 of 74.
  • One clause carries all 74. Kalshi settles a postponed game at last fair price within 48 hours; Polymarket US honours the actual result within two weeks of the original date. A game replayed three to thirteen days late cashes one leg at a price and pays the other on the outcome, which is an unhedged bet on the rulebooks.
  • The kill threshold of 30 settlement-identical pairs was set before the sample was drawn. The committed PLAN.md is dated 28 August, inside the 27 to 28 August run, so that order is stated here rather than checkable from the repo. The sample came in at 74, was not adjusted afterwards, and the study was stopped at its first gate for $0 and one day. The 40-pair minimum quoted in the run notes was never written into the pre-registration, and the notes say so.
  • Reconciliation is printed at every step. 115 Kalshi events retrieved, 111 parsed, 4 dropped with the reason named. A pull that reported 2,499 markets on a short page was caught by that same check and was low by a factor of 217.
What a clone can and cannot check The write-up, the run log and the 74 scored pairs are committed and readable. The roughly 570 MB raw pull behind them is not in the repository, so the comparison script needs that data to re-run and does not execute from a clone alone.
View source & write-up → github.com/anaborne/cross-venue
Freshli
TypeScript · Next.js 15 · Supabase · OpenAI
40 tests CI across 4 timezones bugs published
AI-generated

A food-inventory app: photograph the groceries, get an editable inventory, get recipes from what is actually in the fridge. The part worth reading is the refactor. Pulling the logic that had been written twice into tested pure functions turned up five live defects, and the repository documents them.

Evidence 5 findings
  • Every expiration date read a day early west of UTC. A bare YYYY-MM-DD parses as UTC midnight per the spec and was compared against a local clock, so in New York an item expiring today displayed as expired from 20:00 the evening before. CI now runs the suite under UTC, America/New_York, Asia/Tokyo and Pacific/Kiritimati, because a green run in UTC alone would not have caught it.
  • The day arithmetic mutated the value it then compared against, so the branch on the following line was correct only by accident and only while those two lines stayed in that order.
  • The recipe-line parser matched integer quantities only, so Chicken Thighs (1.5 lb) fell through to the branch meant for salt and pepper and deducted 1 from the user's inventory instead of 1.5. A wrong write, silently.
  • The next/image allowlist for generated recipe images sat in next.config.ts while a next.config.js sat next to it. Next 15 resolves next.config.js first and ignores the .ts when both exist, so the allowlist was inactive and every recipe illustration threw.
  • 40 tests over pure functions, no network, database, API key or component rendering. Typecheck and lint are clean gates.
View source & tests → github.com/anaborne/freshli
Google Write MCP
TypeScript · Node.js · Google Drive API · MCP
MIT CI: Node 18/20/22 live-Drive verified
AI-generated

An MCP server that gives an AI agent in-place writes to Google Drive files, guarded by optimistic-concurrency revision tokens that refuse a stale write. A stale write would otherwise fork the document silently.

Evidence 3 findings
  • Verified with 67 unit tests plus a 41-check suite that runs against a live Drive account.
  • It refuses what it cannot write safely. A binary file, and every native Google type except Docs and Sheets, is refused before any write is issued, so a PDF is never overwritten with the text of its own base64.
  • MIT-licensed and published; CI green on Node 18, 20, and 22.
View source → github.com/anaborne/gdrive-write-mcp
Gmail Multi-Account MCP
TypeScript · Node.js · Gmail API · MCP
MIT CI: Node 18/20/22 81 tests sending off by default
AI-generated

An MCP server that holds several Gmail accounts open at the same time. A connector authenticates one account, so reaching a second means disconnecting and reconnecting or forwarding one inbox into the other, and both lose the distinction that decided to keep them apart.

Evidence 4 findings
  • An account label is checked rather than trusted. On an account's first use the server calls users.getProfile and compares the address it was configured with; a mismatch disables that account and names both addresses. The setup error this catches is pasting the second token under the first label, which otherwise produces a server that works and reads the wrong inbox.
  • Reads and writes are deliberately asymmetric. A read may name any mailbox and says in its header which one it used. A write always names its own account, never inherits the active one, and is refused when the two differ unless the call passes confirmAccountSwitch.
  • 19 tools, 17 of them when sending is off. With GMAIL_ALLOW_SEND unset the send tools are never registered, so the tool list a client sees contains no way to put mail on the wire, and turning them on takes the value true, case-insensitively, with anything else failing at startup on an error that names the variable.
  • Subject, address and threading values are rejected if they carry CR or LF before the message is assembled, so text arriving in an email body cannot add its own Bcc. 81 tests run from a clean clone with no network and no credentials, green on Node 18, 20 and 22.
View source → github.com/anaborne/gmail-multi-mcp

Experience

May 2024 to June 2026

Data Engineering
Research Assistant
Wharton Accounting Department
JUN 2025 to JUN 2026
PHILADELPHIA, PA
  • Cut manual disclosure-classification effort 50% with a RAG pipeline (OpenAI API) mapping unstructured corporate filings to trade codes, plus an event-driven ingestion layer for asynchronous filing updates.
  • Built an async ingestion service (asyncio, BeautifulSoup) that pulled and parsed 60,000+ corporate filings, with Pydantic schema validation and pytest suites gating data integrity before load.
  • Deduplicated and normalized entities across 500+ firms with inconsistent naming using vectorized pandas/NumPy ETL, removing manual entity-matching from the pipeline.
Software Engineer
Intern
GoodRun
MAY 2024 to AUG 2024
NEW YORK, NY
  • Built and shipped an in-house React feedback tool that replaced a paid third-party vendor, cutting license cost and giving the team direct ownership of the data pipeline.
  • Reduced production deployment failures 15% by hardening CI/CD pipelines with GitHub Actions and rollback rules.
  • Owned triage and resolution of 20+ production defects across the feedback and CI tooling codebase on a weekly release cycle.

Background

Education & additional

B.A. Computer Science & English
University of Pennsylvania · QuestBridge Match Scholar
EXPECTED DECEMBER 2026
PHILADELPHIA, PA
U.S. National Team Training Partner
Short Track Speed Skating · US Speedskating
2019 to 2022

Named to the U.S. Junior Development Team in 2019 and 2020; 16th and 15th overall at the 2019 and 2020 U.S. Championships

Technical skills

PythonC++TypeScriptJavaScriptSQL Node.jsREST APIsasyncioPydanticReactMCP servers pandasNumPyRAG pipelinesOpenAI API DockerGitHub ActionsGitpytestCMakeCatch2
↑↓ navigate ↵ open esc close