Skip to content

ML Engineer — Risk & Trust

Build the trust layer for autonomous agent spending: fraud scoring, on a behavioral dataset that doesn't exist anywhere else.

Engineering

Remote (global)

Vertrag

Part-time → Full-time (YC acceptance)

Vergütung

€2,000-3,500/month + 0.5-1% equity

Verfügbarkeit

Offen, bis die Stelle besetzt ist

The role

Payle's authorization engine decides, in milliseconds, whether an AI agent is allowed to spend money. It's deterministic, policy-driven, and live. What it doesn't have yet is a risk layer: the ML that scores transactions for fraud, scores merchants for quality, and builds trust profiles for agents over time.

Phase 0 of that risk engine is rules-based and in final build. Phase 1 is yours: evolve it into real ML on a dataset nobody else on earth has: an append-only ledger of every agent authorization, payment, and outcome, with full context attached.

What you'd build

  • Risk scoring for agent transactions: fraud detection, merchant quality scoring, and per-agent trust profiles that evolve with every verified outcome
  • A feature store with point-in-time correctness: no leakage. Enforced in code, proven by test, not promised in a doc
  • Models: starting with LightGBM (explainable, fast, right for tabular data at 20ms latency budgets), evolving as labeled data grows. If a deep learning model earns its place, bring the evidence
  • Explainability on every decision: SHAP-derived factors, because credit-adjacent outputs must be explainable to users, merchants, and eventually regulators
  • Model discipline: registry, model cards, drift monitoring (PSI), shadow-mode deployments before anything influences a real decision, and fairness testing on underwriting-adjacent models
  • The honest constraint: you cannot train fraud detection on data that doesn't exist yet. Phase 1 is building the pipelines, the feature store, and the evaluation discipline. The models earn their deployment as the labeled data grows. If you'd rather pretend otherwise, we're the wrong company

You

  • Strong classical ML: LightGBM/XGBoost, feature engineering, proper validation. And you can explain why you'd pick classical over deep learning for tabular fraud scoring at 20ms latency
  • You've felt the pain of data leakage, or you're hungry to learn why it silently kills models that look perfect offline
  • You treat AUC 0.99 as a red flag, not a win
  • Python strong, SQL competent, comfortable with point-in-time joins and event-time reasoning
  • Written communication: the team is distributed across four countries and writes everything down. You explain model decisions in text, cleanly

Compensation

€2,000-3,500/month (part-time to start, full-time path at YC acceptance) + 0.5-1% equity (4-year vesting, 1-year cliff). Remote-first, async-friendly, Friday demos for the whole team.

The challenge

Your first artifact: a fraud pattern hunt. We give you a synthetic agent ledger dataset with planted fraud patterns (some obvious, some subtle). Build the scoring model, find the patterns, and defend every feature, threshold, and model choice in a written analysis. We pay for your time. You keep the work.

Erstes Artefakt

Fraud pattern hunt

Find the planted fraud pattern in a synthetic agent ledger dataset, build the scoring model, and defend every feature and threshold you choose.

Lieferumfang: Jupyter notebook + written analysis: features used, model choice defended, explainability output, and the fraud pattern you found.

Passung

Wir passen wahrscheinlich nicht zu dir, wenn

  • Du ein großes Gehalt willst, bevor es ein Produkt mit Nutzern gibt. Wir sind Pre-Seed: Equity ist die Wette.
  • Du Tickets, Sprints und einen Manager brauchst, der deine Stunden prüft. Wir arbeiten zielorientiert.
  • Du glaubst, KI könne den Zahlungsweg entscheiden, weil sie klug genug sei. Wir haben eine feste Regel.
  • Du wegen des Titels hier bist. Hier spricht der Code, auch beim CEO.
Du liest noch? Dann passen wir wahrscheinlich zu dir.