saihaj.ca

Toronto, Canada · open to remote

Saihaj Bath.
First commit to production.

SaaS backends, browser extensions, iOS apps, and the infrastructure underneath them. Currently a contract software engineer at Zencargo.

homelab status: online

01 · work

Selected work

A shared inbox platform built inside Gmail for about 70 users across seven departments, owned from discovery interviews through production rollout.

  • Python
  • TypeScript
  • Cloud Run
  • Firestore
  • Manifest V3
  • OAuth
  • KMS
  • Pub/Sub

The brief was to make team email feel native to Gmail instead of moving people to another tool. That decision drove everything: a Chrome extension as a thin client over shared state, with a backend on Cloud Run and Firestore owning assignments, status, notes, SLA timers, and the audit log. Gmail push sync over Pub/Sub keeps it live.

Roughly 1,000 commits over six months. The backend went through four generations before it settled: Apps Script and Sheets, then n8n, then TypeScript on Cloud Run, then Python in the company monorepo. Each step was forced by a measured ceiling, not taste.

Decisions that mattered

  • Per-inbox OAuth instead of a service account with domain-wide delegation. Least privilege: every shared mailbox grants explicit, revocable consent with narrow scopes. The cost is that operators connect each inbox by hand at kickoff, and that cost was accepted on purpose.
  • Labels-only routing with a fan-out insert. A shared message is fetched raw and inserted byte-identical into each member’s own mailbox, then Gmail labels carry the queue state. No forwarded copies, no header rewriting, native Gmail search keeps working. The shared view is Firestore, not Gmail threading.
  • Gmail thread ids turned out to be per-account, which broke the first right panel. The answer was a four-strategy thread resolver with the RFC 2822 Message-ID as the cross-account key, bounded by a seven-day subject-hash window.
  • The TypeScript to Python rewrite was an organizational decision made technical: the company’s infra and on-call muscle is Python, and the TS code could not clear the reviewer’s conformance bar without a retrofit that cost as much as a port. The extension stayed byte-identical except for the API client; KMS envelopes stayed byte-compatible so tokens carried over without re-consent; every TS test got a pytest equivalent or a written exemption.
  • KMS envelope encryption, not direct KMS calls, for the two token collections. Fresh DEK and IV per write, AES-256-GCM with the collection and document id as associated data, so a ciphertext moved to another document fails to decrypt. Rotation works at the key-version level without re-encrypting payloads.
  • Authorization is per-inbox role checks in per-route decorators, never a global before-request hook. An AST test fails the build if any handler taking an inbox id forgets to call a role helper, and a 31-route by 5-role matrix pins the boundary.
  • Three-tier rate limiting (user, inbox, global) because a per-user tier alone lets one member burn shared Firestore reads for everyone. The global tier is the circuit breaker for Cloud Run and the Gmail quota.

Hardest production bug: a single-instance saturation compound. A label sync for one nine-member inbox ran 85 seconds on the request path, because it rebuilt a Gmail client per thread per member, about 5,600 Firestore round trips and 1,800 KMS decrypts for 200 threads. Fixed by batching per member and moving the sweep behind an atomic claim, roughly 9 KMS calls and 50 Gmail calls instead of 1,800. The lesson that stuck: the saturation was self-inflicted load, not capacity.

A weekly supplier chase for cargo ready dates on open ocean purchase orders, in production for four enterprise customers and nine suppliers in three languages.

  • n8n
  • BigQuery
  • GraphQL
  • Apps Script
  • Gmail API
  • Slack API
  • i18n

Ops teams were chasing supplier confirmation dates by hand. The pipeline closes the loop: it reads order state from BigQuery, emails each supplier a per-language summary with a unique form link, takes confirmations and delays line by line, writes delays back to the platform over GraphQL, and logs every event to an append-only audit sheet that a dashboard grades suppliers from.

Inherited as a prototype in February, rebuilt, and live by the end of March. Over the first three months it sent 124 fully automated emails across 71 weekly chases and recorded 131 line-level actions, with a median time to open of 3.5 hours.

Decisions that mattered

  • Config is a sheet the ops team already uses: recipients, chase window, language, and an active flag per supplier. Config wins over data; a sheet email always overrides what BigQuery thinks the supplier’s address is.
  • The BigQuery layer normalizes messy reality. Suppliers and buyers exist under several org IDs; mapping CTEs fold them to a canonical identity before any chasing logic runs, because a silently missed purchase order is worse than a late one. The wide-net version of that rule later collapsed a multi-division supplier into one contact, so the rule became: narrow by default, verify the product mix before adding an alias.
  • Overdue is latest date wins, applied at the purchase order: a PO is overdue only when every line is past its date, and the query keeps only each line’s most recent revision. Anything stricter chased partially ready orders and trained suppliers to ignore the email.
  • Supplier-facing links route through a separate unauthenticated ingress path, because external recipients cannot pass the internal OAuth proxy. Same cluster, different route, found on go-live morning.
  • Only delays write back to the platform. Confirms are audit-only, flags and cancels alert ops immediately. A confirmed but unbooked order is re-chased next week by design; the remedy is booking, not suppression.
  • Follow-ups pass a freshness gate that re-queries BigQuery before sending, after a follow-up once chased an order that had been fully booked for a day. Reply-in-thread stays, because suppliers answer in-thread, guarded by a UUID check against active config.
  • The email engine is render-then-send: one template registry keyed by language, one Gmail node. Adding Hindi touched five places instead of a graph surgery in two workflows, proven byte-identical for English and Chinese before it shipped.

Hardest production bug: a Friday follow-up replied into one supplier’s real email thread with another customer’s order data. Root cause was data, not code: a demo row built with the supplier’s real address kept its thread id while the contents were swapped for a test. The fix was a pre-send UUID-versus-config guard plus a standing rule that test data is fake or obfuscated, deleted immediately. Measured outcomes worth stating plainly: the most engaged supplier went from a 33-hour response to under one hour and answered nine of ten weeks, and the two chased suppliers at one customer went from 25 to 31 percent late bookings per quarter to zero in the first full quarter under chasing, on a small sample and stated as correlation.

Customer escalation management on Salesforce, with four intake paths, AI triage for email, and 800+ cases tracked. Before it, escalations lived in Slack threads with no ownership and no chasing.

  • n8n
  • Salesforce
  • Gemini
  • Slack API
  • Apps Script
  • SOQL

One escalation is one Salesforce Case. Four intake paths feed that record type: an in-platform form, a shared mailbox, Slack, and a Google Form. Around the Case sit nine coordinated workflows: routing, Slack notifications to raisers and owners, stale-case chasing on a business-day ladder, and drafted closure emails. Slack is the only human interface. Two principles held the whole thing together: Salesforce is the system of record and the automation is glue, and AI goes where text is unstructured, rules everywhere else.

Decisions that mattered

  • Every intake path converges on create a Case, and everything downstream triggers off the Case, not the path. That property was verified before the fourth path was added, and it is why the fourth path shipped without touching the other three.
  • Email gets a three-stage Gemini pipeline: a cheap model extracts addresses from the raw thread, a stronger one produces twelve structured fields against an embedded category table, and a third resolves the customer account with a live SOQL tool. Staged because each step has a different input size, model, and failure gate; a case is never created when the account cannot be resolved, because guessing an account is worse than a DM asking for details.
  • Schema rules learned from two opposite crashes: a strict string type crashes the parser when the model correctly returns null, and a nullable union makes the model reject every call. The pattern now is single types, never null, empty string instead, and gates on not-empty rather than exists.
  • Retries only on idempotent nodes, never on writes. A duplicated case or a double Slack ping costs more trust than a delayed one, so writes hard-fail to an error workflow instead.
  • Automation drafts, humans send. Closure emails are generated for review and never fired at customers, after the model was observed drafting for a still-open case while someone edited it live.
  • I run the Salesforce side as well: record types, dependent and restricted picklists, intake forms, and flows, so the data model and the automation stay in one pair of hands.

Two lessons that cost the most. First, audit the error workflow by its own executions: it had been broken from creation for months, so no alert had ever reached the channel, and nobody noticed because a quiet channel looks healthy. Second, a failed first attempt still burns a dedup key: an email thread whose first forward failed was silently dropped as a duplicate when re-forwarded, and that one was diagnosed read-only from pruned execution data. The fix moved dedup from thread id to message id with a thread-to-case table, and made every skip path tell the forwarder why.

A daily billing chaser for freight finance: timezone-aware Slack DMs, a business-day escalation ladder, and a daily P&L summary to leadership. 337 draft documents chased in the first three weeks, 75.7 percent published after a chase.

  • n8n
  • Slack Block Kit
  • BigQuery
  • Google Sheets
  • Luxon
  • HMAC

Draft billing documents past their revenue date are money earned but not booked. The system reads the finance feed hourly, finds every draft that passes a stack of customer and mode rules, and DMs the responsible coordinator a single digest at 15:00 in their own timezone. A Slack pop-up form captures commitments, blockers, and margin reasons; non-responders get their manager after two business days; a support request opens a private channel with the named blocker; and a daily summary tells leadership what is unbilled, what is below target, and whether chasing is working. Built solo in seven weeks and live across four timezones.

Decisions that mattered

  • Hourly, not daily, because every hour is somebody’s three o’clock. Coordinators span Europe, the US, and China, so one fire time was rejected on day one; the run computes each coordinator’s local hour live from an IANA zone and sends only when it is 15:00. In practice that is four real waves a day, and an empty run at 20:00 is not a problem.
  • The tracker is the source of truth for did we chase this today, not the feed. One persistent row per shipment with clear field ownership between the chaser, the form, and the escalation, so the daily upsert never clobbers what a human wrote, and double DMs are structurally impossible.
  • A Slack modal, not a web app. The form is built entirely from the button payload so it opens inside Slack’s three-second window with no sheet read; later the reveal step moved from a Sheets read to a server-side cache, 641 milliseconds to 65.
  • Committed never auto-escalates. A coordinator who says I will publish resurfaces the next day flagged urgent and sorted first; the leadership nudge is the lever. A whole escalation track was deleted to make that true.
  • Group DMs cannot have a member added later, which killed the original add-the-manager-to-the-thread design. Support became one auto-archiving private channel per issue, the model the infra team already used.
  • The gate rules live in three surfaces that must move together, plus a fourth non-gating mirror. A register tracks all of them, because a rule landing in one surface and not the others is the single most common way this system goes wrong.

Worst bug by what stakeholders saw: fourteen shipments got a false Resolved post and an archived channel because the feed is multi-row per shipment and a last-row-wins lookup let one old published cost line make the whole shipment look published. Worst by invisibility: a hotfix silently reverted three rules in one mirror for three days with every execution green, parking exactly the shipments a stakeholder had just asked to cover. The standing rule since: after any publish, re-fetch and assert the live version is the one you tested.

An iOS app that manages a fleet of self-hosted home-server services from one dashboard.

  • Swift
  • SwiftUI
  • async/await
  • Clean Architecture
  • Keychain

Seven services, each with its own REST API, auth scheme, and data shapes. The app’s core decision is a single connector protocol: every service implements the same four-method interface (test connection, server info, health, dashboard summary), so the dashboard renders whatever conforms and a new integration is one adapter, not an app change.

Decisions that mattered

  • Clean architecture with MVVM keeps connectors, domain models, and views in separate layers. The UI knows nothing about any specific service, and the protocol ships a default dashboard summary so a minimal connector is still a complete one.
  • Aggregation runs in parallel with structured concurrency (TaskGroup), and each service is isolated: a connector that throws renders as a degraded tile with the error attached, never as a failed dashboard. One dead service cannot take the screen down.
  • Credentials live in the Keychain, and every log call passes through pattern-based secret redaction before it is written. A debug log can be shared without scrubbing.
  • Endpoint failover is transparent: a two-second reachability probe tries the local address first, then falls back to the remote one, and local traffic bypasses any system proxy. It works identically at home and away, and the user never picks a network mode.

A production-grade homelab: 20+ containerized services behind zero-trust access, answering the status check on this page live.

  • Cloudflare Zero Trust
  • Docker
  • TrueNAS
  • OIDC / Passkeys
  • WireGuard
  • DNS

A TrueNAS server running 20+ containerized services, published through a Cloudflare tunnel with no exposed ports. Every service sits behind zero-trust access with passkey single sign-on through a self-hosted OIDC identity provider.

Decisions that mattered

  • Ingress is tunnel-only. Nothing listens on the public internet, so the attack surface is the identity layer, which is exactly where I want it.
  • Passkeys over passwords for SSO. Phishing-resistant by construction, and honestly more convenient.
  • DNSSEC, strict TLS, and DMARC enforcement across the board, because infrastructure that only mostly validates is infrastructure that fails quietly.
  • The whole stack recently migrated between domains, identity provider included, with zero unplanned downtime. The sequencing mattered: certificates first, then the IdP, then services, with rollback points at each stage.
  • The same habits carried into paid work: a three-site dental group ran on a WireGuard failover I wrote in Bash against the network controller’s API, recovering in under 30 seconds, and it handled real outages in production.
  • The status pill at the top of this page is answered live by this stack through a health check endpoint. It is the one place the lab touches this site.

02 · experience

Experience

Contract Software Engineer · Zencargo

London, UK · remote

Jul 2025 – Present

Software Engineering Intern · LifeMantra

Remote · registration platform for WWF · code ↗

2025

Technical Support Specialist · Neuteck Technologies

Ontario

2021 – 2025

IT Infrastructure & Systems Manager · Fisher / Nepean / Ottawa Dental

Ottawa · 3 sites

2023 – 2025

Service Desk Specialist · Gosolution Management

Remote

2020 – 2021

03 · about

About

I got my start running infrastructure for real businesses, then moved into building software. That path shows in how I work: I own things end to end, from architecture and code review through deployment, monitoring, and being the person who answers when something breaks.

Computer Science at the University of Toronto. CompTIA A+. Government of Canada Level II (Secret) security clearance.

now: Zenbox at Zencargo · the homelab · Media Stack Manager for iOS

Languages

  • Python
  • TypeScript
  • Swift
  • SQL
  • Bash
  • Java
  • C/C++

Cloud & Infra

  • Google Cloud
  • Cloudflare
  • AWS
  • Docker
  • Terraform
  • Linux

Web & App

  • Next.js
  • React
  • Astro
  • Node.js
  • SwiftUI
  • Chrome Extensions
  • REST & GraphQL

Data

  • PostgreSQL
  • BigQuery
  • Firestore
  • Prisma
  • Salesforce (Admin, SOQL, Flows)

Security & Reliability

  • OAuth/OIDC
  • KMS encryption
  • RBAC
  • OpenTelemetry
  • CI/CD gates

04 · contact

Get in touch.

I read everything and usually reply within a day or two.