Original job description
Product: Agentic AI tutoring platform
Company: Elumai
Stage: Pre-launch, preparing for general availability
Engagement: Fixed-price project. Budget to be proposed.
Duration: Approximately 10-12 weeks.
Location: Remote. Americas time zones preferred.
Stack: Python, FastAPI, PostgreSQL, Google Cloud Platform
About Elumai
Elumai is an agentic AI tutoring platform. The system pairs specialized expert models with per-learner adaptation that captures how each learner actually thinks, not just what they get right. The backend is mature and already in production; we are preparing for general availability and investing in the performance and infrastructure that will carry us there.
This is a hands-on, high-autonomy role suited to someone who prefers a small, experienced team to a large organization, and who is comfortable working in an actively evolving codebase as we approach general availability. We are a small team working without intermediary management so you would work directly with the founder and frontend engineer.
Project overview
This is a finite engagement to deliver a defined set of performance and systems improvements across the Elumai backend. The project spans five workstreams: queue and background worker architecture, database performance, core subsystem performance, API contract refinement, and edge-case latency and production readiness. Each workstream has named deliverables and acceptance criteria. The engagement is for incremental improvement to harden subsystems and refine architecture progressively, not a single large rewrite. The engagement ends when those deliverables are met.
Scope
Queue and background worker architecture
Select and deploy a queue primitive appropriate to our stack (Cloud Tasks, Pub/Sub, Celery, Temporal, or equivalent).
Migrate non-response compute off the critical path with durable retry semantics.
Define retry, failure, and observability behavior for all queued work.
Produce a design document and operational runbook.
Database performance
PostgreSQL query plan analysis and index design on the highest-impact endpoints.
Connection pooling and query shaping.
Pgvector tuning for our retrieval workloads.
Establish and document an eager-vs-lazy loading strategy across the pipeline rather than leaving it to per-endpoint convention.
Core subsystem performance
Implement time-budgeted execution with graceful degradation under load on core retrieval and computation paths.
Measure baseline and improved performance.
Document behavior, budgets, and fallback paths.
API contract refinement
Partner with our frontend engineer to define response shapes, streaming envelopes, and error semantics.
Contracts should be driven by the needs of the user interface, not the convenience of the backend.
Document the contracts so they can be extended after the engagement ends.
Edge-case latency and production readiness
Bring latency in edge-case paths closer to parity with the primary path through pre-warming, parallel speculative execution, and tighter conflict resolution handling.
Instrument diagnostics for latency incidents, database pressure, query plan regressions, and LLM serving edge cases.
Establish patterns that let the team diagnose recurring production issues without contractor involvement.
Deliverables
Queue and worker infrastructure deployed to GCP, with at least three high-impact background tasks migrated off the response path.
Documented retry, failure, and observability semantics.
PostgreSQL and pgvector tuning applied to production, with before/after benchmarks.
Written eager-vs-lazy loading strategy, adopted across the pipeline.
Time-budgeted core subsystem in production with measured performance.
API contract specifications documented and agreed with the frontend engineer.
Edge-case latency improvements shipped to production, with before/after benchmarks on the three highest-impact paths.
Diagnostic tooling and runbooks for recurring production issues: latency incidents, LLM serving edge cases, and query plan regressions.
Design documents and operational runbooks for all subsystems delivered.
Production-ready code meeting the standard of the existing backend (typed, tested, reviewed).
Acceptance criteria
Measurable P95 latency reduction on the primary request path and on previously underperforming edge-case paths. Target: 30% or greater improvement on identified endpoints.
No background work on the response path for the migrated tasks.
Failed queue work retried with documented bounds. No silent drops.
PostgreSQL query plan regressions resolved on the endpoints in scope.
Core subsystems operate within measured time budgets with defined fallback behavior.
API contracts documented, implemented, and consumed by the frontend.
Diagnostic runbooks exist for latency incidents, LLM serving edge cases, and query plan regressions, usable by the team without contractor involvement.
All design documents and runbooks merged to the main branch.
What to send us
A short proposal covering:
Approach. How you would sequence the five workstreams, and which dependencies you see between them.
Timeline. Weeks to completion, with milestones and a proposed order of delivery.
Price. Fixed price for the project.
Acceptance criteria. Any you would adjust, add, or remove.
Risks and dependencies. What could go wrong, and what you will need from us.
Also, briefly answer in your cover note:
What is the most meaningful backend performance improvement you have shipped, and how did you measure it?
Describe a project you scoped down mid-engagement. What did you cut, and why?
How do you approach inherited code you disagree with?
Required
Demonstrated experience shipping production backend services at early-stage companies.
Strong command of Python: FastAPI, async/await, structured concurrency, background workers, and the failure modes of asynchronous systems.
Production experience with at least one queue system (Celery, Cloud Tasks, Pub/Sub, RQ, Temporal, or similar).
Deep PostgreSQL performance instincts, including direct experience diagnosing and resolving query plan issues under production load. pgvector experience preferred.
Experience with performance-sensitive subsystems and techniques for keeping them affordable (time-budgeted execution, pre-computed paths, caching strategies).
Familiarity with LLM pipelines, agentic frameworks, or complex state machines. An understanding that the model is one component within a larger system.
Strong intuition for the distinction between perceived and actual latency, and familiarity with techniques (speculative execution, optimistic streaming, partial results) that bridge the two.
Professional maturity regarding inherited code paired with a high standard for the quality of new work.
Preferred
Google Cloud Platform (Cloud Run, Cloud SQL, Cloud Tasks, GCS).
Server-Sent Events and other streaming protocols.
Self-hosted model serving experience.
PostgreSQL replication and read-replica routing.
Engagement details
Fixed-price project. Milestone-based payment (typical split: deposit on signing, milestone payments per workstream, final payment on delivery).
Code produced under this engagement is assigned to Elumai. Standard IP clauses apply.
Remote. Americas time zones preferred for alignment. Contractors handle their own tax and invoicing in their home jurisdiction.
The codebase is actively evolving as we approach general availability.
Successful engagement may lead to follow-on work. We are not committing to follow-on at this stage.
How to apply
Send your proposal and answers to [email protected] by May 26, 2026. CV or equivalent welcome. Keep submissions under three pages.