This job is closed
Applications are no longer available for this announcement. Explore current related opportunities below.
Job description
Engagement Type: Fractional (Advisory) to start, opportunity to move into full time role following initial engagement.
Compensation: $250–$500 per hour (depending on experience)
Time Commitment: 5–10 hours per week to start
Initial Duration: 4–12 weeks
Location: United States only. Boston area preferred (within a few hours’ drive) for occasional in-person meetings with founder.
About Us
We are a bootstrapped AI company building a high-throughput research and intelligence engine for the public-sector market.
We ingest and analyze public records at scale to identify government agencies entering active buying cycles for our clients’ solutions.
$100K revenue in first 6 months
80%+ retention rate
Annual agreements with brand-name companies
Projecting $500K revenue this year
On track for profitability
Lean, senior team
Our initial internal AI research platform was built by one senior engineer and has successfully supported early customer growth.
Now, as customer volume increases and use cases diversify, we are adding senior talent and need to architect the next generation of our internal and external tools to support 100x current capacity.
The Role
We are seeking a Principal-level AI Infrastructure Engineer to:
Review and pressure-test our current architecture
Design a next-generation, LLM-agnostic system capable of 100x scale
Help guide and support implementation of that architecture
This is an engineering and systems role — not a management position.
There is a clear opportunity to evolve into a full-time lead engineer role after the initial engagement for the right person.
Engagement Phases
Phase 1 – Architecture & Codebase Review
Review current system architecture and codebase
Evaluate LLM usage patterns and token efficiency
Assess API orchestration, rate limiting, batching, queuing, and retry logic
Identify bottlenecks, fragility points, and scaling risks
Deliver a structured architectural assessment
Phase 2 – Next-Generation Architecture Design (100x Scale)
Design a scalable, LLM-agnostic AI architecture
Plan for 100x current throughput
Architect for:
Token and inference cost control
Provider abstraction (closed + open models)
Resilience and fallback routing
Distributed job orchestration
High-concurrency environments
Advise on local vs hosted inference strategy
Evaluate GPU cost, latency, and inference tradeoffs
Phase 3 – Implementation Support
Guide implementation of the new architecture
Review critical technical decisions during buildPressure-test scaling assumptions
Help prevent structural technical debt
Required Experience
Built and scaled LLM-agnostic systems
Scaled AI or API-heavy systems under real production load
Experience operating at billion-token-per-day scale (or comparable throughput environments)Deep expertise in rate limits, retries, batching, queuing, and distributed failure modes
Designed token-efficient architectures
Worked with both closed-model providers and open-source models
Deployed models locally or within controlled infrastructure
Evaluated GPU cost, latency, and inference tradeoffs
Preferred Background
Former CTO, Principal Engineer, or Staff Engineer
Experience at a VC-backed startup with a successful outcome or a major public technology company
History of scaling AI-native or API-intensive systems
Comfortable collaborating closely with a technically involved founder and senior engineer
Systems-oriented, pragmatic, and product-aware
Why This Is Interesting
Strong early product-market fit
Real production workload and scaling pressure
High ownership and architectural influence
Lean team with meaningful upside
Clear path to deeper involvement for the right person
#J-18808-Ljbffr
Worksite address
boston, MA, 02298, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.