<aside>
📍
Role at a glance
- Location: San Francisco, CA. Onsite strongly preferred.
- Commitment: 3 to 6 months. Full-time or 20 hours per week part-time.
- Compensation: $25 to $35 per hour.
- Work Authorization: Astorie can provide standard employer documentation for students who are eligible to obtain CPT authorization through their school.
</aside>
About Astorie
At Astorie, we are building the intelligent canvas where creative agents can see, reason, and act throughout the creative process. They proactively explore ideas with you, offer aesthetic inspiration, and learn your taste and way of working over time. Our goal is to build the creative infrastructure that closes the distance between imagination and expression.
About the Role
This role owns the production integration and operational reliability of model APIs behind Astorie's model gateway. You will onboard providers and model endpoints, maintain routing and fallback behavior, monitor latency and reliability, investigate incidents, track cost and quota constraints, and automate recurring infrastructure workflows.
You will begin with established SOPs and real production workflows. As you build context and demonstrate strong execution, you will help improve, automate, and extend those workflows.
What You'll Do
- Integrate new model providers and model endpoints into the gateway
- Validate capability compatibility and production readiness before rollout
- Configure and maintain routing and fallback behavior across providers
- Tune retry, rate-limit, quota, and error-handling policies
- Benchmark provider latency, cost, throughput, output quality, and reliability
- Monitor logs, dashboards, and alerts to catch failed or degraded model jobs
- Investigate incidents and recurring failures down to root cause
- Turn findings into fixes, regression tests, observability, or clearer playbooks
- Keep records of model capabilities, pricing, quotas, and performance accurate
- Build scripts and internal tools that automate onboarding, benchmarking, and validation
What We're Looking For
- Evidence of personally building or shipping at least one API, backend service, or infrastructure project that you can show and discuss
- Solid understanding of core systems concepts, including APIs, queues, databases, rate limits, retries, and common failure modes