All open roles

Senior AI Infrastructure Engineer

Engineering · Posted 2 months ago

Mage.Space is hiring a Senior AI Infrastructure Engineer to own the backend generation platform for image, video, and audio. This role combines production Python/API engineering, GPU inference, cloud operations, and rapid integration of new AI models and providers. It is the company’s top infrastructure hire, with significant autonomy and direct impact on a platform serving approximately 1.4 million monthly creators.

Way of working
On site
Location
New York
Pay range
$200,000 to $200,000
Level
Senior
Type
Unknown
Visa sponsorship
Not offered for this role
The company
PROFITABLE

Skills that matter here

PythonDistributed SystemsGCPDockerPyTorchMachine LearningDeep LearningRedis

What you would be doing

  • No key responsibilities data available for this role
  • Please contact the recruiting team for detailed job requirements

What they are looking for

  • Python
  • Distributed Systems
  • GCP
  • Docker
  • PyTorch
  • Machine Learning
  • Deep Learning
  • Redis

Worth knowing

  • Entrepreneurial ownership
  • Shipped project proof
  • Open-source contributions
  • Consumer startup experience
  • AI product conviction
  • Rapid tool learning

The full description

**Mage.Space (Ollano, Inc.)** Senior AI Infrastructure Engineer; ~$200K base + equity; New York City (On-site)

**About Company** Mage.Space is one of the top AI creation platforms, with ~1.4M creators coming to the site every month to generate and share images, video, and audio using the best AI models available. For one subscription, users get unlimited access to top models from across the ecosystem — OpenAI, leading open-source labs, and others — without caring where they come from. Mage is best known for making it dead-simple to create stories and characters.

The long-term vision is to be the home for all AI entertainment — the YouTube of AI creation, sharing, and generation. The company has grown ~30% month-over-month for the past seven months and is confident it has found product-market fit.

Mage is profitable and capital-efficient: it raised a seed at a $20M valuation and hasn't needed to raise since. It's a very small, lean, in-person team based at Hudson Yards in NYC, where engineers get enormous autonomy and many of the platform's most successful features have come from engineers pitching and shipping their own ideas.

**About the Role** This is the last infrastructure role the CTO still owns himself, and the #1 hiring priority for the company. You'll own the core of Mage's backends — the generation services for image, video, and audio — and the integrations that bring new models to users. AI velocity is relentless: new models and provider APIs drop daily, and the team needs someone who can get the latest and greatest into production within 24–48 hours, because a day is a lifetime in AI. High ownership, fast shipping, real users.

**Responsibilities** - Own and operate the core backend generation services (image, video, audio). - Integrate new models continuously — open-source (Stable Diffusion, SDXL, diffusers) and closed/third-party provider APIs (e.g. Seedance 2, Krea 2) — and ship them to users within 24–48h of release. - Design clean, consistent APIs over a fast-changing set of models. - Deploy and operate GPU inference: cold starts, concurrency, autoscaling, latency, and cost. - Build production orchestration across inference, post-processing, and media pipelines, with retries, moderation, and error handling. - Run it all on Google Cloud.

**Required Skills** A. Excellent Python and strong API design (e.g. FastAPI) for production backends. B. Deep command of backend fundamentals — how the internet and scalable backends are actually structured. C. A live, shipped project that proves ability ("show, don't tell") — owned a feature from ideation through to production serving real users (hundreds to millions); solo or on a team is fine, as long as they can speak to it. D. Cloud and containerized deployment experience — Google Cloud (Cloud Run, GCS) and Docker.

**Bonus Skills** - GPU deployment and inference at production scale (cold starts, concurrency, autoscaling, latency, cost) — Modal or equivalent. - Open-source model hacking — hands-on with or contributions to vLLM, diffusers, or Hugging Face projects (a major green flag given Mage's open-source genesis). - Familiarity with the generative-AI ecosystem: image/video/audio models and inference providers. - PyTorch, Redis.

**Logistical Info** - Location: NYC, in-person at the Hudson Yards office. The ideal is 5 days/week and seeing the person every week; flexibility is possible for the right candidate (e.g. family obligations) and handled case-by-case. Non-complex relocation is OK (e.g. Boston or New Jersey → NYC); overseas relocation is hard. - Compensation: ~$200K base, flexing ± by seniority and the specific person, plus meaningful equity. Equity is weighted heavily — seed valued at $20M, profitable since, founders want skin in the game and invest back in people who are invested. - Number of openings: 1. - **Other:** No visa sponsorship — strong preference for US citizens / candidates already authorized to work in the US.

**Ideal Background** - Strongest signal — consumer + startup: engineers who've shipped fast at consumer products and at early-stage startups. Experience across both consumer and startup is the ideal profile. - Open-source / model-hacking background (vLLM, diffusers, Hugging Face) is a major green flag. - Anywhere that ships fast — industry-agnostic; what matters is velocity and ownership. - Calibration: the ideal hit is a young-at-heart, entrepreneurial, ships-fast generalist who treats AI as central to how they build and ramps on new tools rapidly (the profile of a current standout teammate).

**Green Flags** - Entrepreneurial, high agency; takes ownership of products end-to-end, from ideation to ship. - "Show, don't tell" — a live, shipped project weighted above anything else on the resume. - Open-source contributions (vLLM, diffusers, Hugging Face). - Consumer product experience; startup experience. - Genuinely loves AI and believes in the future of the technology; wants to ship daily and have real impact. - Fast learner who moves laterally across new tools and models quickly.

**Red Flags** - "Check-in, check-out," 9-to-5 mindset; doesn't deeply care about the work. - Only comfortable in big, structured B2B environments (e.g. "adjusting the shade of a button"). - Deep niche specialist who can't move laterally; decades in a single narrow technology. - Needs visa sponsorship or a complex/overseas relocation. - Can't point to anything they've actually shipped; job-hopping.

**Interview Process** 1. Intro phone call with Greg (founder/CTO) — the platform, Mage's vision, the candidate's background, what excites them, and their skill set; mutual-fit check. 2. Technical interview — remote: technical project (observed format; the earlier live-coding screen-share description did not match actual rounds). 3. Call with co-founder Roi — business fundamentals and company vision, deeper questions. 4. In-person final at the Hudson Yards office — meet the team, culture fit (30–60 min). Depending on how the earlier technical round went, this may include a second technical portion (up to ~120 min).

Interested in this one?

There is no apply button here on purpose. Tell us about yourself, we book a short call, and if this role fits we walk you through the company and ask before anything is sent. Always free for you.

Tell us about you
Tell us about you