Skip to main content

Baseten

AI inference and infrastructure

Serve and scale open-source and custom AI models on the fastest, most reliable inference platform.

Hiring now

1open role

Is Baseten hiring?

Yes. As of October 6, 2026, Baseten has 1 open role on Metaintro.

Open roles

Newest first. Pay shown when the posting lists it.

Customers

All 39 companies Baseten names as customers.

Case studies27

  • Writer
  • OpenEvidence
  • You.com
  • Superhuman
  • Wispr Flow
  • Apogee Behavioral Medicine
  • Sully.ai
  • Speechify
  • Scaled Cognition
  • Rime
  • Praktika
  • Patreon
  • OpenCode
  • Posit
  • Latent Health
  • SpeechifyAI
  • Parallel
  • Parallel Web Systems
  • Notion
  • Gamma
  • EliseAI
  • Poolside
  • Hebbia
  • AlliumAI
  • Oxen AI
  • Latent
  • World Labs

Other customers12

  • Zed Industries, case study
  • Clickup
  • Wispr
  • Zed Industries
  • AI Core Team
  • Bria
  • Subconscious
  • Inception
  • Praktika AI
  • Amp
  • SullyAI
  • Scaled Cognition in grayscale

What their customers say

Quotes from people at Baseten's customers.

“With the launch of Brain MAX we've discovered how addictive speech-to-text is - we use it every day and want it everywhere. But it's difficult to get reliable, performant, and scalable inference. Baseten helped us unlock sub-300ms transcription with no unpredictable latency spikes. It's been a game-changer for us and our users. I want the best possible experience for our users, but also for our company. Baseten has hands down provided both. We really appreciate the level of commitment and support from your entire team.”

Nathan Sobo · Co-Founder · Zed Industries, case study

“Inference for custom-built LLMs could be a major headache. Thanks to Baseten, we're getting cost-effective high-performance model serving without any extra burden on our internal engineering teams. Instead, we get to focus our expertise on creating the best possible domain-specific LLMs for our customers.”

Waseem Alshikh · CTO and Co-Founder · Writer

“With the launch of Brain MAX we've discovered how addictive speech-to-text is - we use it every day and want it everywhere. But it's difficult to get reliable, performant, and scalable inference. Baseten helped us unlock sub-300ms transcription with no unpredictable latency spikes. It's been a game-changer for us and our users.”

Mahendan Karunakaran · Head of Mobile Engineering · Clickup

“With Baseten Embeddings Inference, we immediately saw 3x speed improvements. Doctors rely on speed when treating patients, and that improvement has been critical to our product experience. 160 millisecond latency is crazy.”

Jagath Jai Kumar · Full Stack Engineer · OpenEvidence

“With Baseten, we gained a lot of control over our entire inference pipeline and worked with Baseten's team to optimize each step.”

Sahaj Garg · Co-Founder and CTO · Wispr
Show 25 more quotes

“I want the best possible experience for our users, but also for our company. Baseten has hands down provided both. We really appreciate the level of commitment and support from your entire team.”

Nathan Sobo · Co-Founder · Zed Industries

“We’re really appreciative of the support and just moving so quickly on this. The latency improvements have been so impressive, and on such a short timeline.”

Antonio Scandurra · Co-founder

“Our engineering team just wants to work with you guys. They told me really straightforwardly. And that goes a long way.”

Nathan Sobo · Co-founder

“The ‘let’s get this thing going and then figure out the contract details later’ attitude is so much better than some of your competitors who wanted to tie us down first. We’re really impressed with the level of commitment and support from your engineers.”

Nathan Sobo · Co-founder

“We think open-weight models plus web search is a powerful alternative to routing everything through frontier model providers, at a fraction of the price and comparable quality for our developers.”

Saahil Jain · CTO · You.com

“Inference for custom-built LLMs could be a major headache. Thanks to Baseten, we’re getting cost-effective high-performance model serving without any extra burden on our internal engineering teams. Instead, we get to focus our expertise on creating the best possible domain-specific LLMs for our customers. The improved performance means Writer customers can enjoy faster response times and higher token throughput when interacting with Palmyra models.”

Writer

“Superhuman is all about saving time. With Baseten, we're delivering a faster product for our customers while reducing engineering time spent on infrastructure.”

Loïc Houssier · CTO

“Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.”

Loïc Houssier · CTO

“The deployment mechanism is so good that I was able to self-serve 95% of what I needed, and the Baseten team was incredibly responsive every time I had a question.”

Agustín Bernardo · Senior AI Engineer

“We have ambitious goals for a best-in-class customer experience powered by consistent low-latency inference, but I didn’t think it would be a good use of our team’s bandwidth to build an inference platform in-house.”

Loïc Houssier · CTO

“We no longer have to reserve GPUs just in case we see usage spikes during viral moments. This scalability is critical as Flow gains wider adoption with its new team and enterprise offerings.”

Wispr Flow

“We no longer have to reserve GPUs just in case we see usage spikes during viral moments.”

Sahaj Garg · Co-Founder and CTO

“With Baseten and AWS, we have providers that we trust, and more importantly, our users can trust.”

Sahaj Garg · Co-Founder and CTO

“Llama is controllable and customizable, which lets us focus on the output.”

Sahaj Garg · Co-Founder and CTO

“We measure latency on a p90 or p99 basis for each user; we don’t care at all about p50. We’re optimizing the p99 experience for the p99 user.”

Sahaj Garg · Co-Founder and CTO

“I write differently when texting my mom, my wife, and my colleagues. One early challenge was encapsulating this behavior and making it controllable, matching Flow’s output to the user’s context and preferences.”

Sahaj Garg · Co-Founder and CTO

“We’re getting paid faster, our compliance risk is much lower, and our administrative overhead is down. Most importantly, we have a happier, more effective clinical team that’s providing better care to more engaged patients. Sully has become a major competitive advantage for us in the marketplace.”

Derek Ayers · CMO · Apogee Behavioral Medicine

“At our scale, inference efficiency matters as much as model quality. Baseten enabled us to cut costs by 90% while delivering significantly faster, more predictable performance. That efficiency is what allows us to add millions of productive minutes back into our customers.”

Ahmed Omar · Co-Founder and CEO · Sully.ai

“The open-source ecosystem is moving incredibly fast, and Baseten moves just as fast with it. Having access to newly released models like GPT OSS 120b within days, already optimized for production, gives us a real competitive edge. It means we can continuously improve model quality without slowing down product development.”

Amit Kumthekar · Head of Research · Sully.ai

“Working with the Baseten team was a no-brainer. Together, we decreased our model latency by over 50%, reduced our cost per million characters by over 44%, and delivered the highest uptime of any inference provider we know of. Baseten has enabled Speechify to provide the highest-quality, lowest-latency, and most cost-efficient text-to-speech AI voice models in the world to consumers, developers, and enterprises.”

Rohan Pavuluri · Chief Business Officer

“Baseten collapsed a per-model tower of Terraform, Envoy, and Filestore into a single "truss push".”

Kai Krause · VP of Engineering and AI

“A researcher shipped SIMBA 3.0 vLLM into production by themselves. That would have taken days of work from our entire AI Platform team with our old infrastructure stack.”

Kai Krause · VP of Engineering and AI

“The reason we came to Baseten in the first place was the complexity of managing our own infrastructure. Our priority is continuing to deliver the best TTS platform for our 60M+ users, and we didn’t want our inference infrastructure to stand in the way of that.”

Kai Krause · VP of Engineering and AI

“Scaled Cognition has always been known for the quality of its agents, but performance is just as important. We partner with Baseten to ensure the lowest possible latency for our models and agentic workflows. It’s been a major differentiator for us in the market. Quick implementation was key to maintaining momentum with existing prospects, so Baseten’s engineers collaborated with Scaled Cognition to benchmark various solutions that leveraged their existing workloads and rapidly optimized them for active POCs.”

Scaled Cognition

“We really appreciate the collaboration with Baseten. The Hybrid Cloud solution and access to cutting-edge GPUs while working with our existing AWS commitments were key to our success on launch day and beyond. Beyond that, Baseten’s developer experience has been a favorite across our team.”

Jordan DeLoach · VP of Engineering

AI launches and partners

What Baseten has launched and who it works with on AI.

  • LaunchSOTA text-to-speechBaseten built real-time audio streaming for AI phone calls, voice agents, and translation.
  • LaunchPre-optimized Model APIsBaseten offers optimized AI models for production use.
  • LaunchInference PlatformBaseten offers infrastructure for serving open-source, custom, and fine-tuned AI models.
  • LaunchTranscription and diarizationBaseten offers transcription and diarization, including streaming support for real-time voice AI use cases.

Skills they ask for

Most frequent across open postings.

Tech stack

CMS
HubSpot CMS
Libraries
Tailwind CSS
Analytics
Google Analytics 4
Frameworks
Next.js

Company facts

Founded
2019
Industry
AI inference and infrastructure
Sector (NAICS)
Information
Website
baseten.co

Questions about Baseten

Is Baseten hiring?
Yes. As of October 6, 2026, Baseten has 1 open role on Metaintro. The most are in Security (1).
Who are Baseten's customers?
Baseten names 39 customers, including Writer, OpenEvidence, You.com, Superhuman, and Wispr Flow.
Return to navigation