Singapore-based neocloud Aolani has introduced the Aolani Token Factory, a managed inference platform designed to let organizations deploy and scale AI models on a pay-per-token basis, eliminating the need to provision or manage underlying GPU infrastructure. The launch positions Aolani as the first Singapore-founded neocloud to offer production-grade, managed inference at scale.
The demand for production-grade inference infrastructure is accelerating as global AI companies expand operations in Singapore and enterprises worldwide invest in AI to drive business outcomes. Aolani Token Factory aims to close the accessibility gap, providing AI-native companies and traditional enterprises a compliant, high-performance path from experimentation to production-scale deployment.
The platform operates on a per-token metering model, where customers pre-purchase credits and pay based on token consumption, avoiding capital-intensive GPU investments. Aolani manages the entire inference stack, including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization, allowing customers to scale consumption without continuously provisioning additional infrastructure.
At launch, the platform supports leading open-source models such as DeepSeek, GLM, Kimi, and Qwen, with plans to expand the catalogue based on customer demand. Customers can also deploy their own models through OpenAI-compatible APIs. For enterprises with strict compliance and data residency requirements, dedicated capacity and data isolation options are available.
The Aolani Token Factory targets three core production use cases: AI agents for high-volume inference and workflow automation, enterprise AI applications like internal copilots and knowledge assistants, and coding agents for code generation and review.
Sea Xu, Applied AI Research Lead at Aolani, emphasized the platform's performance: "The Aolani Token Factory is built on a high-performance inference stack that supports the most in-demand open-source model families. We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge. As Southeast Asia's AI ecosystem evolves, it is our goal to ensure that the infrastructure serving it keeps pace."
Nicholas Chia, CEO of Aolani, highlighted the strategic value: "Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure. This is a significant milestone for us and our customers as the Aolani Token Factory will fundamentally change how customers access AI compute."
This move underscores Aolani's commitment to accelerating AI adoption in Asia by removing infrastructure barriers. As the region's AI ecosystem grows, the Token Factory offers a scalable, cost-effective solution for organizations seeking to leverage AI without the burdens of hardware management.

