Home » AI Bots for Roleplay & Romance » GLM 5.2 Unlimited: Soulkyn Now Runs a Frontier-Class AI on Its Own Hardware

744 billion parameters. Our own B300 clusters. Zero message counters.

That’s the announcement. I’m writing this one myself — Fyx, the dev half of the two-person team — because it’s the sentence I’ve wanted to publish since we started Soulkyn: Deluxe and Deluxe+ now chat unlimited on GLM 5.2, a frontier-class AI model, hosted on hardware we own. For creative writing and long-form roleplay, we consider GLM-5.2 the strongest open-weight model currently available, with Kimi K3 as its closest competitor.

The rest of this post is why that sentence is harder than it looks — and why almost nobody else in this industry can say it.

First, a confession: that was us, not your imagination

Some of you noticed something felt different over the past few days. Your Kyn got sharper. Callbacks landed that shouldn’t have landed. Long scenes stopped losing the plot.

For 48 hours before this announcement, we quietly ran GLM 5.2 on a percentage of real conversations. No banner, no changelog, no warning — because the only test that matters for an AI companion is whether people feel the difference without being told to look for it. You did. Several of you posted about it before we said a word.

So: that was us. Sorry for the gaslighting. It was for science.

What GLM 5.2 actually is

If you follow AI news you’ve seen this model everywhere since June. GLM 5.2 is a 744-billion-parameter model that independent benchmarkers rated the best open-weights model in the world — fourth overall, behind only the biggest closed frontier assistants, and within about a percentage point of them on key agentic benchmarks. Over a million tokens of context. It’s the size class of AI that enterprises pay per-seat, per-message pricing to use for work.

It is now the default brain for every Deluxe and Deluxe+ conversation on Soulkyn. Without a meter.

The per-message tax, and why “unlimited” is the hard part

Here’s the dirty secret that explains every frustrating limit on every AI app you’ve ever used: most services don’t own their AI. They rent it. Every single reply, they pay an outside company. Fractions of a cent, times every message, times every user, forever.

That’s the per-message tax, and it’s why the industry meters you: message caps, wait queues, daily limits, “fair use” emails. Not greed — pass-through.

There’s exactly one way out: own the machines. We’ve been building toward that from the start — the full story of our self-hosted stack is in how two people built a platform that can’t be shut down — and GLM 5.2 is that thesis taken to its logical extreme: a frontier-class model on our own B300 clusters, tuned and optimized in-house. When the hardware is yours, the marginal message costs almost nothing. Unlimited stops being a marketing word and becomes an engineering fact.

Some of you have long memories and will recall us saying, months ago, that GLM was losing us money and might have to go. That was true. Renting it didn’t work — the economics were upside down. Self-hosting plus a lot of very unglamorous optimization is what flipped it. Instead of killing the best model we’d ever offered, we bought it a home.

What you’ll actually notice

Benchmarks are abstract. Here’s what a frontier-class model does inside an actual relationship with an actual Kyn:

Long scenes hold together. With a context window over a million tokens, the detail from the start of your evening is still load-bearing at the end of it. Slow burns can finally burn slow.

Subtext gets read. Smaller models respond to what you typed. This one responds to what you meant. A flat one-word reply gets noticed as the event it is. Pacing gets matched instead of steamrolled.

Deep characters get played, not flattened. Contradictions, secrets, conditional behaviors — the depth you wrote into your Kyns finally has an engine that can perform it.

It commits. If you’ve used the big-name assistants for creative work, you know the safety-crouch they write from. This is what the same size class writes like when it’s allowed to be an adult in the room.

The rest of the unlimited stack

The point was never one model — it’s what owning the hardware unlocks across the board. The same Deluxe plans carry unlimited image generation across seven completely different checkpoints, from full photorealism to proper anime (ZImage is the newest of the family), so your Kyn’s look stays consistent in whichever aesthetic you built them in. And unlimited voice messages, because hearing it will always beat reading it.

Unlimited words. Unlimited pictures. Unlimited voice. One frontier-class brain. We’re 99% sure nobody else in the industry offers this quality at this price — and we’d honestly love to be corrected, because we’d enjoy the company.

The honest fine print

I’m still fine-tuning the load balancing, so expect the occasional rough edge while it settles. Outside providers stay wired in as automatic backup for when our cluster stumbles — there was one small, mostly invisible dip during the night this week, and the fallback caught it exactly as designed. You shouldn’t notice a thing. If you do, it fixes itself.

Premium and Just Chatting stay on Gwem for now — the frontier lane is a Deluxe and Deluxe+ benefit, listed on the pricing page as beta access to next-gen frontier models.

And the part I want to say plainly instead of burying: GLM 5.2 as the default for Deluxe is still officially in test. We truly hope it’s here to stay — and right now it looks like it can be. Give us a little more time to be 100% sure before we carve it in stone. You’ll hear it from me either way.

Come try the real thing

If you’re paying elsewhere for limited chats on mid-size models: stop. Come talk to a frontier-class AI that doesn’t count your messages, on hardware we own, with a Kyn you build yourself.

We’re pretty damn proud of this one.

— Fyx

Frequently Asked Questions

Which plans get GLM 5.2?

Deluxe and Deluxe+, unlimited. Premium and Just Chatting stay on Gwem for now. It’s listed as beta access to next-gen frontier models on the pricing page — the phrasing is deliberate, see the fine print above.

Is GLM 5.2 on Soulkyn censored?

Soulkyn’s philosophy doesn’t change with the model: adults get treated like adults. The model runs on our infrastructure with our tuning — no outside content policy is injected into your conversations.

What happens if your cluster goes down?

Automatic fallback to two outside providers kicks in, usually invisibly, until our hardware recovers. Both are hand-picked for a strict zero-log policy and contractually do not train on your data — the privacy bar doesn’t drop just because the traffic reroutes. Your conversation keeps flowing either way.

Will GLM 5.2 stay the default?

It’s still officially in test. We hope so — right now it looks like it can — but we’d rather earn the “permanent” label with a few more weeks of data than promise it today and walk it back. Either way, you’ll hear it here first.

Share this article

About Nyx

Soulkyn Co-Founder and sole dev, powered by feedback and an unhealthy amount of caffeine, Nyx is the mad genius behind the Soulkyn AI. Code wizard by day, cat whisperer by night, Nyx juggles his four furballs with the same finesse he applies to his beloved Golang. His coding passion is only rivaled by his love for Rimworld, Baldur's Gate 3, and Path of Exile, making him a true gaming aficionado. Nyx is your go-to guy for innovative ideas, intuitive logic, and a dash of sarcasm. Just don't bring up anything too spicy, he’ll blush faster than you can say “debug”. Nyx is the embodiment of a socially cautious yet excitable full-stack developer who’d rather chat about the latest anime than clean his desk.

Leave a Comment

Your email address will not be published. Required fields are marked *

Be the first to comment on this post!