GLM Coding Plan 2026: Prices, Limits and Setup
October 5, 2026 · 56 min read · Ox Alpha
Z.ai's subscription for coding agents, taken apart: what Lite, Pro and Max cost in October 2026, how the credit meter really runs, the rules that can cost you the account, and when paying per token is the better deal.
The cheapest GLM Coding Plan costs $18 a month, and by Z.ai's own estimate its weekly allowance covers 48 to 97 million tokens of the flagship GLM-5.3 - Z.ai. Bought at Z.ai's pay-as-you-go prices, that much GLM-5.3 would cost roughly $70 to $140 a month, which is the whole pitch in one comparison. It is why the plan has become the default way developers run GLM inside Claude Code, Codex, Cline and a dozen other agents, and why so many of them arrive at this site asking what the plan really includes.
The trouble is that the plan most reviews describe no longer exists. In 2026 alone it moved from prompt counts with no weekly cap, to prompt counts with 5-hour and weekly caps, to a token-metered credit system that took over on July 30 - Z.ai. Prices moved too: the Pro tier was $72 a month on monthly billing in April and is $80 now. And the headline promise, savings of "up to 92%", is measured against Z.ai's own list price, which for a cache-heavy coding workload is one of the most expensive ways to buy GLM-5.3 on the open market.
This guide does the work the pricing page does not. It explains exactly how credits are charged, converts each allowance into dollars at today's prices, shows where the plan beats pay-as-you-go and where it does not, walks through the setup for Claude Code, Codex and Cline, and spells out the usage rules that get accounts restricted. Every price, limit and rule below was checked against Z.ai's live documentation and subscription page on October 5, 2026. For the models themselves, our pages on GLM-5.3 and the whole GLM family carry the specs and benchmarks; this guide is about the subscription that sells access to them.
Contents
- What the GLM Coding Plan is, and why it exists
- Prices and tiers in October 2026
- How credits work: the formula, the multipliers and peak hours
- What the quota is really worth in dollars
- The models you actually get
- Setting it up in Claude Code, Codex, Cline and other tools
- How to make the credits last
- The rules that can cost you the subscription
- Privacy, data and jurisdiction
- The plan against the alternatives
- The Team plan
- How the plan changed in 2026, and what comes next
- The verdict: which tier, if any, to buy
1. What the GLM Coding Plan is, and why it exists
The GLM Coding Plan is a monthly subscription from Z.ai, the international brand of Zhipu AI, that pays for GLM models inside AI coding tools instead of per token - Z.ai. You create an API key, point a supported agent such as Claude Code or Cline at Z.ai's coding endpoint, and every call draws on a fixed allowance instead of your card. All three tiers include the same two models, the flagship GLM-5.3 and the faster GLM-5.3-Flash, plus four MCP tools for image understanding, web search, web reading and repository reading. The tiers differ only in how much you can use, and one allowance is shared across every tool you connect.
What the plan is not matters just as much, because it decides whether the plan can serve you at all. It is not an API plan for your own software: Z.ai's subscription terms forbid using the quota to call models "from your own applications, bots, websites, SaaS products or other systems" - Z.ai subscription terms. It cannot be shared with a colleague, it does not fall back to your account balance when the allowance runs out, and it is not refundable. If what you need is GLM inside a product you are building, you need the pay-as-you-go API or a third-party host, and our API migration guide walks through the drop-in settings for both.
Why a flat-rate coding plan exists at all
A coding agent is an unusual customer for a model provider. Every step of an agent loop re-sends the conversation so far (the system prompt, the tool definitions, every file it has read), so a single task can push millions of tokens through the model, and nearly all of them are identical to the previous call. Providers price those repeated tokens as cached input: Z.ai charges $1.40 per million fresh input tokens for GLM-5.3 but $0.26 per million cached ones - Z.ai pricing. Under per-token billing, a heavy agent user faces a bill that is large, dominated by re-reads, and impossible to predict a month ahead.
Developers dislike that kind of bill, and they have an obvious comparison: Anthropic sells Claude Pro and Claude Max as flat subscriptions that include Claude Code - Claude pricing. A capped subscription solves the problem from both sides. The developer gets a known monthly price; the provider gets committed revenue and a bounded worst case, because the 5-hour and weekly caps stop any one subscriber from consuming unlimited compute. The credit multipliers, explained in section 3, then let Z.ai price each model and each hour of the day separately without changing the sticker price. And Z.ai gets something less visible but just as valuable: a developer whose agent, prompts and habits are tuned to GLM is a developer who keeps renewing.
That logic tells you who the plan is built for:
- Daily agent users whose GLM usage at API prices would run past the plan's price within a week or two.
- Claude Code users who want a cheaper model for routine work without leaving the tool.
- Developers in the Americas, whose working day falls almost entirely in Z.ai's discounted off-peak window.
All three groups get real value because the plan's economics reward steady, heavy, cache-friendly use, which is exactly what an agent produces when you work with it every day. The same logic explains who the plan quietly does not serve. Light users pay for an allowance they never touch: a developer who runs a handful of agent tasks a week will usually spend less on the pay-as-you-go API. App builders are excluded outright, because anyone hoping to power a chatbot or SaaS feature with a cheap subscription is breaking the terms on day one.
Keep both of those exceptions in mind as you read the price section, because "cheaper than the API" is only true above a usage threshold, and section 4 calculates that threshold precisely for each tier. The threshold is lower than most people expect at Z.ai's own prices and much higher against the cheapest third-party hosts, which is the single most important nuance in deciding whether to subscribe at all.
2. Prices and tiers in October 2026
Z.ai sells three individual tiers, Lite, Pro and Max, each available monthly, quarterly or yearly. The subscription page lists Lite at $18 a month, Pro at $80 and Max at $168 on monthly billing, and shows a 20% discount for quarterly billing and 30% for yearly - Z.ai subscription page. On the yearly setting the page displays the monthly equivalents as $12.60, $56 and $117.60. Every tier includes the same models and tools; what changes is the size of the two credit windows and a few service promises.
The allowances come from Z.ai's own documentation, and they are the numbers that actually matter, because the price only buys you a meter. Each plan has a 5-hour credit limit and a weekly credit limit, and both apply at once - Z.ai. The table puts the prices and the allowances side by side.
| Tier | Monthly billing | Yearly billing (per month) | 5-hour credits | Weekly credits | Z.ai's description |
|---|---|---|---|---|---|
| Lite | $18 | $12.60 | 2,000 | 10,000 | Lightweight iteration on small repos |
| Pro | $80 | $56 | 12,000 | 60,000 | Day-to-day development on mid-sized repos |
| Max | $168 | $117.60 | 28,000 | 140,000 | Advanced users on mid-to-large repos |
This is what the pricing cards looked like on the day this guide was checked, with yearly billing selected. The crossed-out figures are the monthly-billing prices, and the "6× Lite usage" and "14× Lite usage" labels are the ratios of the weekly credit allowances.
Beyond the allowance, the cards promise different service levels. Lite advertises "rolling access to the latest flagship models", support for 20+ agent tools and "default data privacy". Pro adds priority access to new models, "a curated selection of MCP tools" and "faster generation speeds". Max adds first access to new flagship models and "dedicated resources during peak times", which matters more than it sounds if your working day overlaps Z.ai's busy hours. None of these promises is quantified, so treat them as tie-breakers rather than reasons to buy a tier.
The arithmetic that is quantified favours the bigger tiers. Per 1,000 weekly credits, Lite costs $1.80 on monthly billing, Pro $1.33 and Max $1.20, so Pro buys each credit about 26% cheaper than Lite, and Max about a third cheaper. That does not make Max the rational default, because unused credits are worth nothing. It means that once you know you will use several times Lite's allowance, the bigger tier is the cheaper way to buy it: six Lite subscriptions would cost $108 a month for the same weekly credits that Pro sells for $80.
Discounts, referrals and team pricing
Two other price levers are worth knowing before you pay. The first is Z.ai's referral program: a new subscriber who signs up through someone's invitation link gets 10% off their first subscription order, and the person who invited them earns credits worth 10% of what the new subscriber paid, released once three invitees have paid - Z.ai referral rules. Those credits can pay for renewals or API calls but cannot be withdrawn. The second lever is timing: Z.ai runs promotions often enough (two ran from September until October 7, covered in section 3) that the capacity you see in any given week may include a temporary deal.
Teams buy seats instead of individual plans. The Team tab of the same subscription page prices a Standard Seat at $88 per seat per month and a Premium Seat at $188, with a minimum of two seats and a 10% saving advertised for annual billing - Z.ai subscription page. Section 11 covers what the seats include, which is more than a bigger allowance: central billing, usage dashboards, an optional pay-as-you-go overage, and a default exclusion of team data from model training.
Finally, a word on what the prices used to be, because it explains a lot of conflicting advice online. In April 2026, Z.ai's own migration notice listed monthly prices of $18, $72 and $160 - Z.ai migration notice. Lite has held its price; Pro and Max rose by about 11% and 5% when the credit system arrived, and the way usage is measured changed completely. Any review written before July 30, 2026 is describing a different product, even when the tier names match.
3. How credits work: the formula, the multipliers and peak hours
Credits are Z.ai's unit for metering usage, and unlike the old prompt counts they are tied directly to tokens. The formula is published, which makes the plan unusually predictable once you understand it: credits equal (input tokens × input multiplier + cached input tokens × cached multiplier + output tokens × output multiplier) / 10,000 - Z.ai. MCP tool calls are charged separately, as a flat number of credits per call. The multipliers differ by model, and they are the most important numbers in this guide.
The multipliers are effectively an internal price list. For GLM-5.3 they are 6.9 for fresh input, 1.7 for cached input and 24 for output, so a million cached tokens cost 170 credits, a million fresh input tokens 690 credits, and a million output tokens 2,400 credits. Output is fourteen times as expensive per token as cached input, which roughly mirrors the API price sheet, where GLM-5.3 output costs $4.40 per million against $0.26 for a cache hit. GLM-5.3-Flash costs about a third as many credits per token as GLM-5.3 across the board.
| Usage | Input multiplier | Cached input multiplier | Output multiplier |
|---|---|---|---|
| GLM-5.3 | 6.9 | 1.7 | 24 |
| GLM-5.3-Flash (also the vision MCP) | 2.3 | 0.56 | 8 |
| Web Search, Web Reader, Zread MCP | - | - | 1.2 per call |
Two consequences follow directly from that table. The first is that cache hits are the plan's lifeblood: a coding agent that re-sends 120,000 tokens of context on every step pays mostly at the 1.7 rate, which is what makes the advertised token allowances possible at all. The second is that reasoning is expensive: GLM-5.3's thinking tokens are output tokens, billed at 24, and the model defaults to its maximum reasoning effort. Section 7 turns both facts into habits that stretch an allowance.
Peak hours and the 50% off-peak rate
The meter does not run at the same speed all day. Outside peak hours, model usage is charged at 50% of the standard rate, and peak hours are Monday to Friday, 14:00 to 18:00 Singapore time (UTC+8) - Z.ai. That is four hours a day on weekdays, or twenty hours a week out of 168, so most of the week is discounted. Whether your own working hours fall inside the expensive window depends entirely on where you live.
| Location (time zone on October 5, 2026) | Peak window on weekdays | Peak window after daylight saving ends |
|---|---|---|
| Singapore, Beijing (UTC+8) | 14:00 to 18:00 | unchanged |
| India (UTC+5:30) | 11:30 to 15:30 | unchanged |
| Berlin, Paris (CEST, UTC+2) | 08:00 to 12:00 | 07:00 to 11:00 |
| London (BST, UTC+1) | 07:00 to 11:00 | 06:00 to 10:00 |
| New York (EDT, UTC-4) | 02:00 to 06:00 | 01:00 to 05:00 |
| San Francisco (PDT, UTC-7) | 23:00 to 03:00, the night before | 22:00 to 02:00, the night before |
The pattern is stark. A developer in the Americas working normal hours almost never touches the peak window, so in practice they get the off-peak rate on nearly everything and their allowance is effectively doubled. A developer in Europe pays full rate through most of the working morning, and one in India pays full rate across the middle of the day. If you are in Europe and can schedule long autonomous agent runs for the afternoon or evening, you halve their credit cost for free.
The two windows: 5 hours and 7 days
Both limits apply at the same time, and either one can stop you. The 5-hour window is rolling: credits are "dynamically refreshed", and each chunk of usage frees up again five hours after you spent it. The weekly window starts when you subscribe and resets every seven days from that moment, not on a calendar week - Z.ai FAQ. When either is exhausted, calls from your tools fail until it refreshes; the plan never dips into your account balance.
The relationship between the two numbers is deliberate. On every tier the weekly allowance is exactly five times the 5-hour allowance (2,000 and 10,000 for Lite, for example), so a subscriber who empties the 5-hour window five times has used the whole week. In practice that means the plan supports roughly one long, intense session a day on a five-day week, or a steadier pace spread across seven days, and a developer who works in bursts will hit the 5-hour wall long before the weekly one.
The flow below traces one request from your agent to the moment its credits are deducted.
The routing at the top of that diagram surprises people. Z.ai's documentation states that requests for GLM-5.2 or GLM-5.1 are automatically routed to GLM-5.3, and requests for GLM-4.7 to GLM-5.3-Flash - Z.ai. Configuring an older model id does not give you the older model; it gives you its replacement, billed at the replacement's rate. Section 5 explains why that matters for anyone who chose GLM-5.2 for its license.
A worked example: what one agent call costs
Abstract multipliers become useful once you price a real call. Take a typical step in a Claude Code session: the agent re-sends 120,000 tokens of context, 95% of which the provider recognises from the previous call, and writes back 2,000 tokens of reasoning and code. On GLM-5.3 that call costs (6,000 × 6.9 + 114,000 × 1.7 + 2,000 × 24) / 10,000 = 28.3 credits at the peak rate and 14.2 credits off-peak. At Z.ai's API prices the same call would cost about 4.7 cents.
Divide the windows by that figure and the tiers become concrete. Lite's 2,000-credit window covers about 71 such calls in five hours at the peak rate and about 141 off-peak; Pro covers about 424 and 847; Max about 989 and 1,977. Z.ai's own older documentation estimated that one prompt to a coding agent triggers 15 to 20 model calls - Z.ai. On that estimate, Lite supports about four substantial prompts per 5-hour window at the peak rate (3.5 to 4.7) and seven to nine off-peak, while Max supports somewhere between 50 and 130.
The context size is the variable that moves those numbers most, and it moves them a lot. Because every call re-sends the whole context, a session that has grown to 400,000 tokens costs about 83 credits per call at the peak rate even when 95% of it is cached, three times the cost of the same step at 120,000 tokens. The chart shows how quickly that climbs.
The practical reading is that GLM-5.3's 1M-token context is something you can afford to use occasionally, not by default. A session allowed to grow to 800,000 tokens burns about 160 credits per step at the peak rate, so Lite's entire 5-hour window would last about a dozen steps. Compacting or clearing the conversation between tasks is the single biggest lever on how long a plan lasts, and section 7 shows how to do it in each tool.
The autumn promotions, and how they skew a trial
Two temporary offers ran when this guide was first published, and both ended on October 7, 2026. From September 25, all-day usage on every plan was charged at the off-peak rate. And under a separate campaign that began on September 3, GLM-5.3-Flash used through Z.ai's own ZCode and AutoClaw apps between 23:00 and 09:00 Singapore time consumed no quota at all, while other supported agents got double the usual allowance in that window - Z.ai campaign rules. That campaign was first due to end on September 20 and was extended once. When this section was updated on October 7, Z.ai's documentation announced no successor offer, so the standard rates in the tables above are what a new subscriber pays.
Offers like these come back often, and they distort the trial that most people run before deciding on a tier. A week spent testing Lite during an all-day off-peak promotion shows you roughly twice the capacity you will get afterwards if you work during Asian afternoon hours, and unlimited overnight Flash in ZCode says nothing about what a Claude Code session costs. If you are evaluating the plan around a promotion, measure your consumption in credits (the plan overview page shows it) and price your normal week at the standard rates from the tables above.
4. What the quota is really worth in dollars
Z.ai publishes its own translation of credits into tokens, and it is the right place to start because it shows the assumptions behind the marketing. The documentation estimates that at a 95% cache hit rate, Lite covers 48 to 97 million GLM-5.3 tokens a week, Pro 290 to 580 million and Max 676 to 1,352 million - Z.ai. The low end of each range assumes every call lands in peak hours, the high end assumes every call is off-peak. On GLM-5.3-Flash the same credits stretch about three times further, to 146 to 292 million tokens a week on Lite.
Those ranges are not arbitrary. Working backwards from them, they correspond exactly to a workload in which about 99.5% of tokens are input (95% of that served from cache) and about 0.5% are output, which is a fair picture of an agent that re-reads a large context and writes a little code on each step. That matters because your own mix decides your real allowance. An agent that reasons at length or writes long files spends a larger share on output at 24 credits per 10,000 tokens, and a session that keeps changing its early context loses cache hits and pays the 6.9 rate on more of its input.
The same allowance at Z.ai's API prices
Converting tokens into dollars needs one more input: what the same tokens would cost on Z.ai's pay-as-you-go API. GLM-5.3 lists at $1.40 per million input tokens, $0.26 cached and $4.40 output, and GLM-5.3-Flash at $0.15, $0.03 and $0.50 - Z.ai pricing. Applied to the workload that Z.ai's own table implies, a million GLM-5.3 tokens cost about 34 cents, and a million Flash tokens about 3.8 cents.
Multiply that by each tier's monthly allowance and you get the plan's value at list price, if you use all of it. The chart puts that value next to the monthly-billing price of each tier.
Read at Z.ai's own prices, the plan is a strong deal for anyone who uses it. Fully used, Lite delivers roughly 4 to 8 times its monthly price in GLM-5.3 usage, Pro 5 to 11 times and Max 6 to 12 times, and yearly billing pushes each multiple up by another 40%. Z.ai's claim of saving "up to 92%" against the standard API corresponds roughly to the best case on this chart: yearly billing, every call off-peak, every credit spent.
The more useful number is the break-even point, because nobody uses every credit. At Z.ai's API prices, Lite pays for itself once you would otherwise have bought about 53 million GLM-5.3 tokens in a month, which is about a quarter of its peak-rate allowance. Pro breaks even at about a fifth of its allowance and Max at about a sixth. If your honest estimate is that you would use less than that, the pay-as-you-go API is cheaper, and you also keep the right to use it in your own software.
The Flash asymmetry
The picture changes completely for GLM-5.3-Flash. The plan charges Flash about a third of GLM-5.3's credits per token, but Z.ai's API sells Flash at roughly a ninth of GLM-5.3's price, so the plan's discount on Flash is much thinner. Fully used on Flash alone, Lite is worth about $24 to $49 a month at API prices, only 1.3 to 2.7 times its $18 price, and breaking even takes about three quarters of its peak-rate allowance.
The practical rule that falls out of this is simple: spend plan credits on GLM-5.3. Using Flash inside the plan for quick subagent work still makes sense, because it stretches a tight 5-hour window about three times further, but a developer who mostly wants Flash is usually better off paying per token. Our GLM-5.3 Flash page lists every host's current Flash price, and many of them sit below Z.ai's own.
Against the cheapest hosts, not just Z.ai's list price
There is a second, bigger catch in the "up to 92%" headline: it compares the plan with Z.ai's own API, and Z.ai's API is not where cost-conscious developers buy GLM-5.3. Because GLM-5.3's weights are open, dozens of providers serve it, and OpenRouter listed 32 providers (41 endpoints) for it on October 5, 2026 - OpenRouter. Priced on the same cache-heavy workload, Z.ai's own endpoint ranked 31st of 32, because its $0.26 cached-input price is high next to providers charging a few cents.
| GLM-5.3 host on OpenRouter (October 5, 2026) | Input / cached / output per 1M tokens | Cost per 1M tokens of a 95%-cache agent workload | Pro's peak-rate allowance at this price, per month |
|---|---|---|---|
| Z.ai (FP8) | $1.40 / $0.26 / $4.40 | $0.34 | $424 |
| DeepInfra (FP4) | $0.5625 / $0.125 / $2.50 | $0.16 | $199 |
| Novita (FP8) | $0.42 / $0.078 / $1.32 | $0.10 | $127 |
| InferenceNet (precision not stated) | $0.12 / $0.04 / $4.40 | $0.07 | $83 |
Against the best of those hosts, the plan's advantage shrinks from "five to eleven times" to something much closer to break-even. Pro's peak-rate allowance would cost about $127 a month at Novita's FP8 prices, so Pro at $80 still wins for a heavy user, but by a factor of about 1.6 rather than 5. At the cheapest listed price the comparison is roughly even. These host prices move every week, which is why our GLM-5.3 page pulls them live rather than quoting them.
Three caveats keep the cheapest hosts from being an automatic win. Cache discounts only apply when your requests keep landing on a host that still holds your context, which a router cannot always guarantee, so real cache hit rates on a third-party host can be lower than on a single provider. Hosts that run lower-precision builds (FP4, or an undisclosed format) are cheaper to serve but can behave slightly differently from Z.ai's own deployment. And the plan bundles things a raw API key does not: the vision, search, reader and Zread MCP tools, and a fixed monthly cost that does not grow on a bad week.
What the legacy plans were worth
One more comparison explains why long-time subscribers grumbled about the July change. Z.ai's description of the legacy prompt-based plans put their monthly quota at roughly 15 to 30 times the monthly subscription fee in API terms - Z.ai. By the calculation above, the credit-based tiers deliver about 4 to 12 times their price in GLM-5.3 usage at Z.ai's current API rates. The comparison is not perfectly like-for-like (the older figure predates GLM-5.3 and its pricing), but the direction is clear, and it leads to one firm piece of advice: if you still hold a Legacy Plan V2, which can keep renewing, think hard before letting it lapse.
5. The models you actually get
Every tier includes the same two models, and the documentation is explicit that they are the only two: "All plans support GLM-5.3, GLM-5.3-Flash" - Z.ai FAQ. The subscription page's title still mentions GLM-5.2 and GLM-5-Turbo, but the routing rules decide what you actually get: requests for GLM-5.2 and GLM-5.1 are served by GLM-5.3, and requests for GLM-4.7 by GLM-5.3-Flash. There is no faster Prime variant on the plan and no way to pin an older release.
GLM-5.3 is Z.ai's current flagship, released on August 14, 2026 - DataNorth. It uses the same base model as GLM-5.2, and Z.ai says every improvement came from post-training, which made it much stronger at long-horizon coding and, unexpectedly, at security work. It has a 1M-token context, a reasoning control with three levels, and open weights under Z.ai's own GLM-5.3 License. Z.ai's launch chart shows how it compares with its predecessor and with several closed frontier models on the benchmarks it chose to highlight.
The chart is Z.ai's own and the results are vendor-reported, so read it for direction rather than as proof. The direction is still meaningful for a subscriber: on Terminal Bench 3.0, where no model in the chart scores above 35, GLM-5.3 reaches 28.3 against GLM-5.2's 4.6, and on DeepSWE it reaches 66.9 against 46.2 - Hugging Face. It leads the chart on AutomationBench and GDPval-AA, trails the closed models on Terminal Bench 3.0 and DeepSWE, and lands within about two points of them on the rest. Our GLM-5.3 page carries the full table, and our earlier comparison of Ox Alpha against the frontier explains why vendor-reported agent scores deserve a discount until someone else reproduces them.
GLM-5.3-Flash is a different and much smaller model: 320 billion parameters with 18 billion active, natively multimodal, released under the plain MIT license in August 2026. It is the model that ran anonymously as the stealth "Ox Alpha" before Z.ai claimed it, which is the story this site was built to document; our guide to what Ox Alpha was and the investigation into who built it cover that history. On the plan, Flash is the cheap workhorse: it costs a third of GLM-5.3's credits per token and makes sense for quick lookups, simple edits and subagents.
The 1M context, effort levels and the MCP tools
Three plan features change how you should configure your tools. The first is the 1M-token context: in Claude Code you enable it by adding a [1m] suffix to the model name, as in glm-5.3 [1m], and by setting Claude Code's compaction window to a million tokens - Z.ai. As section 3 showed, a long context is affordable only occasionally, so the suffix is something to switch on for whole-repository work rather than leave on by default.
The second is reasoning effort. GLM-5.3 accepts three levels, low, high and max, and defaults to max when you do not set one; Z.ai maps the effort names used by other tools onto those three (anything from "minimal" to "low" becomes low, "medium" and "high" become high, and "xhigh", "max" and "ultra" become max). In Claude Code, the /effort command switches the level in a running session. Because thinking tokens are billed as output, this is the most direct control you have over credit consumption.
The third is the set of MCP tools included in every tier:
- Vision Understanding, a local MCP server that lets agents read images and video, billed as GLM-5.3-Flash usage.
- Web Search, for current documentation and API changes, at 1.2 credits per call.
- Web Reader, which fetches a full web page as structured content, at 1.2 credits per call.
- Zread, which searches and reads public GitHub repositories (docs, structure and files), at 1.2 credits per call.
Z.ai says it offers no other way to call the vision, search and reader tools than through the plan - Z.ai FAQ. For agents that browse documentation a lot, the flat 1.2-credit price is cheap next to the model calls around it: ten searches cost 12 credits, less than half of one 120,000-token GLM-5.3 call at the peak rate. The vision server is the one to watch, because image analysis runs on GLM-5.3-Flash and its credits come out of the same allowance as everything else; the documentation asks for version 0.1.2 or later of the server to get the GLM-5.3-Flash capability - Z.ai.
6. Setting it up in Claude Code, Codex, Cline and other tools
Setting the plan up is a configuration change rather than a migration, but it has one hard requirement that trips up a lot of first-time subscribers: the quota only counts when the request reaches Z.ai through a coding endpoint from a supported tool. Point the same API key at Z.ai's general API instead, or call a model the plan does not include, and the call is billed to your pay-as-you-go balance or rejected. Z.ai's FAQ gives exactly that diagnosis for the error new subscribers most often describe, "1113 Insufficient Balance" after buying the plan - Z.ai FAQ.
The plan speaks three protocols, each with its own base URL, so the first step for any tool is knowing which protocol it uses - Z.ai. Claude Code and Goose use Anthropic's Messages format; most editors and extensions use OpenAI's Chat Completions format; and OpenAI's Codex uses the newer Responses format.
| Protocol | Base URL | Typical tools |
|---|---|---|
| Anthropic Messages | https://api.z.ai/api/anthropic | Claude Code, Claude for IDE, Goose |
| OpenAI Chat Completions | https://api.z.ai/api/coding/paas/v4 | Cline, Kilo Code, Roo Code, Crush, Cursor and most others |
| OpenAI Responses | https://api.z.ai/api/v1 | Codex |
The list of officially supported tools is longer than most people expect. Z.ai's integration page names sixteen coding tools: its own ZCode environment (marked "1.5× usage") plus Claude Code, Claude for IDE, Codex, OpenCode, Cursor, Cline, Kilo Code, Roo Code, Crush, Goose, TRAE, Qoder, Droid, Pi and Eigent. A second group of general-purpose agents, Z.ai's own AutoClaw (also marked "1.5× usage"), OpenClaw, Hermes Agent and SillyTavern, is supported "on a best-effort basis" and may be rate-limited under high load. Anything outside those lists is outside the terms, which section 8 covers.
A good overview of what the switch looks like in practice, and of what carries across when you change the model under an existing harness, comes from Nate B. Jones's walkthrough of running GLM-5.3 inside Claude Code and Codex, published a week after GLM-5.3 launched.
The parts most relevant to this guide are the Claude Code launcher from 07:09, the Codex setup from 12:09, and the section on plan limits and whether the switch is actually cheaper from 16:13. His segment at 06:19, "What 96% reused input taught me", is a useful real-world check on the 95 to 98% cache hit rates Z.ai assumes in its allowance tables, and on why cached-input pricing dominates the plan's economics.
Claude Code
Claude Code, Anthropic's terminal coding agent - Claude Code docs, is the most common pairing, and Z.ai documents it in the most detail. You install Claude Code itself with npm (npm install -g @anthropic-ai/claude-code, which needs Node.js 18 or newer, per Z.ai's Claude Code guide), create an API key in the Z.ai console under the Coding Plan's overview page, and then tell Claude Code to send its requests to Z.ai instead of Anthropic. The cleanest way to do that is the env block of ~/.claude/settings.json, which keeps the change out of your shell profile and applies to every project.
Z.ai's documented manual configuration looks like this, with your own key in place of the placeholder. It maps all three of Claude Code's model slots to GLM, enables the 1M context with the [1m] suffix, and raises the timeout for long agent turns.
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "your_zai_api_key",
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash [1m]",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3 [1m]",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3 [1m]",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": 1,
"API_TIMEOUT_MS": "3000000"
}
}
Two of those lines deserve a second thought once you have read section 3. The [1m] suffixes and the million-token compaction window let a session grow to a million tokens before Claude Code summarises it, which is great for whole-repository work and expensive in credits; leaving the suffix off for everyday sessions keeps each call smaller. Mapping the Haiku slot to GLM-5.3-Flash is the right call either way, because Claude Code uses that slot for its background functionality - Claude Code docs, and Flash costs a third of the credits.
After restarting Claude Code, run /status to confirm the switch took. Z.ai's own screenshot shows what a working setup looks like: the base URL points at Z.ai's Anthropic endpoint, and the model line names GLM-5.3 with the 1M suffix.
If the status screen shows the right model and requests still fail with error 1113, the usual cause is a key mix-up rather than a setup error. Team plan members must use the separate Team Plan Key from the Team section of the console, which is not interchangeable with an individual key, and any key only draws on the plan when its calls go to a coding endpoint rather than Z.ai's general API. Z.ai also ships an automated alternative, the Coding Tool Helper, which you can run with npx @z_ai/coding-helper to detect, install and configure Claude Code, Codex, OpenCode, Crush and Factory Droid in one guided pass - Z.ai.
Codex and the OpenAI-compatible tools
Codex, OpenAI's coding agent, is the one tool that needs the Responses endpoint, https://api.z.ai/api/v1, and it needs a little more setup than the others - Z.ai. Z.ai's guide has you declare the GLM models in a ~/.codex/models.json catalog (including the three supported reasoning levels, with max as the default) and then point Codex's provider configuration at Z.ai. It is fiddlier than Claude Code, which is one reason the Coding Tool Helper exists, but once done it behaves like any other Codex model.
Most other tools on the list use the OpenAI Chat Completions protocol with the same three values: the base URL https://api.z.ai/api/coding/paas/v4, your Z.ai key, and a model id. In Cline, an AI coding extension for VS Code - Cline, you choose "OpenAI Compatible" as the provider and fill those three fields in, as Z.ai's own example shows.
Two details in the OpenAI-compatible tools catch people out. First, use glm-5.3 or glm-5.3-flash as the model id: older ids still work but are silently served by those two models, so naming them only adds confusion when you read your usage. Second, Z.ai's Cline guide has you set the context window size by hand, so set it to what you are willing to pay for per call rather than the largest window the model allows. Z.ai's guides for Kilo Code, Roo Code, Crush and Cursor use the same base URL, and OpenCode needs none at all: running opencode auth login and choosing the built-in Z.AI Coding Plan provider does the job - Z.ai.
7. How to make the credits last
When the plan seems to run out too fast, the credit formula points to three habits that cost nothing to change: letting the context grow without limit, running every task at maximum reasoning effort, and doing heavy work in peak hours. The credit formula makes each of them measurable. Context size drives the cost of every call, output tokens cost fourteen times as much as cached input, and the peak window doubles the price of everything inside it.
The order of priority follows from the numbers in section 3. A developer who keeps sessions around 100,000 tokens instead of letting them drift to 400,000 cuts the cost per call by more than two thirds; switching routine work to a lower effort level shrinks the most expensive token category; and moving long autonomous runs out of the peak window halves their cost outright. None of it requires a bigger tier.
Keep the context short and the cache warm
The cheapest call is one that re-sends little, and the second cheapest is one whose context the provider has already cached. In Claude Code, /clear starts a new conversation with an empty context and /compact summarises the conversation so far to free space, while /context shows how full the window is - Claude Code docs. Clearing between unrelated tasks is the single most effective habit on the plan, because a fresh session costs a fraction of a long one for the same amount of work.
Caching rewards a stable beginning. Z.ai's caching is automatic: repeated system prompts and conversation history are recognised and billed at the cached rate without any configuration - Z.ai. That works because the start of each request matches the previous one, so anything that changes the beginning of the context mid-session (editing your project instruction file, adding or removing MCP servers, switching models) can force the next call to pay full input price on everything after the change, because prompt caches generally match a request from its start. Make those changes between sessions, not during them.
Turn the reasoning dial down for routine work
GLM-5.3 thinks at maximum effort unless you tell it otherwise, and every thinking token is billed at the output multiplier of 24. Most coding steps do not need that: renaming a variable, writing a test from a clear specification or explaining a function benefits little from a long internal monologue. In Claude Code, /effort switches the level in a running session; Z.ai maps Claude Code's levels onto GLM's three, and recommends staying at max for complex coding tasks - Z.ai. A sensible pattern is to plan and debug hard problems at max, and to drop to high or low for the mechanical work in between.
Deciding which tasks deserve the full dial is a skill in its own right, and it applies to every reasoning model rather than to GLM alone. founden.ai, which is also run by the person who publishes Ox Alpha, has a task-to-effort framework for reasoning models that shows where extra thinking pays for itself and where it is pure waste. On the GLM Coding Plan the payoff is unusually direct, because the credits you do not spend on unnecessary thinking are credits left in the 5-hour window.
Use Flash for the small jobs and schedule the big ones
Routing work between the two models is the next lever. GLM-5.3-Flash costs a third of GLM-5.3's credits per token, so sending lookups, file searches, simple edits and subagent tasks to Flash stretches the window, while the planning and the hard debugging stay on GLM-5.3. In Claude Code the Haiku slot already does some of this automatically if it is mapped to Flash; in other tools, a second model profile for quick tasks achieves the same.
Scheduling is the last lever and the easiest for anyone outside East Asia. Because off-peak usage costs half, a European developer who starts long autonomous runs after midday local time, or leaves a large refactor to run in the evening, halves the cost of that run. Concurrency follows the tier: Z.ai recommends Lite for one project at a time, Pro for one or two in parallel and Max for two or more, and raises concurrency limits for all plans during off-peak hours - Z.ai usage policy. If you hit rate limits running parallel subagents on Lite, that is the plan working as designed rather than a fault.
8. The rules that can cost you the subscription
The plan's terms are stricter than most developer subscriptions, and Z.ai enforces them with automated "risk control", so they are worth reading before you build a workflow around the plan. The core restriction is about where the quota can be used: only within officially supported tools, and never for "general-purpose API access", including calling models "from your own applications, bots, websites, SaaS products or other systems" - Z.ai subscription terms. The terms also name "SDK-based access" as a pattern the system detects, which rules out scripting the coding endpoint from your own code even for personal automation.
The second restriction is about who can use it. A subscription is "licensed only to the individual natural person associated with such account", and sharing it with "colleagues, friends, customers or any organization", by any means, is forbidden; reselling, proxying or offering the plan's model access as a service is forbidden too. If Z.ai suspects sharing or resale, the terms allow it to restrict features, reduce the quota, suspend the service and reclaim any remaining quota. A team that wants shared access needs Team seats, which section 11 covers.
Enforcement is graduated but real. Z.ai's usage policy says violations "may trigger risk control measures, including rate limiting, account freezing, or other restrictions", and that "accounts with more than three violations may be banned"; a flagged account sees a notice on its plan overview page, where an appeal can be filed - Z.ai usage policy. The rules most likely to catch honest users are mundane:
- Running an unlisted tool because it speaks the OpenAI protocol and "seems to work".
- Wiring the key into a script or internal bot, which counts as SDK access.
- Lending the key to a colleague for a day, which counts as sharing.
None of those feel like abuse from the inside, which is exactly why they cause trouble. A fourth habit is not a violation but produces errors that look like one: running many parallel agents on Lite, which Z.ai sizes for one project at a time, runs into concurrency limits rather than a breach of the rules. The first two are the ones to watch most closely, because they look like ordinary engineering: a developer who builds a small internal tool on top of their coding plan key has, in the terms' language, used the quota for general-purpose API access. The safe pattern is to keep the plan key inside the supported tools and to use a separate pay-as-you-go key, or a third-party host, for anything you write yourself.
Refunds, cancellation and renewal prices
Money rules are equally firm. Subscriptions are non-refundable once purchased, "even if you have not used up your plan" - Z.ai usage policy. They renew automatically, and the documentation is inconsistent about how early you must cancel: the FAQ and the subscription terms say at least 24 hours before the next billing date, while the usage policy says at least 3 days. Cancel three days ahead and both versions are satisfied. A cancelled plan keeps working until the end of the period you paid for.
Two quieter clauses matter for anyone planning a long subscription. Automatic renewal charges "the price, applicable rules, and promotional arrangements displayed on the page" on the renewal date, not the price you first paid, so a price rise reaches existing subscribers at their next renewal - Z.ai subscription terms. And Z.ai caps its total direct liability to you at what you spent in the most recent calendar month. Neither is unusual for a cloud service, but together they argue for shorter billing cycles until you are sure the plan fits.
Changing plans follows its own logic. Moving to a longer term on the same tier (Lite monthly to Lite annual, for example) does not take effect immediately: the terms stack, so that switch would give you 13 months in total. Moving up a tier takes effect at once, and the unused value of the old plan is converted into account balance and credited against the new price - Z.ai FAQ. Changing only the billing cycle waits until the current cycle ends.
9. Privacy, data and jurisdiction
Every request you send through the plan carries your code, so where it goes and what happens to it is a fair question to settle before subscribing. The international Z.ai service is run by JINGSHENG HENGXING TECHNOLOGY PTE. LTD, a Singapore company, and its privacy policy states that it generally provides the services from Singapore, so your data is "generally processed in Singapore" - Z.ai privacy policy. That entity is distinct from Beijing Zhipu Huazhang, the mainland company behind the GLM models.
On training, the documents are careful rather than categorical. The privacy policy describes prompts and other "User Content" as information "processed in real-time to provide you with the Service", and the purpose it lists for training and improving models covers account, communication, usage, log and device data rather than naming user content; the data processing addendum for business API customers goes further and states that the company does not store the content customers input. For subscriptions, the Team plan states plainly that code, prompts and conversations are "excluded from model training by default" - Z.ai. The individual Lite card promises "default data privacy" without defining it.
The jurisdiction question sits one level up. Beijing Zhipu Huazhang Technology, Zhipu AI's mainland entity, was added to the US Commerce Department's Entity List in January 2025 - Federal Register, and the company listed on the Hong Kong Stock Exchange in January 2026 - CNBC. The listing restricts certain US exports to that company; whether it affects your use of a subscription sold by Z.ai's Singapore entity is a question for your compliance team, and our profile of Z.ai sets out the corporate background in more detail.
In practice, three precautions cover most of the risk:
- Keep secrets out of context: exclude
.envfiles, keys and credentials from what your agent reads. - Match the data to the terms: use the plan for code you would be comfortable sending to any third-party API.
- Self-host when it matters: for regulated work, GLM's open weights let you run the model on your own hardware.
The third option is more realistic than it sounds for some teams. GLM-5.2 is published under the plain MIT license and GLM-5.3 under Z.ai's own license, so an organisation that cannot send code to an outside provider can run the same family on its own servers, at the cost of serious hardware; our GLM-5.2 page lists what that takes. The plan and self-hosting are not mutually exclusive either: plenty of teams use the plan for open-source and personal work and keep a self-hosted or tightly governed setup for sensitive repositories.
The first precaution is easy to make concrete, because an agent reads whatever its permissions allow, and a .gitignore entry does not stop it. In Claude Code, a deny rule in ~/.claude/settings.json blocks reads of environment files for every project, and the documentation uses exactly this example - Claude Code docs:
{
"permissions": {
"deny": [
"Read(./.env)",
"Read(./.env.*)"
]
}
}
The rule sits in the same file as the Z.ai variables from section 6, so one edit configures both the provider and the guardrail. It is worth adding before the first session rather than after, because the first thing many agents do on a new repository is read the project files to orient themselves, and a secret that has been sent once cannot be recalled. Other tools have their own ignore or permission settings; whichever you use, check them before pointing the agent at a repository that holds credentials.
10. The plan against the alternatives
The plan is one of four ways to put a GLM model behind a coding agent, and one of several ways to pay a flat price for an AI coding agent in general. Choosing well means comparing it on the dimension that actually differs: how your usage pattern maps onto each pricing model. The table summarises the four GLM routes with the figures from this guide.
| Route | What you pay | Limits | Use inside your own apps? | Best for |
|---|---|---|---|---|
| GLM Coding Plan | $18 to $168 a month, less on longer terms | 5-hour and weekly credits | No | Steady daily use in supported tools |
| Z.ai pay-as-you-go API | GLM-5.3 at $1.40 / $0.26 / $4.40 per 1M tokens | Rate limits only | Yes | Light or irregular use, app features |
| Third-party hosts | Often a third of Z.ai's price or less on cache-heavy work | Per host | Yes | Cost-sensitive, flexible use |
| Self-hosting the weights | Your GPU bill (756 GB of FP8 weights for GLM-5.3) | Your hardware | Yes | Data control, very high volume |
The comparison that matters most is the plan against pay-as-you-go, and section 4 gives the threshold: below roughly a quarter of Lite's peak-rate allowance, or a fifth of Pro's, the API is cheaper even at Z.ai's own prices, and against the cheapest third-party hosts the threshold is much higher. Our GLM API pricing table shows every GLM model's Z.ai price next to its cheapest host, refreshed every few hours, which is the quickest way to price your own usage at today's rates. If you used this site's free GLM-5.3 Flash endpoint before it closed, our migration guide has the base URLs and model ids for each provider.
Self-hosting deserves a sober note. GLM-5.3's published FP8 checkpoint alone is 756 GB, and Z.ai's own example deployment for this architecture shards the model across 8 GPUs - Hugging Face. That is a server-class commitment that only pays off at very high volume or when data control is non-negotiable. GLM-5.3-Flash, at 320 billion parameters, is far easier to host, and our GLM-5.3 Flash page lists its sizes and serving options.
GLM Coding Plan or a Claude subscription?
The comparison most people actually make is with Anthropic's plans, because the GLM plan is so often used inside Claude Code. Claude Pro costs $20 a month billed monthly or $17 a month billed annually, and Claude Max starts at $100 a month, with options for five or twenty times Pro's usage; all paid plans include Claude Code - Claude pricing. What those plans buy is access to Anthropic's own models rather than GLM. In Z.ai's own GLM-5.3 comparison, the Anthropic model listed as Fable 5 scores above GLM-5.3 on five of the seven coding benchmarks both report - Hugging Face.
So the honest framing is not "GLM instead of Claude" but "which work goes where". The GLM plan is far cheaper per unit of agent work, and Z.ai pitches it explicitly as a way to run more volume for less; Anthropic's plans buy the strongest models Z.ai itself benchmarks against. Many developers run both: Anthropic's models for planning and the hardest debugging, and GLM for the high-volume routine work, switching with a separate Claude Code settings profile or shell alias that sets Z.ai's base URL. Our model-by-model comparison table makes the same point about GLM-5.3 Flash: the value is in matching model to task, not in crowning one model.
The decision tree below condenses the choice.
The tree deliberately puts the usage question before the price question, because the plan's value collapses when usage is light and its rules exclude whole categories of use outright. Once you are in the "heavy use most days" branch, the plan is very likely the cheapest route, and the remaining choice is between tiers, which section 13 covers.
11. The Team plan
The Team plan is a self-service subscription bought by the seat, and it is more than individual plans with a shared invoice - Z.ai. An administrator buys seats, assigns them to members, and each member gets their own Team Plan Key with their own allowance. Two seat types exist: a Standard Seat with 15,000 credits per 5 hours and 66,000 per week, and a Premium Seat with 35,000 and 155,000. The credit multipliers, the off-peak discount and the supported tools are the same as for individual plans.
Priced per seat, the Team plan costs a little more than the matching individual tier: $88 a month for a Standard Seat against $80 for an individual Pro plan, and $188 for a Premium Seat against $168 for Max, with a minimum of two seats - Z.ai subscription page. For that premium, each Standard seat gets 10% more weekly credits than Pro, and the team gets features individual plans do not have. Administrators manage seats, roles and permissions in one place and see usage dashboards by member and period. Code, prompts and conversations are excluded from training by default. Billing and invoicing are consolidated for the whole organisation, and verified companies can request VAT invoices.
The overage option is the feature that changes the economics for a business. Individual plans simply stop when a window is empty, which is tolerable for a solo developer and expensive for a company with a deadline. When an administrator enables on-demand overage, a seat that exhausts its allowance keeps working at pay-as-you-go rates, with per-member spending limits to keep surprises bounded, and for a limited time that overage is billed at 10% below the API list price. The team gets the plan's discount for most usage and pays close to list price only for the peaks. Premium seats add first access to new flagship models and priority resources during peak hours, which matters most for teams based in Europe or Asia.
A few seat rules are worth knowing before the first purchase. Standard and Premium seats cannot be mixed on one subscription, a Standard subscription cannot be upgraded to Premium, and seats can be added mid-cycle at a pro-rated price but not removed until the cycle ends. The administrator who buys the plan does not occupy a seat unless they assign one to themselves, and a member can hold an individual plan and a Team seat at the same time.
12. How the plan changed in 2026, and what comes next
The GLM Coding Plan has been rebuilt twice this year, and the direction of travel says something about where it is going. Until spring 2026 some subscribers still held legacy plans with no weekly limit at all. Z.ai's April 21 notice phased those out: their auto-renewal was cancelled on April 30, affected users received two free months of the equivalent current tier, and a 50% migration discount valid until three months after the end of their legacy period - Z.ai migration notice. The current plans at that point were priced at $18, $72 and $160 a month and measured usage in prompts.
The second rebuild arrived on July 30, 2026, when Z.ai moved new subscriptions to the credit system described in this guide - Z.ai. Under the previous scheme (now called Legacy Plan V2), Lite allowed up to about 80 prompts per 5 hours and 400 per week, Pro about 400 and 2,000, and Max about 1,600 and 8,000, with GLM-5.3 consuming quota at 1× off-peak and 3× at peak. Existing subscribers kept their plans until the end of their billing cycle; V2 holders can keep renewing, while V1 holders move to the new plans when their cycle ends.
Two details of that change are easy to miss. The peak penalty became milder: under V2, GLM-5.3 cost three times as much at peak as off-peak, while credits cost only twice as much at peak (the full rate against the 50% rate). And measurement became honest about context: a prompt that drags a 400,000-token context through twenty model calls used to count as one prompt, and now costs what its tokens cost. That is better for light, focused sessions and worse for long, sprawling ones, which is consistent with the much lower dollar value per subscription that section 4 calculated.
What to expect next
The first-principles reading of these changes is that Z.ai is moving the plan from a marketing loss leader towards a priced product. Prompt-based quotas are generous to heavy users in ways the provider cannot control, because the cost of a "prompt" varies by a factor of a hundred depending on context size. Token-metered credits fix that, and per-model multipliers let Z.ai price each new model as it ships without changing the sticker prices. The promotions Z.ai ran this autumn (all-day off-peak rates, free overnight Flash in Z.ai's own apps) look like the other half of the same strategy: generous temporary capacity to win habits, on top of a meter that is now sustainable.
For a subscriber that suggests three expectations. Allowances and multipliers will keep being re-tuned as new GLM models arrive, as they were for GLM-5.3. The plan's own apps, ZCode and AutoClaw, will keep getting the best deals, as the 1.5× usage markers and the overnight campaign already show. And renewals will follow whatever the page says on the day, so locking in a yearly plan buys a 30% discount at the cost of flexibility you cannot get back through a refund. The switching cost the other way is low: because GLM is open-weight and the endpoints speak the Anthropic and OpenAI protocols, leaving the plan for an API host is a configuration change.
13. The verdict: which tier, if any, to buy
The GLM Coding Plan is a good product with a narrow sweet spot, and the right decision depends more on how you work than on how much you want to spend. Reduced to its mechanics, the plan sells GLM-5.3 at roughly a quarter to a twelfth of Z.ai's own API price to people who use it steadily, inside approved tools, mostly outside Asian afternoon hours. Everything outside that description erodes the deal, and the alternatives section showed how fast.
Three variables decide which side of that line you are on: how many hours a week your agent actually runs, how large its context grows, and whether those hours fall inside Z.ai's peak window. Price matters less than any of them, because the cheapest tier is the expensive choice if it runs out every afternoon, and the most expensive tier is waste if half its credits expire unused each week. The table turns those variables into a recommendation.
| Your situation | What to buy | Why |
|---|---|---|
| A few tasks a week | Pay per token, ideally a cheap host | Usage stays below the plan's break-even |
| Daily, in the Americas | Lite to start, Pro when the 5-hour wall bites | Almost all your hours are off-peak |
| Daily, in Europe or India | Pro, with long runs moved out of peak | Peak hours overlap your working day |
| All-day, multi-project work | Max | Credits a third cheaper than Lite, more concurrency |
| Companies | Team seats with overage enabled | Plans cannot be shared, and overage avoids hard stops |
Whichever tier you choose, start on monthly billing for the first cycle and measure. The console shows credits consumed, and a week of real work at standard rates (not during a promotion) tells you whether you are a Lite, Pro or Max user far better than any estimate. Keep a pay-as-you-go key or a third-party host for anything the plan's terms exclude, keep secrets out of the agent's context, and remember that the cheapest credit is the one a cleared session, a lower effort setting or an off-peak hour lets you keep. If the plan is your first contact with GLM, our earlier guides on using the model in practice and reading its benchmarks honestly are a good second read.
This guide reflects the GLM Coding Plan as of October 5, 2026, checked against Z.ai's documentation, subscription page and terms on that date; the promotions in section 3 were updated on October 7, when both ended. Prices, credit allowances, multipliers, promotions and supported tools change often, and third-party host prices change weekly, so verify the current figures on Z.ai's pages and our live GLM pricing table before you subscribe.