The short version
UVU's instinct is right: at $25 per seat per month, covering 48,669 students runs $14.6 million a year — $16.4 million with all 6,123 employees added (54,792 × $300) est — four to five percent of the university's entire Education & General budget. Nobody should sign that.
But the real choice is not "expensive cloud or nothing." Three facts, all verified this week, change the picture:
- Tools that add no license cost already exist. UVU's own guidance pages list Copilot Chat as included for students and staff under the current Microsoft license, and Google's Gemini for Education base tier costs education institutions nothing. Whether Gemini is actually switched on for UVU's account is something UVU's Digital Transformation office must confirm. fact license inclusion · tenant scope
- The biggest signed education deals run ~$19–30 per user per year. Cal State pays ~$19.20/user·yr across 675,000 people; Colorado ~$20; Maine ~$23 per billed FTE. Against those, the $25/month fear is off by 10–16× — though mid-size schools have been quoted far more, which is why week one requests UVU's own numbers. fact
- Owning the floor is now cheap. The best free-to-download models score six points behind the best paid ones on the 0–100 capability score explained above, run on a $2,139 Mac mini, and improve at no new license cost — each upgrade still earns a qualification pass. Faculty/staff serving hardware is roughly $64k upfront (~$80k over three years with power and care); operating staff and adoption support are the recurring costs, and they're priced in the open throughout this guide. fact est
$0
The no-added-license-cost starting option — already theirs
Copilot Chat appears to be included in UVU's Microsoft license, and Gemini for Education's base tier costs education institutions nothing; UVU must confirm whether Gemini is switched on for its account. Together they can put a competent assistant in front of most people at no added license cost, whether or not this plan proceeds. This is the no-added-license-cost starting option: it removes the reason to wait, and nothing built later depends on it. fact the license terms · unknown whether Gemini is enabled
OWN
Basic on-campus service — a Mac fleet UVU owns
Mac minis and Studios running free-to-download models, behind UVU's existing campus login. Today's pick is a model with 27 billion learned settings — written “27B” — about the largest that fits comfortably on one Mac mini. Work on this route stays on campus; if the campus service is unavailable, a request for sensitive work stops rather than being sent outside — the written rules that decide which data must stay on campus are in §08. There is no charge per person. Model upgrades are free downloads, and each one must pass testing before students see it. It is built as a self-contained service — its own front door, its own operations, its own budget — so it stands whether or not any existing UVU system is available to it. If Digital Transformation wants it, the fleet can also connect to UVU's existing artificial-intelligence entry point: a connection, not a dependency. est
RENT
The pay-as-used research pool — for work the campus machines cannot do
Intense mathematics, physics and deep engineering get access to the strongest paid models (Claude, GPT‑5.6, Gemini) through a fund with a spending limit on it: roughly $22,000 a year covers 200 heavy researchers, allowing each about 20 million pieces of text. Programs that cost nothing — Anthropic scientist seats, research credit grants of up to $50,000, and the US National AI Research Resource — are claimed first, and this line is only what is left to buy. est
The one deadline that matters
Apple replaced its entire Mac line on August 25. The new machines ship September 22; nobody has production benchmarks yet. The move: order four minis as a technical validation — not a service launch — to arrive with the first wave. Two weeks of hard measurement — how fast it replies, how many people it serves at once, whether one person's work can reach another's, and whether it stays stable over days — while quotes and shared-resource applications run in parallel, and every later step passes its gate before people touch it. fact dates · est plan
The words this plan uses, in plain English
Every specialist term on this site is defined here, and each is glossed again where it first appears. If a word on any page is not in this list and not explained beside it, that is a fault in the writing — tell us.
- Conversation at once (elsewhere “stream”) — one person waiting for one answer. Capacity is counted in these, never in people: 36,629 people do not all type at the same moment. The capacity rule.
- Model — the program that writes the answers. Free-to-download models can be run on your own machines; paid-only models exist only as a rented service.
- Model score — the Artificial Analysis Intelligence Index: an independent lab runs every major model through the same reasoning, mathematics and coding tests and combines the results into one 0–100 number. Higher is better.
- Parameters, or learned settings — how big a model is. “27B” means 27 billion of them. Bigger usually answers better and needs more memory.
- The strongest paid models (elsewhere “the frontier”) — the best commercial services of the moment: Claude, GPT‑5.6, Gemini.
- Router — the piece of software that decides which model should answer a given request, and whether that request is allowed to leave campus.
- Entry point, or gateway — UVU's existing front door for artificial-intelligence services, run by Digital Transformation. This plan can connect to it and does not need it.
- FTE, full-time equivalent — one person's full working year. “1.25 FTE” means one and a quarter full-time positions, however many people that is.
- MFA, multi-factor authentication — signing in with a password and a second proof, such as a phone prompt.
- Privacy-separation tests (elsewhere the “isolation canary battery”) — planting a unique secret in one person's session and then trying, from another, to reach it.
- Approved outside destinations (elsewhere “egress allowlist”) — the short, written list of places the service is permitted to send anything at all.
- Dx — UVU's Digital Transformation office. SCET — the Scott M. Smith College of Engineering and Technology. HERFP — Utah's Higher Education Research Fund Program. IRB — the university committee that approves research involving people. CISO — the university's chief information security officer. RACI — a chart naming, for each task, who does it, who is accountable, who is consulted and who is told.
- FERPA — the US federal law protecting student records. ADA — the Americans with Disabilities Act. WCAG — the international standard for accessible web content. USHE — the Utah System of Higher Education.
Why this fall
Downloaded models have become useful enough for daily campus work. Two years ago, running a useful model on a desktop was a hobby. Today the best free-to-download models — Kimi K3 and GLM‑5.3 — score 60 on the Artificial Analysis index against 66 for the best paid-only model, and the lag behind the strongest paid models has fallen from about twelve months to three or four. The strongest paid models still do some work better. fact The chart below is the whole hardware argument: the score of the best model that fits in a 64GB Mac mini, measured on the same index.
The hardware just reset. Apple announced new Mac minis (M6 / M5 Pro) and Mac Studios (M5 Max / M5 Ultra) on August 25, first arrivals September 22, with a 512GB flagship following in late October. Buying last-generation hardware right now would be the one unforced error — the previous models are no longer sold new, and the survey in §04 and Appendix O found only single refurbished units and reseller backorders; timing the validation to the first shipment wave costs nothing extra. fact
Utah is funding exactly this. The 2026 Legislature put $15M one-time into a shared higher-ed AI research data center (administered under the University of Utah's budget item), created a research fund that reserves 15–25% for institutions like UVU, and UVU's own research office lists a $5M "Utah AI Moonshot" and a $45M Higher Education Research Fund with full proposals due October 6. A new president took office August 10. The October 6 round is available only if UVU filed the required notice of intent by August 28; if it did not, the next round is the target. fact the dates · unknown whether the notice was filed
How to read the scores in this guide
Whenever this guide gives a model a number like 52, 60, or 66, it is the Artificial Analysis Intelligence Index: an independent lab runs every major model through the same set of reasoning, math, and coding tests and combines the results into one 0–100 score, updated as models are released. "Open" means the model is free to download and run on your own machine; "closed" means it exists only as a paid service you rent. A higher score is better; a six-point gap is small — roughly the difference between this year's and last year's best. fact source · est "small" is our reading
What cloud AI really costs
| Option | Per user / year | All-campus (54,792) | Evidence |
|---|---|---|---|
| The fear: list-price seats at $25/mo | $300 | $16.4M | hypothetical est |
| Microsoft 365 Copilot (academic list) | $216 | $11.8M | Microsoft fact |
| Google AI Pro for Education (list) | $240 | $13.2M | Google fact |
| Cal State × OpenAI — renewal | ~$19.20 | $1.05M | CSU's own $1.60/user·mo figure fact |
| U. Colorado × OpenAI — signed | ~$20 | ~$1.1M | $2M ÷ 100k fact |
| U. Maine System × OpenAI — signed | $22.64/billed FTE | ~$1.24M | $1.4M / 2yr ÷ 30,912 FTE pricing basis fact |
| Mid-size reality check: UC Davis / U. Iowa internal rates | $144–156 | $7.9–8.5M | published campus rates fact |
| Copilot Chat (UVU's A-license) & Gemini for Education | $0 | $0 | included / free tier fact |
| Hosted open models via API (gpt‑oss‑120b class) | ~$0.20–0.75 | ~$10–40k | live price tables fact usage est |
So cloud is not unaffordable. The honest problem is different:
Owning trades that recurring bill for operating work: hardware is mostly one-time money (the kind universities actually have), capacity grows as free model upgrades land, conversations stay on campus under UVU's own policies, and the machines double as teaching objects for the AI programs UVU already runs — but somebody has to run them, and that staffing cost appears beside every hardware number in this guide. Rent where elasticity matters. Own where usage is daily and data is sensitive. Decide with UVU's own quotes and telemetry, not anyone's brochure — including this one.
The local floor: what to buy and what it does
The serving unit is unglamorous and right-sized: the new Mac mini M5 Pro with 48 GB of shared memory, a 512 GB disk and 10-gigabit networking — $2,399 retail, $2,229 education — likely purchasable on Utah's existing cooperative contract, with Procurement confirming the route (§09). The plan prices the 10-gigabit machine because its own rack kit assumes 10-gigabit switching; the same mini without it is $2,299 retail / $2,139 education and ships three weeks sooner. Published tests on this exact chip class measured the fast-lane model family reading prompts at 2,100–3,700 tokens per second (a token is roughly a word piece) and writing replies at 97–105. The higher-quality 27B model runs slower per word, so the router sends each job to the lane built for it (§11). Every number here gets re-proven on the delivered minis before any purchase beyond the four-unit validation. fact prices & published measurements · est service planning
Measured by Utlyze on its own machine — not quoted from someone's blog
Utlyze ran the exact software and model this plan proposes, on a different and larger machine than the one the plan would buy: our own 512 GB-class Mac Studio, deliberately while it carried other work, so these are conservative floors. (“t/s” is tokens a second, and a token is roughly a word piece, so 27 t/s is comfortably faster than reading speed.) One conversation: 27 t/s. Four at once: 4 × 10.7 t/s. At eight, the queue behaved as designed — everyone still got 10.7 t/s, in two waves. A 4,563-token document was read in at 389 t/s. The privacy-separation tests — a unique planted secret per user, probes that try to reach another user's, and follow-up questioning — found zero leaks. These tests run before every future update. The four-computer test must confirm all of this on the planned Mac minis, which are smaller machines. fact — measured 2026‑09‑02 by Utlyze; method in the appendix pack
| Machine | Memory / storage | Price (education) | Model it serves | Conversations at once, mixed work | Role |
|---|---|---|---|---|---|
| Mac mini M5 Pro | 48 GB / 512 GB, 10-gigabit | $2,229 | Qwen3.8‑27B class (score 52, handles pictures, Apache‑2.0) | 4 | The workhorse. Buy many. |
| Mac Studio M5 Max | 128 GB / 512 GB | $4,639 | 120B-class (Mistral Small 4 / gpt‑oss‑120b) | 4 | Heavyweight node for Engineering & Science |
| Mac Studio M5 Ultra | 256 GB / 1 TB | $9,869 | GLM‑5.3‑Flash (320B, 18B active; score 57; MIT) — handles pictures and video | 4 | Later, if demand earns it. |
| Model mirror | one node moved from 512 GB to 2 TB | +$720 | holds every approved model on disk, 715–721 GB | — | So a rebuilt machine copies over the campus network instead of pulling 720 GB from the internet again |
| Mac Studio M5 Ultra — research node | 512 GB / 1 TB | price unknown | GLM‑5.3 (score 60) — the strongest model that runs with no cloud at all | 2, in a research lane of its own | Conditional. Apple showed it on September 3 as “coming late October”, with no price. |
Two definitions this plan uses on every page
1. What a computer can carry — the capacity rule. One machine does not have one capacity; it has a different one for each kind of work. This plan therefore publishes one planning figure and the table behind it, and never mixes them:
| Kind of work | Mac mini | 128 GB Studio | Why |
|---|---|---|---|
| Plain chat and tutoring | 8 | 8 | Limited by how fast the machine writes the answer |
| Coding help | 4 | 8 | A whole repository is far more to read than a chat message |
| Questions about a document it has not read | 2 | 4 | Limited by how fast it can read the file; 4 per mini once indexed |
| A coding agent working on its own | 1 | 2 | One agent occupies a machine by itself |
| Research writing | 2 | 4 | Uses a model tested to hold the stated answer quality |
| Mixed-work planning limit | 4 | 4 | One average across the mixture above. This, and only this, is what the budget model buys against, and what every “conversations at once” figure on this site means unless it says otherwise. |
2. What a fleet is made of — the fleet definition. A machine count on this site always breaks down into five named parts, and never into anything else:
- Active minis — Mac minis assigned to a college and carrying conversations.
- Studios — 128 GB Mac Studios, one per heavy college, carrying the larger everyday model.
- Ultra — at most one 256 GB Mac Studio Ultra, shared campus-wide through the router.
- Spares — about one machine per seven, held unconfigured and switched off until another fails.
- Shared equipment — the rack kit (one per six machines: shelf, 10-gigabit switching, battery backup) and one node's storage upgraded to hold the model mirror. These are money, not extra computers.
So “25 computers” at the Phase one budget of $120,000 means 19 active minis + 2 Studios + 1 Ultra + 3 spare minis, with the rack kit and the model mirror priced beside them. A DGX Spark, where one is bought, is a teaching computer and is counted in the cost and never in the capacity. est capacity figures; fact the prices behind the equipment. Detail: Appendix R.
The newest lineup beside the generation it replaced
A fair question from the first read-through: the plan is priced on Apple's newest machines, so what did the generation before them cost, what does it cost today, and how much faster should the new one really be? Here is the answer, priced both ways, with the grades. The short version: the previous machines are no longer sold new; single refurbished units exist at prices near or above what they cost new; and the new ones should carry about 10% more conversations per box (40% more for the Ultra) — credits we set well below Apple's headline claims. fact prices, availability, published measurements · est speed credits and like-for-like deltas
| Role | Previous generation | What it cost then · what it costs today | Newest lineup (announced Aug 25, 2026) | Price now (retail · education) | Price change, like for like | Expected speed | Our call |
|---|---|---|---|---|---|---|---|
| Workhorse mini, 48GB | Mac mini M4 Pro · memory speed 273 GB/s fact | Family launched Oct 2024 at $1,399 ($1,299 education) fact; this configuration's launch price unknown. Today: not sold new; Apple Refurbished $1,949 with 10GbE, listed one unit at a time fact | Mac mini M5 Pro · 307 GB/s fact; ships Sept 22 | $2,299 · $2,139 fact; with 10GbE like the refurbished unit, $2,399 · $2,229 est | New (education, 10GbE) vs refurbished: +$280, about +14% est | +10% conversations per box est — memory speed +12%; same-build tests +11% reading, +21% writing fact | Buy new. Supported supply, warranty, contract price, and a modest speed gain. Old units only against a written quote for the exact quantity. |
| Heavyweight Studio, 128GB | Mac Studio M4 Max · 546 GB/s fact | Launched Mar 2025 at $3,699 fact. Today: Apple Refurbished $4,159, single units fact — more than it cost new | Mac Studio M5 Max · 614 GB/s fact; ships Sept 22 | $5,099 · ~$4,639 est (512GB drive; $5,399 at 1TB fact) | New (education) vs refurbished: +$480, about +12% est | +10% per box est — memory speed +12%; same-build tests +12% reading, +25% writing fact | Buy new for planned nodes; refurbished only if the exact quantity and drive size are confirmed. |
| Ultra, 256GB | Mac Studio M3 Ultra · 819 GB/s fact | Top chip launched Mar 2025 at about $7,099 est. Today: Apple Refurbished $8,149 (lower 28/60 chip, 2TB), single units fact; CDW-G $11,640, backordered fact | Mac Studio M5 Ultra (top 36/80 chip) · 1,200 GB/s fact; ships Sept 22; 512GB version late October | ~$10,799 · ~$9,900 est — 256GB needs the top chip; $11,299 at 2TB fact; education price not yet posted unknown | New vs refurbished at 2TB: about +39% est | +40% per box est — memory speed +47% fact | Buy new only when models need more than 128GB, and get an institutional quote first. The 512GB version: wait for its price and a delivered-unit test. |
| Fleet | Previous, refurbished | Newest, education | Newest, retail | Writing speed carried, previous → newest |
|---|---|---|---|---|
| 4-mini validation | $7,796 | $8,916 (+14%) | $9,596 (+23%) | 480 → 528 tokens/s (+10%) |
| 26-mini faculty & staff fleet | $50,674 | $57,954 (+14%) | $62,374 (+23%) | 3,120 → 3,432 tokens/s (+10%) |
| Hardware per unit of capacity (15 tokens/s) | $244 | $253 (+4%) | $273 (+12%) | — |
What it means. At education pricing the newest fleet costs about 14% more than refurbished previous-generation units and should carry about 10% more — roughly 4% more per conversation carried — and it is the only one UVU can order in quantity, with a warranty, on contract. That is why the plan is priced on the newest lineup, and why we call buying the old generation the one unforced error. One caution for Facilities: Apple's maximum power ratings rose (mini 140 → 155 W; Studio 145 or 270 → 480 W). Those are electrical ceilings for circuit planning, not expected draw — the one measured AI load on the previous mini was 46 W. The budget simulator now has a generation switch and the configurator prices every configuration both ways, so no comparison is hidden. est
What each machine can actually run — the matrix, not one number
A fair challenge from the read-through: why not put the most intelligent open models on 512GB Studios, and did we test every model against every machine, capacity included, with two models on one box? We had not. Appendix Q now does it — ten open models × six machines, each cell with memory fit, speed, conversations at once, cost per conversation, and an intelligence-per-dollar index. The short version is below; scores are the same 0–100 index explained in §02. fact measured or published · est derived, basis in Appendix Q
| Model (score) | Memory needed | Mini 48GB · $2,139 | Mini 64GB · $2,499 | Studio 128GB · ~$4,639 | Ultra 256GB · ~$9,900 | Ultra 512GB · price not posted |
|---|---|---|---|---|---|---|
| Qwen3.5-9B (22) — small fast lane | 7 GB | 8 | 8 | 8 | 8 | 8 |
| Qwen3.8-27B (52) — the daily driver | 18 GB | 8 (1 at 20+) | 8 (1 at 20+) | 8 | 8 | 8 |
| Qwen3.6-35B-A3B (32) — sparse fast lane | 21 GB | 8 | 8 | 8 | 8 | 8 |
| gpt-oss-120B (24) · Mistral Small 4 (20) | 64–69 GB | — | — | 8 (6–7 at 20+) | 8 | 8 |
| Qwen3.8-Flash-Next (56) — license on hold | 107 GB | — | — | 3 (memory-limited) | 8 | 8 |
| GLM-5.3-Flash (57) — multimodal, MIT | 182 GB | — | — | — | 8 (4 at 20+) | 8 (4 at 20+) |
| GLM-5.3 (60) — the open ceiling | 431 GB | — | — | — | — | 2 (none at 20+); 18 words/s each |
| Kimi K3 (60) · DeepSeek V4 Pro (53) | 850–930 GB | Do not fit any single Mac — frontier pool only | ||||
Same money, both ways
Five 48 GB minis cost about the same as the priced 256 GB Ultra. Five minis carry about 40 conversations of a score-52 model; one 512 GB node carries two conversations of a score-60 model. Multiplying quality by conversations gives 2,080 against 120 — that is a rough Utlyze comparison, not an industry measure. The same-money comparison against the 512 GB machine cannot be completed at all until Apple posts its price. The eight points matter for advanced coding (Terminal-Bench 88.2 vs 73.0; DeepSWE 66.9 vs 42.2) and complex professional documents (a 217-point lead on a work-products test); for tutoring chat, summaries, and ordinary drafts no matched test shows a difference. So the campus default stays minis and the frontier lane is the metered cloud pool — until a 512GB node earns its place. fact task results · est capacity
When a 512GB node earns its place — decided in advance
Path 1, recommended: a conditional shared research node. Kept in the plan as an unpriced option behind the router; bought only when the hardware pool reaches about $100,000, it takes no more than 15% of hardware spend, Apple has posted the price (late October), and a delivered-unit test passes. Path 2: a department or grant buys one once the price is posted (earlier access, low utilization, no redundancy). Path 3: no 512GB node — the same money buys minis, 128GB nodes, and cloud. What breaks first on the big node: reading long documents (a 131,000-token prompt took about 23 minutes on the previous generation) and a single point of failure. est
Two models on one machine — yes, and it is why 64GB exists
The serving software keeps several models loaded at once (llama-server's router mode, oMLX, LM Studio, vLLM-MLX), sharing the box's memory speed. A 48GB mini holds the 27B daily model plus a 9B fast model comfortably; the 64GB mini is the one that holds the 27B plus the sparse 35B fast model with production margin — that, not speed, is what its extra $360 buys. A 128GB Studio holds a 120B-class model plus the 27B; a 256GB Studio holds GLM-5.3-Flash plus the 27B. A 512GB node running GLM-5.3 at the preferred quantization has no room for a companion. Swapping from disk: the 27B in about three seconds, a 120B in eleven to thirteen, GLM-5.3 in about seventy-five — so small models are pinned and the big ones stay resident or stay off. fact software support and swap measurements · est pairings
What the mixed-work planning limit of four turns into
The same mini carries about 8 tutoring chats, 4 coding-help sessions, 2 sessions on a document it has not read (4 once that document is indexed), 1 coding agent, or 2 research-writing sessions at comfortable speed — a 22-fold spread in requests an hour hidden inside one number. Documents break first: a 30-page upload read cold by the 27B model takes one to three minutes before the first word, so uploads are indexed once and then served from the fast lane. A coding agent uses about 1,200 times the text of a chat, so it gets its own queue and a usage limit per person; public-class work may go to an approved cloud service under the written routing rule, never automatically or silently. The budget model buys against four, the mixed-work planning limit — the full table is in the two definitions above and in Appendix R. est
~$12,700
Four-computer test: equipment only. Four minis, the shared rack shelf, 10-gigabit switching, battery backup, and one node's storage upgraded to hold the model mirror. Staff time to run the validation (about a quarter to half of one full-time person) is real money and is priced visibly in the simulator — never hidden in a hardware number. est
~$66,900
Faculty and staff: equipment only, sized for exam week. 26 minis ($57,954) plus shared kit ($8,250) and the model mirror ($720) — $66,924 to buy, and about $81,600 over three years once AppleCare and power are added: about $13 per employee over three years, for the machines alone. That fleet covers the exam-week stress case (~104 simultaneous conversations); an ordinary day needs a fraction of it. Operating staff (evidence band: one to one-and-a-half full-time people) is the larger cost and is shown beside it, always. est
~$250,000+
Everyone: the broader program estimate. Students included, heavyweight machines in Engineering and Science, the research pool and the full adoption program funded, and headroom sized for exam week. This figure includes staff and training, unlike the two above. The simulator explores every point between — including the staffing line. est
Want to try a different mix of machines? The hardware configurator starts from the computers instead of the dollars: pick how many of each, and it shows people served on an ordinary day and in exam week, the model sizes it can run, cost to buy and to keep for three years, staff, and power — and prices that same configuration on the previous generation beside the newest.
Reality check kept in view: peak-hour demand for a 5,000-person population models at ~7 simultaneous conversations (base case); for the whole campus, ~65–70. Demand, not headcount, sizes the fleet — which is why the pilot's telemetry matters more than any projection in this guide. est
The budget simulator
There is no single right budget. There is a curve: every dollar buys coverage in a particular school, and every cut has a name. We built the curve into an instrument. Type any budget — or set each school's own budget, since money inside a school stays inside the school — and it solves the per-school supply-and-demand equations live: who is covered, where the first bottleneck forms, what falls below the line, and what the same coverage would cost to rent forever.
The four worked budgets
Every tool marks the same four budgets, and calls each one by the same name. Pick the one nearest the money you actually have.
- Four-computer test — $12,700. Four Mac minis in one college. It tests the planned software and hardware before anything bigger is bought. est
- First colleges — $45,000. The first colleges get their basic on-campus service, and the work of getting people to use it is funded. est
- Phase one — $120,000. Faculty, staff and degree-seeking students across the colleges, with a heavyweight computer where research needs one. est
- Ceiling — $260,000. Everyone, with spare machines and extra capacity for exam weeks. The planning model uses only part of this budget and shows the rest as an unspent reserve. est
Open the budget simulator
Slide the one-time equipment and launch money from a $12,700 four-computer test to a quarter-million-dollar campus build; the yearly staff cost and the three-year total are shown beside it throughout. Toggle students in or out, stress-test exam-week peaks, override per-school budgets, and print the scenario you like as a one-page sheet.
Launch the simulator → Or plan on the campus map →
The map view puts the same engine on real overhead imagery of the Orem campus: each college is a circle sized by its modeled demand, colored by what the budget covers — click to include or exclude colleges and watch the money re-flow.
In print: the budget scoping is printed on the "Budget scenarios at a glance" sheet that follows this guide — the coverage curve from $0 to $300,000 and four worked budgets, college by college. The live simulator and the campus planning map (college selection on real aerial imagery) are on the web version.
Prefer paper? Budget scenarios at a glance freezes the simulator onto one printable sheet: the coverage curve and four worked budgets.
Who will actually use it
These are estimates of UVU demand, not UVU measurements: no UVU usage data exists yet, and the pilot's own figures will replace every number here. They are anchored to the best published whole-campus results in the country — a full-campus service at Cedarville (82% of eligible people activated, 75% used it in a given week, 49 messages per active person per week), Virginia Tech's controlled pilot (78% weekly, median 11 messages a week), and 2026 national surveys (32% of students and 25% of faculty now use AI daily). The busiest hour of the day carries about 8% of that day's traffic. fact the other campuses' figures · est everything applied to UVU
| School | People (phase 1) | Peak streams (base) | Why it skews |
|---|---|---|---|
| Engineering & Technology (SCET) | 8,365 | 14.2 | CS/engineering: 62% regular use nationally; Burns's own college |
| Woodbury Business | 6,808 | 10.6 | Business: 51% regular use; case-work heavy |
| Humanities & Social Sciences | 6,669 | 8.5 | Big population, moderate skew; writing support |
| School of Education | 5,887 | 7.5 | Teacher-prep AI literacy mandate rising statewide |
| College of Science | 3,479 | 5.4 | Math/science tutoring + research pool users |
| Health & Public Service | 3,037 | 3.7 | Strong privacy constraints favor local serving |
| School of the Arts | 2,384 | 2.9 | Lowest text-AI use; vision models change this |
| Phase-1 total | 36,629 | 52.7 | generated by the same engine the simulator runs |
The planning insight: for the phase-one population, base-case peak load is about 53 simultaneous conversations — roughly 65–70 if concurrent-enrollment students join later. The financial risk is not under-buying — it is over-buying before measuring. est
The growth story, from the federal database — with a twist that matters. UVU's headline growth is real (32,670 → 48,669 since 2010; Utah's largest and fastest-growing big institution). But the government data shows 76–87% of the last decade's net growth came from concurrent-enrollment high-schoolers; the phase-one population — degree-seeking students plus staff — grew only ~0.6–1.3% a year. So year-three capacity planning is driven by adoption deepening, not headcount: the sourced band is 1.3–2.4× year-one demand, base case 1.9×. Utah is also more resilient than the national "enrollment cliff" (high-school graduates projected −5.6% by 2041 vs −10.3% nationally). Design integration and facilities for the 2.4× envelope; buy compute on the measured triggers. fact IPEDS/USHE/WICHE · est band
We stress-tested this model until it broke, and it broke in instructive places. The base rate is anchored to measured whole-campus telemetry, not surveys: Cedarville's full-year deployment works out to 2.15 requests per eligible person per day — 1.43 peak streams per 1,000 people against our 1.44 — and Berkeley's course bot and Michigan's token volumes land inside the same envelope. Two cases demand far more capacity: a 60-second reasoning reply doubles the streams needed (response time is the single biggest lever), and one 300-student class starting an assignment at 2pm needs ~143 slots to keep waits near a minute. So the design shows the queue, spills only public-class overflow to approved cloud routes, and carries numeric reorder triggers: buy six more minis when the weekly 95th-percentile load — the level exceeded only 5% of the time — holds above 70% of capacity for two weeks; open the overflow valve when any 15-minute wait passes 30 seconds. Automated/agent traffic gets its own queue class entirely — the one European campus that published full telemetry saw 99% of volume come from scripts, not chats. est model · fact anchors
One more honest correction, from the engineering review (Appendix R): the model's fixed 30-second service time is right for chat and wrong for a finals-week mix of chat, coding, documents, and research writing — the weighted service time is about 51 seconds, so the same arrivals occupy about 93 conversations instead of 54, a 70% understatement. That sits inside the 9× stress envelope the fleet is already sized for, so the numbers stand; the reason is now stated plainly. What fits inside no envelope is coding agents admitted freely: one agent turn uses about 580,000 prompt tokens, roughly 1,200 chats' worth, and if 5% of users started agents in the peak hour a 26-mini fleet would run at 133% of capacity with 16-minute waits. Agents therefore run in their own queue with a metered allowance and overflow to the frontier pool before utilization reaches 85%. est · fact agent telemetry (GitHub production trace, 3.2 million users)
Getting professors to actually use it
Machines are the easy half. The national record is brutal about the hard half: Cal State bought 500,000 seats and had activated at least 250,000 by spring 2026; voluntary training completion was 0.7% of students and 16% of faculty — and activation is not the same as use. UK government trials left ~20% of Copilot licenses untouched. Seats do not create usage. fact
What does? The programs that worked share five moves, and the evidence is consistent:
- Pay for a finished artifact, not attendance. CSU Bakersfield's $500-on-completion program hit 86% completion with a waitlist. Hawai'i paid 50 faculty $1,000 each to redesign an assignment. Small money, real deliverables. fact
- Gate access with a short training, then support weekly. Virginia Tech's trained pilot ran 78% weekly-active. Manchester reported 90% adoption in 30 days with cohort support. fact
- Teach inside the discipline. Faculty ignore generic prompt classes; they show up for "AI for your anatomy course." (Ithaka S+R, 19 institutions.) fact
- Design adjunct-first. In fall 2025 UVU counted 1,310 part-time and 802 full-time instructional faculty fact. That is 62% part-time est. It counts people, not full-time jobs. Any program built around salaried faculty misses most of the people who teach. Asynchronous modules, evening clinics, paid completion. UVU Common Data Set 2025–26, §I-1 (fall 2025); the narrower count is in Appendix J.
- Address integrity fears head-on. In every faculty survey we found, academic dishonesty outranks job loss as the top concern by a wide margin — but no single survey pins one share, so this guide no longer quotes one. The integrity evidence itself is calmer than the fear: no clean post-2022 jump in proven cases where institutions published totals, and AI detectors too unstable and uneven to prove anything (Appendix V). Lead with assessment redesign and authorship policy, never with "AI is inevitable." fact survey ranking · unknown any single share
Which departments first — evidence, not guesswork
A dedicated receptivity study of all seven colleges (courses taught, named faculty champions, policies published, grants won, Academy participation) lands on a clear sequence: wave one: Engineering & Technology (AI degrees, an AI lab in the new Smith building, the NVIDIA partnership, published department AI protocols) and Woodbury Business (AI courses in the 2026-27 catalog, the Institute's director on its faculty, three ready-made syllabus stances). Wave two: Education (a quiet powerhouse — AI in its accreditation mission, three faculty on UVU's UNESCO AI-and-education committee, statewide teacher-prep relevance) and Humanities & Social Sciences (largest writing load, biggest assessment strain). Science follows with research and lab pilots; Health & Public Service gets a separate safety lane (synthetic cases before anything clinical); Arts gets a rights-aware multimodal pilot once consent and licensing controls exist. fact signals · est sequence
The behavioral engine — engineered adoption, honestly
The adoption program is built on two published playbooks — Influencer's six sources of influence and Nudge's choice architecture — and then audited against thirty cases, from hospital hand-washing campaigns to Canvas rollouts, where mass adoption was actually achieved cheaply (Appendix N). That audit sharpened it into four vital behaviors: (1) run one real, bounded task, check the result, and apply, edit, or reject it within seven days — a reasoned rejection counts; (2) before the first graded assignment, publish an assignment-level AI rule with a disclosure example and a checking requirement, keeping a no-AI route wherever AI is not the learning outcome; (3) reuse and improve the workflow at its next natural occurrence within thirty days — grading, course copying, recommendation letters — rather than on an artificial weekly cue; (4) each peer catalyst personally helps three to five colleagues and shares one useful result and one rejected one. The main moment is not the first week of term, which is too late and too busy, but course-shell creation, four to six weeks before, where a removable Canvas block offers an active choice — permitted, limited, or prohibited — one example, and one support link. Defaults are engineered but honest — visibility is opt-out, participation is always opt-in; social-norm messages run only when true, current, and drawn from groups of twenty or more; catalysts are chosen by confidential peer nomination using the two questions the evidence favors ("who spreads useful teaching practices?" and "whom do you consult when a tool may not be ready?") — including respected skeptics and adjuncts — not by enthusiasm. Success is measured as verified adopters — applied use within seven days, one course rule published, a repeat within thirty — never as activations, and never in a way that could reach an evaluation file. The first wave costs about $20,000 (ten catalysts paid $1,000 each for a reusable artifact plus two clinics; twenty $500 artifact awards for adjuncts and faculty) and is expected to produce 110–150 verified adopters at roughly $133–$182 each; it scales to a $60,000 second wave only if cost per adopter, adjunct parity, and trust all hold. No honest plan promises a majority of instructors in one term at these budgets — reaching the 1,057 that a majority requires would take roughly $140,000–$190,000 directly, less if peer diffusion is real, which the first wave measures. est design · fact book frameworks & cited cases
The people already exist. Before any nomination round, the public record already names faculty in every one of the seven colleges who have built something with guardrails and said so out loud — a revision coach that refuses to draft, virtual patients for nursing, a marketing "boss" that argues back. The roster, with sources, is Appendix J. Peer nomination starts from real people, not hope. fact
Two safeguards this plan would require before launch — assessment assurance, and learning outcomes
Assessment assurance (Appendix V). One standard replaces the three templates: every material assessment carries an assignment-level Secure or Open label, the allowed and prohibited AI functions, a disclosure rule, and an individual proof point where the learning outcome demands it — set at the course-shell moment, with a report of unlabeled high-stakes tasks. Secure checks stay sparse: a few oral, written, coding, or practical proofs where UVU must show individual skill; everything else open, applied, staged, and explicit about allowed help — blanket redesign would burden honest students, adjuncts, and online learners most. Detector output alone is never evidence and cannot by itself cause a grade change, referral, interview, or sanction; a fair follow-up protocol gives the student the rule and the evidence, admits drafts and assistive-tool explanations, asks questions aligned to the original task, gives written reasons, and preserves appeal. Each paid catalyst's deliverable becomes one redesigned signature assessment, run and measured for workload and student trust — a code defense in Engineering, a changed-fact decision task in Business, staged research plus close reading in Humanities, microteaching in Education, a lab defense in Science, a synthetic practical in Health, a provenance-rich portfolio with live critique in Arts. Success counts substantiated cases, overturns, validity, workload, access, and trust — never detector flags. fact detector evidence · est standard
Learning outcomes (Appendix W). Access to a chatbot is not an intervention: the strongest trials that helped used course-aligned material, structured practice, feedback, and active student work; the strongest warning trial found unrestricted AI raised assisted practice scores and lowered the next unaided exam, and refusing premature answers removed that harm without adding learning by itself. No credible study shows a campus-wide deployment moved retention, completion, or DFW rates — so "verified adopters" stay an adoption measure, never a learning measure. The router becomes purpose-first (learn, produce, check, decide); learning-first tutoring — attempts, progressive hints, self-explanation, retrieval, transfer — is the course-practice default, with a labeled productivity mode where producing the product is the outcome and an attempt gate removable without shame or disclosure. A pre-registered, section-randomized pilot with delayed rollout, masked grading, and delayed unaided tests on problems the tutor never rehearsed is funded as part of the service. The contract: proceed to a larger trial only with no credible harm on delayed unaided learning; expand after the first credible year only if the delayed-learning estimate is positive with its lower 95% bound above −0.10 SD, transfer is not negative, DFW is no more than 2 points worse, and no subgroup shows repeated harm near 0.20 SD; reshape if assisted scores rise while delayed or transfer scores do not; stop a mode if its upper confidence bound sits below no effect. No adoption target overrides these rules. fact trial findings · est thresholds
UVU's head start
This is the part most universities would envy: the Kahlert Institute already runs an AI Academy (120 completers), curriculum grants, twice-monthly Power Hours, and a confidential one-on-one faculty coaching pilot. Board-reported employee AI use rose from 61% to 76% in a year. The plan here is not a new program — it is scaling what exists, paying for reusable artifacts and peer help rather than attendance, and measuring behavior instead of sign-ups. The 120 Academy completers are the candidate pool for nominations, not an automatic champion list. fact
The coach bot — phase two, built only when the support log proves the need
The idea: a UVU-tuned assistant inside the approved environment that asks a professor's discipline and course and walks them through a ten-minute build of something they will use this week, then hands hard cases to the Institute's human coaches. The audit's verdict was blunt and we accepted it: no precedent shows a bot alone moves faculty adoption unknown, and making it a launch dependency would turn a cheap behavior program into a software project. So the launch ships three faculty-reviewed templates, the existing Power Hours, and human help beside the task; the bot is built only when support logs show a repeated problem that templates and people cannot solve cheaply, with a clear test when it does ship — beat the self-service path on thirty-day repeat use, or be retired. Under the telemetry charter, its prompts are shielded from supervisors and barred from evaluations, with the charter's narrow legal exceptions stated in it, not hidden. The templates ship on day one regardless: allowed-use levels, disclosure and citation language, and an explicit rule that detector output alone is never misconduct evidence. est
Budget shape: adoption money is its own budget, never carved from hardware numbers. A one-college test costs $20–40k (30 faculty × $500 + two paid department leads); a cross-campus program runs $150–250k. At scale budgets the simulator holds an adoption line (default 18%) — because the evidence says the machines fail without it. est
The rights that make faculty willing — written in, not implied
Everything here is voluntary and faculty-governed: instructors decide whether and how AI enters their courses; students always keep an equivalent non-AI path with no penalty. Required adjunct work is paid at a stated hourly equivalent with no effect on rehire, and the intellectual-property terms for paid artifacts are handed to participants before enrollment. Usage telemetry lives under a published charter: aggregate reporting only, small groups suppressed, and a hard rule that no metric touches evaluation, tenure, discipline, or rehire. If pilot results are headed for publication, the IRB determination comes first. These are launch gates, and Faculty Senate, department chairs, adjunct representatives, and student government sit inside the governance from day one. est — commitments the plan binds itself to
How it's built
Every component is open, inspectable, and swappable; the stack fits UVU's existing systems like it was designed for them — because it was:
Front door
LibreChat (free, open-source, MIT license) supports UVU's existing Microsoft login (Entra) with two-step verification, per-user workspaces, and per-department budgets — capabilities of the product; wiring them into UVU's tenant is a named validation-phase task with Dx. Students never see an API key (the technical password that connects to a model). Two conditions from the reviews: its accessibility must be proven on the exact pinned build with real assistive technology, and UVU's existing Wilson assistant may be the better student front door — the fleet can serve behind Wilson first, and no competing student front door launches until UVU decides Wilson's role (Appendices X, Y). fact product · est UVU wiring
Traffic control
LiteLLM routes each request to the right machine and the right model, enforces quotas, fails over automatically within the routing table's data-class rules, and exposes live capacity — the "how busy are we" gauge that feeds the simulator's real-world twin. fact product · est config
Serving
llama.cpp pinned builds on each Mac — the most battle-tested engine on Apple hardware — with popular models held resident (no cold-start roulette) and cold models swapped on demand. fact
Fleet care
Enrolled in UVU's existing Jamf Pro (already mandatory for university Macs), monitored with the same dashboards Digital Transformation already runs, updated in staged rings with tested, timed rollback paths for the app, the models, and the database. fact Jamf · est ops design
The one hard problem: isolation
Each of the major Mac serving engines has had at least one documented cache bug where text from one user could surface for another (llama.cpp #27148 — open; MLX‑LM #965 — fixed; Ollama's MLX path #17599 — open; LiteLLM response cache #29955 — open). This system therefore ships with every shared cache disabled, one active generation per user, and a standing "canary" test — synthetic users with unique markers hammer the system across concurrency, restarts, retries, and failover before every update, and any cross-user leak fails the release. The battery's first production run already happened (§04). Privacy here is an engineering discipline with a regression test, not a promise. fact named issues · est controls
| UVU data class (official scheme) | Local Mac route | Labeled cloud routes | On overflow / outage |
|---|---|---|---|
| Public — general coursework, published material | Yes (default) | Yes — destination shown before send | May spill to free layer or approved cloud |
| Sensitive — student information, most institutional data | Only route | Never | Fails closed — waits or declines, never reroutes |
| Restricted — the most protected categories | Excluded from the pilot entirely | Never | Not applicable — this system does not accept it until Dx approves a design for it |
Built in after the engineering, safety, and law reviews
Lanes by kind of work
Short chat and coding turns ride the fast lane; uploads are indexed once and answered by retrieval (2,000–6,000 relevant tokens per question, never the whole file); exact course prefixes are cached and a student's follow-ups stay on the box holding that cache; a large prompt the machine has not seen before goes to the fast, lighter model rather than the slow, heavy one. Coding agents get their own queue, a usage limit per person, and are kept on the machine holding their cache (“session affinity”); public-class work may go to the paid research pool under the written routing rule before a machine reaches 85% of its capacity, never automatically or silently (Appendix R). est
Crisis handoff, before broad student use
A separate safety detector — rules plus a tested classifier, never keywords alone; an honest message ("I'm an AI tutor, not a crisis service — if anyone is in immediate danger call 911 or UVU Police; for a suicide or emotional crisis call or text 988"); the smallest useful packet to an acknowledged human queue, not an email; a restricted case record; graduated handling for low-confidence signals; no discipline on classifier output alone. The trained person, not the model, decides Title IX, Clery, abuse-reporting, and police questions (Appendix S). est design · fact duties
The first pilot leaves the dangerous features out
No tools, no web fetching, no shared document retrieval, no automatic cloud spill; outbound network denied from inference and parsing workers; uploads quarantined — allowed formats only, real type verified, size capped, active content disarmed, parsed in a no-egress sandbox. Each feature returns only behind its own control and test gate. est
Labelled, accessible, on the record
"UVU AI assistant — not a human" shown before and throughout every chat, the safe harbor under Utah's AI disclosure law; counseling, clinical, legal, and financial advice kept out of the general assistant. Accessibility (WCAG 2.1 AA) is a launch gate tested with assistive technology on the real chat, uploads, streaming answers, and generated documents — the federal deadline for UVU is April 26, 2027. A records map before launch says which chats are education records, which retained data are open-records material, and what retention applies (Appendix T). fact law · est design
A finished control plane — the settings and small databases that run the service
Redundant routers that hold no state of their own with tested failover; identity outage fails closed for new sessions with a status page outside the login path; approved model weights mirrored on every compatible node plus a checksum-verified mirror; a same-class spare or a documented degraded route for the one heavy node; backups of configuration, prompts, manifests, and required state with a proven restore; recovery targets — a failed mini in five minutes, public-data fallback to the metered route in thirty, the restricted route fails closed and rebuilds in four hours; a change freeze from seven days before finals through the grade deadline (Appendix S). est
Model admission record
Every model enters through a record: repository and immutable revision, download source, cryptographic hashes, exact license and notices, model card, conversion recipe, scan and test receipts, named approvers, a read-only local copy and an independent mirror. Licenses are pinned by exact checkpoint, not by family — the daily driver and both 128GB workhorses are Apache-2.0; Qwen3.8-Flash-Next is on hold under its community license; GLM-5.3 and Kimi K3 sit behind counsel review (Appendix T). est
One trust contract, everywhere
The same source states, uncertainty states, refusals, report route, and human route across Wilson, LibreChat, and every routed model. Every sourceable claim opens the exact supporting passage; every answer shows whether it came from approved UVU material, an uploaded file, the open web, or model memory; an evidence-sufficiency gate lets the assistant say "not found" or "sources conflict" instead of filling the gap; no raw confidence percentages until UVU validates them per model, task, and course. Errors become repair cases with a receipt, an owner, a status, and a visible correction notice. Tutoring is student-first and hint-first, enforced outside the model prompt; calculators and code tests are used wherever the task permits. Results are published by task and route, never as one blended "accuracy" score (Appendix X). est design · fact calibration studies
Accessible text chat first; the rest behind its own gate
The front door is a conditional candidate until the exact pinned build passes an assistive-technology audit (VoiceOver, NVDA, JAWS, switch access, magnification) on complete conversations; streaming is treated as a visual effect — short states and complete messages are announced, never every token; focus stays under the student's control. Uploads, math, diagrams, PDF export, and voice each pass a separate gate later. Interface language, content language, and reply language are separate student choices, and every deployed model passes a paired UVU English/Spanish benchmark before Spanish is promised. Paid disabled and bilingual student testers under an IRB determination; a funded accessibility owner; a release-specific conformance report and a barrier-response route; phone, weak-Wi-Fi, and interrupted-upload tests; generated documents as accessible HTML first (Appendix Y). fact standards and dates · est priorities
Logging, precisely: chat history lives in the user's own workspace on UVU-controlled systems (that's what makes conversations resumable), with user deletion and a published retention rule — stated exactly: "Chat content logging is off by default. UVU retains only the data listed in its records notice, except when a user saves a conversation or a scoped legal or records hold requires preservation." A saved chat tied to a student is a FERPA education record; retained chats and metadata are records under Utah's open-records law — classified and, where required, produced, never promised "confidential" (Appendix T). Operational analytics see metadata only — tokens, latency, model, department — never prompt text; capacity planning and department accounting run entirely on those counters. Security-event logs required by Policy 447 capture system events, not conversation content, and the forensic procedure for a suspected leak is a written runbook with CISO sign-off. The full data-flow diagram, retention table, and runbooks are the first deliverables of the validation phase — drafted by us, owned and approved by Dx before any pilot user logs in. The control plane (the small databases and router state behind the curtain) is named, priced, backed up, and owned by the service's operator — with Dx holding approval, audit, and takeover rights — in that same package; no invisible single points of failure, and no borrowed Dx staff time. est
Stands on its own; fits the house: this service is designed to exist independently — its own hardware, front door, model operations, staffing, and budget — because everything UVU already runs (the employee AI Gateway launched in July, Ask Wilson in 44 courses, the shared state compute) is spoken for by the people using it today. Nothing here borrows their capacity. What it needs from UVU is small and specific: a room with power and a network drop, sign-on federation with the campus Entra tenant, a data-class ruling, and a sponsor. Three harness tiers ride the same router: a chat front door for everyone (LibreChat), a sandboxed coding agent for builders (OpenCode or Cline class), and a governed research search for scholars (Onyx class) — compared product by product in Appendix L, with the whole market of 171 harnesses — adoption, growth, funding, stability, and the dead-or-declining list — in Appendix M. The surfaces it plugs into are already mapped from public sources — identity, Canvas, portal, ticketing, security tooling, network — and the questions only UVU can answer are down to five (Appendix K). Where UVU wants the two worlds to meet, the connection points are ready — the fleet can serve as a local engine behind the Gateway, and course assistants like Wilson can call it — as optional integrations, never as dependencies. Dual sponsorship still fits the governance pattern UVU's documents describe: the college owns the academic mission; Dx owns identity, security review, and the Policy 445/447 package (data flows, classifications, runbooks), which arrives drafted for Dx and the CISO to own, change, and approve before any pilot user logs in. fact Gateway/Wilson · est arrangement
The path through the university
- No new bid expected. Apple minis and Studios are on NASPO ValuePoint contract MA 23003 / Utah addendum PA4282 — the competition requirement is already satisfied through June 2027; UVU Procurement confirms the SKUs and process (their call, made easy). fact
- Utah's trial-use statute (§63G‑6a‑802.3) was built for this: minimum quantity, defined test measures, no obligation to scale. Last-published signature levels: dean/executive at $25k+, VP at $50k+ — a $12k validation sits comfortably low (current schedule to be confirmed; the live PDF is down). fact verify current schedule
- Week-one, in parallel, at zero cost: request binding seat quotes at four scopes (500 / 2,112 instructors / 6,123 employees / all-campus) so the cloud comparison becomes like-for-like; apply for a Redtail early-access allocation (Utah's new 264-GPU statewide AI system — UVU faculty are eligible today) for research workloads; ask USHE whether a system-wide seat deal is in motion (Burns sits on the statewide AI task force); and stand up the governance charter (§07) with Faculty Senate at the table. fact eligibility · est plan
- Reviews to plan for, not fear: academic tech committee (ATSC), IT Oversight, security (Policies 445/447), and accessibility as a tested release gate — the web accessibility standard (WCAG 2.1 AA) with disabled-user testing, a funded remediation owner, and an equally effective alternative for every required task (Policy 452). UVU's campus AI policy (draft 441) is still in the approval pipeline; the pilot charter follows the enacted policies and adopts 441 when it lands. Timeline estimate for a funded on-premise pilot with no student-record data: 8–16 weeks to users, owned by those reviewers, not promised by us. Campus-wide: 6–18 months, phased. est
- Two gates added by our own gap-hunt: a one-page Facilities + Network + CISO signoff (named room, measured power/cooling headroom — 26 minis draw ~4kW and ~1.15 cooling tons — and the 10GbE build-to-order option specced, +$100/box) before any purchase order; and counsel-approved interim service terms (acceptable use, a "report this response" route, an incident taxonomy that carries UVU's existing 24-hour reporting duties into the service, and a narrow rule that the assistant never makes official decisions — it explains, cites, and hands off) before the first user. est
- One scheduling courtesy: Dx cuts the campus portal over to Pathify on November 16 — the pilot goes live before or after that window, not during it. fact
- Three rules that arrived in 2024–2026 (Appendix T): Utah's AI disclosure law — the service labels itself "UVU AI assistant — not a human" before and throughout every chat, the statute's safe harbor, and keeps counseling, clinical, legal, and financial advice out of the general assistant; since May 6, 2026 a state-funded university may ask Utah's Office of AI Policy for a joint interpretation of the one open question (is a free campus assistant a "consumer transaction"?). The federal ADA Title II web rule — WCAG 2.1 AA for the real chat, uploads, streaming answers, and generated documents by April 26, 2027, tested with assistive technology, not a page scanner. Records — a records map before launch: which chats are FERPA education records, which retained data are open-records material, which retention series applies (advising records: five years after separation), and how a legal hold stops deletion. fact law and dates · est design
- Ride the state's programs; count none as funding yet: the Utah Board of Higher Education's statewide AI direction (December 2025) and its AI Task Force (May 2026 — Dr. Burns is a named member); the free statewide AI Workforce Credential (available since July 1, 2026) goes into faculty and student onboarding instead of a competing certificate; the 2026 budget's $15 million shared AI research data center for Utah's public universities and $3 million ongoing AI workforce accelerator — UVU's share is unknown until an allocation exists; the NSF State and Regional AI Infrastructure Hubs solicitation (deadline November 4, 2026; $4–12 million over five years; one award per state) as a Utah consortium, not a solo request. fact programs and dates · unknown UVU's share
- Who pays, and who runs it (Appendix U): $0 from general student fees — Utah's fee rule bars them for instruction, academic support, and administration; central IT funds the shared base (fleet, identity, security, monitoring, refresh reserve, baseline research pool, the accountable owner); colleges and grants buy marginal nodes and reserved capacity; showback first, chargeback only for reserved or unusually heavy use; a college joins only after a signed three-year commitment covers its hardware and half its marginal staff time. Staffing: a professional owner with a paid student operations team beneath — more coverage at the smallest band, about a quarter off staffing cost at the larger bands (the simulator's staffing switch shows both). Buy the first fleet outright unless a written education-lease quote beats ownership; plan a year-three review and a year-four refresh, with three-year resale values near half of purchase price. est · fact fee rule, wage anchors, financing terms
- Buy in gated waves, never 26 at once (Appendix Z): four minis after the delivered-unit test (load the model; 1k/8k/32k-token tests; four streams at the 10-token-per-second floor; batch 1/2/4/8, cold-load, peak-memory, and 24-hour soak; identity, logging, isolation, accessibility, and recovery checks); up to eight more only after half the pilot returns at their next natural task within 30 days, measured demand uses at least 60% of capacity for four teaching weeks, and support stays inside the staffed plan; up to the planned 26 after spring results and a refreshed local-versus-cloud comparison. Waiting for a rumored next chip means about 15 months — 60 validation device-months and 390 fleet device-months forgone — while a machine 25% faster would repay no more than 9.6 months of waiting inside a 48-month life; so delay only for a failed must-have, an official Apple announcement within 90 days that clears a stated bar (the required memory, or 35% more measured throughput for no more than 20% more cost), a dated grant award within 60 days, or a blocking accessibility or security result. The 512 GB line stays unpriced until Apple posts it and two approved research workloads need it. A dated 24-month watchlist — Apple events, model checkpoints, grant deadlines, state milestones, the April 26, 2027 accessibility deadline — closes the appendix. fact cadence and arithmetic · est gates
- The full engineering design lives in its own document: the architecture dossier — component contracts and a substitution ledger, the data-class routing with per-contingency behavior, the model-provenance and release-integrity rules, 24 pre-decided contingencies ("if the CISO rejects a component… if quotes come back cheap… if one department eats 80% of usage…"), five ready decision memos, and the eight weekly tripwires with owners. est
- Money paths beyond the college budget: the $5M Utah AI Moonshot and $45M HERFP — full proposals due Oct 6, but only for teams whose notice-of-intent was filed by Aug 28, so the first question for Burns is whether UVU filed one (if not, target the next cycle); the state's $15M shared AI-compute appropriation; NSF NAIRR compute; the Kahlert relationship for a named lab; and the free research programs in §11. fact
The 14-week sequence — every step behind its gate
Weeks 0–2: dual sponsorship (college + Dx), governance charter with Faculty Senate, order 4 minis on the co-op contract for the Sept 22 shipment wave — accepted only after one delivered unit passes the acceptance test (Appendix Z) — and fire the free parallel actions: four-scope seat quotes, Redtail application, USHE system-deal question, IRB determination request. Weeks 2–6: technical validation. Stand up the stack on the real hardware, benchmark the exact models, run the isolation battery and a soak, wire up Microsoft Entra sign-on — UVU's own campus login, with multi-factor authentication — and deliver the written security package (data flows, how long anything is kept, written procedures for operators, and a chart naming who does each task and who is accountable for it) for Dx and CISO approval. Weeks 6–14 (opens only when those gates pass): a voluntary instructor cohort — proportionally adjunct, paid, IP terms in hand — plus an opt-in student cohort in SCET, telemetry under the charter. Week 14: the build/rent/hybrid decision made with UVU's own curves and the returned quotes in the same spreadsheet — with a Moonshot/HERFP application already in flight. est
The risks, weighted and named
| Risk | L | I | D | Score | Safeguard required before launch |
|---|---|---|---|---|---|
| A crisis disclosure lands in a tutoring chat and no trained person sees it | 3 | 5 | 3 | 45 | Separate safety detector, honest crisis message (988 / 911 / UVU Police), an acknowledged human queue rather than an email, restricted case record, graduated handling of low-confidence signals — built before broad student use (Appendix S) |
| Assisted scores rise while unaided learning falls — the strongest warning trial's exact pattern | 3 | 5 | 3 | 45 | Learning-first tutoring by default (attempts, hints, self-explanation); delayed unaided tests on unrehearsed problems in every pilot course; the learning-outcomes contract's expand / reshape / stop rules; adoption never overrides them (Appendix W) |
| The front door fails the assistive-technology audit, or Spanish quality is worse and nobody notices | 3 | 4 | 3 | 36 | Front door is a conditional candidate; exact-build audit with real assistive technology; accessible text chat first; paired English/Spanish benchmark per deployed model; paid disabled and bilingual testers; an alternative front door held ready (Appendix Y) |
| Protected data pasted into the wrong tier — in one 2026 clinical study, 23.5% of health students who used AI on placement admitted disclosing a patient identifier; education has the same shape with minors (HB 273), business with licensed case content | 3 | 4 | 3 | 36 | The local Sensitive lane exists precisely so protected work has a safe home; department onboarding states permitted use per course (the appendices map all 51 units); Public-cloud tiers carry paste-guards and no-PHI defaults; placement work follows the partner's rules |
| Adoption stalls; hardware idles | 4 | 4 | 2 | 32 | Adoption fund with a $20k first wave (catalysts paid for a reusable artifact plus peer help; artifact awards for adjuncts and faculty; clinics); a "verified adopter" measure — applied use within seven days, repeat within thirty — instead of activations; scale gates on cost per adopter, adjunct parity, and trust; the coach bot deferred until support logs show a repeated need |
| Cross-user privacy leak | 2 | 5 | 3 | 30 | Caches off, per-user isolation, canary battery on every release, content-free logging |
| A minor's data or a patient record enters the service before the rules are settled | 2 | 5 | 3 | 30 | Concurrent-enrollment students stay out until eleven named gates pass — parental permission alone is not one; placement patient data fails closed; synthetic clinical cases only; the FERPA-or-HIPAA status of UVU's own health records resolved with counsel (Appendix S) |
| First-mover risk: no public campus Apple-serving precedent found | 3 | 3 | 3 | 27 | Pilot-gated scale, hosted lane running beside it, published exit thresholds, fallback documented (§12) |
| New M5 hardware underperforms same-chip projections | 2 | 3 | 4 | 24 | Nothing beyond 4 boxes ordered until 2-week validation on real hardware |
| Coding agents admitted freely swamp the shared queue — one agent turn ≈ 580,000 prompt tokens; 5% of users starting agents in a peak hour puts a 26-mini fleet at 133% of capacity | 3 | 4 | 2 | 24 | Agents run in their own queue with a metered allowance, session affinity, and overflow to the frontier pool before 85% utilization; chat is protected from long work (§06, Appendix R) |
| Long documents read cold take minutes to the first word; users conclude "it's slow" | 4 | 3 | 2 | 24 | Uploads indexed once and served by retrieval from the fast lane; exact course prefixes cached; follow-ups stay on the box holding the cache; cold large prompts never hit the dense model (Appendix R) |
| Accessibility deadline missed — the ADA Title II web rule requires WCAG 2.1 AA by April 26, 2027 | 3 | 4 | 2 | 24 | Accessibility is a launch gate: assistive-technology testing of the real chat, uploads, streaming answers, and generated documents; tagged PDFs; a funded remediation owner (§09, Appendix T) |
| A student is accused on an AI-detector flag | 3 | 4 | 2 | 24 | Binding rule: detector output alone is never evidence and cannot trigger a grade change, referral, interview, or sanction; fair follow-up protocol; Secure/Open labels set before term (Appendix V) |
| Harmful output or misuse incident in the early weeks | 3 | 3 | 2 | 18 | Interim service terms before first user; "report this response" route; incident taxonomy with urgent lanes; narrow no-official-decisions rule; pre-written response memo |
| A model's license is misread — the family is not the checkpoint | 3 | 3 | 2 | 18 | Licenses pinned by exact checkpoint in the model admission record; Qwen3.8-Flash-Next on hold under its community license; GLM-5.3 and Kimi K3 behind counsel; the daily driver and both 128GB workhorses are Apache-2.0 (Appendix T) |
| The one Ultra node, the router, or identity fails during finals | 3 | 3 | 2 | 18 | Redundant stateless routers; identity outage fails closed with a status page outside the login path; a same-class spare or a documented degraded route for the heavy lane; recovery targets (a failed mini: 5 minutes; public-data fallback: 30 minutes; restricted route: 4 hours); a change freeze from seven days before finals (Appendix S) |
| Stress peaks swamp capacity (compound envelope ≈9× base) | 4 | 2 | 2 | 16 | Visible queue + Public-class overflow valve + numeric reorder triggers (§06) |
| Ongoing ownership unfunded after year one | 3 | 4 | 1 | 12 | Operator staffing priced inside the plan (~0.25–0.5 FTE pilot), no borrowed Dx time; Utlyze support offer; Dx holds takeover rights if it ever wants them |
| Open-model releases slow (policy/geopolitics) | 2 | 3 | 2 | 12 | Today's models already exceed the use case; fleet still serves them; frontier pool unaffected |
| Approval sequencing slips | 3 | 2 | 2 | 12 | Dual sponsorship, drafted security package at first review, bounded-validation framing |
| Software-stack churn (fast-moving projects) | 3 | 2 | 2 | 12 | Pinned versions, staged rings, rollback manifests; every component swappable |
| Wave order reads as favoritism between colleges | 3 | 2 | 2 | 12 | The scoring formula is published (Appendix A) and arguable; the planning map lets any college be included or excluded live; the free layer reaches every college on day one; waves 2–3 carry calendar dates, not "later" |
| A "free" tier the campus leans on changes terms, rotates its models, or trains on the traffic | 4 | 2 | 1 | 8 | Free tiers are never on the critical path: sandboxes and Public-class burst only, behind the data-class rules; contributor/training tiers barred for UVU data; every outside route has a pre-qualified substitute (Appendix L) |
| Arts, music, and language programs need speech and media the core stack doesn't serve | 4 | 2 | 1 | 8 | Named openly: the local stack is text + vision; routes R5/R6 attach separate approved speech and media services, rights charter first — no pretending one box does everything |
| Model licenses shift under a key model | 2 | 2 | 2 | 8 | Daily driver is Apache-2.0; immutable approved checkpoint retained; two alternates pre-qualified |
Said plainly: in every search we ran, no public example surfaced of a university running campus-wide AI on Apple hardware — the schools that self-host (Münster, UCSD, Florida, Indiana, Purdue) all run Linux/NVIDIA on central research computing. UVU would likely be first, which is exactly why the plan buys four boxes before it buys twenty-six, measures everything, keeps the hosted lane running beside it for honest comparison, and publishes exit thresholds up front. If the validation says the Macs lose, the same gateway, sign-on, governance, and adoption work carries straight onto whichever backend wins — nothing is wasted but the price of four excellent computers UVU keeps anyway. fact search result · est framing
The models
| Model | Role | Score | License | Runs on |
|---|---|---|---|---|
| Qwen3.8‑27B | Daily driver — chat, tutoring, coding, vision | 52 | Apache‑2.0 | mini 48GB |
| Qwen3.6‑35B‑A3B | Fast lane — classrooms, batch, overflow, cold documents; pairs with the 27B on a 64GB mini | 32 | Apache‑2.0 | mini 48GB |
| Mistral Small 4 (119B) | Heavyweight workhorse, stable (gpt‑oss‑120B, score 24, Apache‑2.0, is the alternate) | 20 | Apache‑2.0 | Studio 128GB |
| Qwen3.8‑Flash‑Next | Quality ceiling for 128GB — on hold: only three conversations fit, and its license needs counsel (Appendix T) | 56 | Qwen Community License — separate license for hosted services | Studio 128GB |
| GLM‑5.3‑Flash (320B/18B) | The 256GB shared node — multimodal, premium coding lane, eight conversations at once | 57 | MIT | Studio 256GB+ |
| GLM‑5.3 (753B/40B) | The open ceiling — conditional research node only: two conversations at once at about 18 words/s (§04) | 60 | Custom GLM license — counsel review | Ultra 512GB (late Oct; price not posted) |
| Kimi K3 · DeepSeek V4 Pro | Considered — do not fit any single Mac (850–930GB at 4-bit); frontier pool only. K2 Horizon (Sept 3) is a watch item | 60 · 53 | Custom · MIT | Cloud only |
| Muse Spark 1.3 (Meta) | Frontier-pool candidate — coding and agent route, standard no-training tier only; never the contributor tier | 61–62 | Closed — API only (weights "on the roadmap") | Meta cloud, metered |
For the mathematicians and physicists — Burns's hardest requirement — the frontier pool routes heavy research to the strongest current models: Claude and GPT‑5.6-class models now solve International Math Olympiad sets essentially perfectly, score ~94% on graduate-level science exams, and produced at least one passing referee grade on 7 of 10 unpublished research problems in a controlled trial this summer. Metered access for 200 heavy researchers models at ~$22k/year — and the free lanes come first: Anthropic's scientist program (free seats + up to $50k API credits per project), OpenAI's researcher seats, NSF NAIRR compute. fact scores & programs · est pool cost
The alternatives, priced honestly
Any hardware review will ask "why not NVIDIA?" Here is the honest comparison with NVIDIA's own campus-friendly box, the DGX Spark, using verified prices and published measurements:
| System | Price | Memory | Bandwidth | $ / GB |
|---|---|---|---|---|
| NVIDIA DGX Spark (4TB) | $4,699 | 128GB | 273 GB/s | $36.71 |
| Mac mini M5 Pro 48GB | $2,299 | 48GB | 307 GB/s | $47.90 |
| Mac Studio M5 Max 128GB | $5,099 | 128GB | 614 GB/s | $39.84 |
| Mac Studio M5 Ultra 256GB | $10,799 | 256GB | 1,200 GB/s | $42.18 |
Where Spark genuinely wins — and earns a place in this plan: it runs CUDA, the language of every datacenter GPU a UVU graduate will ever touch. For the ML courses, fine-tuning labs, and NVIDIA-certification work the Smith College already teaches (UVU has an active NVIDIA partnership), one or two Sparks as named lab machines are the right purchase — for teaching, not for serving. For serving inside a Jamf-managed campus, the published evidence favors the Macs on management fit, power, and single-user speed, while NVIDIA/Linux keeps the more mature high-concurrency serving stack — which is precisely what the validation phase measures head-to-head before any fleet money moves. fact capabilities · est verdict
Every option on one page — where each honestly wins
| Option | Cost basis (population · horizon) | Data control | Ops burden | Wins when |
|---|---|---|---|---|
| Local Mac floor (this plan's Layer 1) | ~$80k machines 3-yr + ~1–1.5 FTE/yr · 6,123 staff · stress-sized | Best — never leaves campus | Real: fleet + model ops | Data rules bite, usage is daily, teaching value counts, staffing is funded |
| Hosted open-model APIs (zero-retention routes) | ~$2.5k/yr (5k users) to ~$23k/yr (45k) + gateway ops | Contractual only; data leaves | Light-moderate | Variable load, Public-class data, fastest start — runs as our comparison lane |
| Negotiated frontier seats (ChatGPT Edu class) | $118–184k/yr · 6,123 staff if mega-deal rates; $880k+/yr at quoted mid-size rates | Vendor contract | Lightest | The quote comes back near $19–30/user·yr and polish beats control |
| Utah shared resources (Redtail) + AI Gateway | Unknown — quote requested week 1 | State-governed | Application + queue realities | Research bursts, training, classes — complement, not yet a daily-chat backend |
| DIY 128GB PCs (AMD Strix Halo class, ~$3.7k) | ~30% less hardware $ than Studios · same horizon | Same as local | Heavy — Linux/driver fleet outside Jamf | 128GB under $4k matters more than management fit and warranty |
| Free tiers (OpenRouter free models, AI Studio, OpenCode Zen "contributor" access to Meta Muse Spark 1.3) | $0 · capped at 50–1,000 requests per day per account, ≈1.4% of phase-1 daily demand with the three most generous recurring offers combined under one institutional account; multiplying accounts is barred by the providers' terms | Weakest — consumer terms; contributor tiers train on the traffic | None — and no admin plane, SLA, or continuity | Sandboxes, model evaluation, Public-class burst — never the floor (Appendix L) |
| NVIDIA DGX Spark teaching computer (1–2 units) | $4.7k each · one-time | Local | Medium (separate OS world) | Teaching NVIDIA's own software and training small models — always, in a named laboratory role |
About this work
How this was made. Twenty-seven independent research passes ran against the live web across September 2–3, 2026 — model leaderboards, Apple's store, signed contracts, Utah statutes, UVU's own policy manual, published campus telemetry — plus direct pulls from the federal IPEDS database (enrollment trajectories for UVU and eight Utah peers; UVU degree production by field). Every research thread is logged as its own tracked issue. Claims carry grades: fact verified at a primary source that day; est derived, with the math shown; unknown honestly unresolved. Then the work attacked itself, twice: an adversarial pass re-verified the eighteen most load-bearing numbers (ten confirmed, eight sharpened, zero broken); a paid-skeptic pass built the best case against this plan (its strongest points now live inside §09 and §12); three independent adversarial reviews — a blocking CIO, a skeptical CFO, a wary Faculty Senate — attacked the draft; a second audit round then graded the fixes and re-checked every number, catching an engine bug and several inconsistencies, which were corrected before publication. A fresh-eyes review of the published site on September 3 found more, and those corrections are in this version; the calculators and the written figures were reconciled against one another as part of it. The demand model was stress-tested to its breaking points (§06), and the proposed stack was physically benchmarked on our own hardware, isolation battery included (§04). Known remaining gaps are stated where they live, not hidden. Failed searches are logged, not papered over. This is how we believe AI-era consulting should be done — and the method is itself a demonstration of what UVU's own people can build with these tools.
About this work, and an offer
Utlyze prepared this research and plan because we believe Utah's largest university deserves a first-rate answer to the most important infrastructure question of the decade — and because the person asking it is asking exactly the right questions.
If UVU wants help making it real — architecture, the isolation test battery, the coach bot, the pilot build, training, or ongoing operations — we would be honored to help, and we will work inside whatever budget the university actually has. If UVU builds it without us, this guide is still yours, and we will still cheer.
— The Utlyze team · utlyze.com
Key sources
Live leaderboards: Artificial Analysis, LMArena, Epoch AI lag analyses. Hardware: Apple newsroom (Aug 25, 2026), Apple education store configurations, NVIDIA DGX Spark documentation, llama.cpp/oMLX published benchmarks. Contracts & pricing: CSU–OpenAI signed agreement, CSU renewal reporting, U. Maine System, U. Colorado, Microsoft education pricing, Google education pricing. UVU & Utah: UVU institutional data, Kahlert Applied AI Institute, UVU Policies 445/447/452, NASPO Apple contract, Utah Code §63G‑6a, UVU HERFP / AI Moonshot notice. Adoption evidence: Virginia Tech pilot report, Cedarville telemetry, Tyton 2026, Ithaka S+R, CSU rollout reporting. Precedents: Münster uniGPT, UCSD TritonGPT, U‑M ITS reports. Teaching, trust, accessibility, and timing reviews (Appendices V–Z): randomized trials and meta-analyses on AI tutoring 2023–2026 (Harvard, Turkey, Nigeria, Khan Academy evaluations), the Challenge Success and ICAI integrity data, independent AI-detector tests, the AI Assessment Scale, W3C WCAG 2.1 and DOJ guidance, LibreChat's public issue tracker, Apple's release history and newsroom, the Artificial Analysis index history, Wikimedia Commons image licenses. Engineering, safety, law, and money reviews (Appendices Q–U): Artificial Analysis model pages and methodology, the oMLX and llama.cpp benchmark databases, MLX-LM quantization benchmarks, GitHub's production coding-agent telemetry (arXiv 2608.00101), Utah Code Titles 13, 53E, 53H, 63G and 78B, the DOJ ADA Title II rule and its 2026 interim rule, Department of Education FERPA/PPRA guidance, HHS HIPAA guidance, USHE and Governor's Office announcements, Apple Financial Services terms, USHE R516, the NSF 26-513 solicitation, EIA electricity data. The complete source log — several hundred dated citations with the searches that failed as well as succeeded — is available on request.