UTLYZE
An independent opportunity plan for Utah Valley University

Own the hardware.
Cloud for when you need it.

Utah Valley University can put a capable AI assistant in front of its faculty, staff, and degree-seeking students in three layers: the free cloud tools it already licenses, a small fleet of Mac minis it owns for daily and data-sensitive work, and pay-as-used access to the strongest paid models where research needs them. It starts with a four-computer test that costs about $12,700 for the equipment package. Staff, training and any later expansion cost extra, and every one of those costs is priced in the open, next to it.

Prepared independently by Utlyze. Not a publication of, endorsed by, or affiliated with Utah Valley University.

Prepared for Dr. Barclay Burns, Chief AI Innovation Officer, and the Kahlert Applied AI Institute
Prepared by Utlyze
Evidence date 2026-09-02/03 · every claim graded fact estimate unknown
Print Press Print in the menu. This page comes out as a clean PDF.
$2,139
education price of the serving unit — Mac mini M5 Pro 48GB fact
$12.7k
complete 4-mini validation bundle: hardware + rack/network + AppleCare + 3-yr power (staff time priced beside it, in the simulator) est
~$19/user·yr
Cal State's 675,000-user renewal rate ($1.60/mo, their figure) — a scale benchmark, not a UVU quote fact
60 vs 66
how close free-to-download models are to the best paid ones — on Artificial Analysis's 0–100 capability score, an independent lab that tests every major model the same way. The overall test scores differ by six points; quality, safety and speed may still differ task by task. fact the scores · est that six points is small
9 → 52
on that same 0–100 score, the best model that fits on one $2,139 Mac mini went from 9 to 52 in two years. The machines stay; the software inside them keeps improving at no license cost, though every upgrade still needs testing and staff time. fact the two scores
01

The short version

One page, for the meeting where there are only five minutes

UVU's instinct is right: at $25 per seat per month, covering 48,669 students runs $14.6 million a year — $16.4 million with all 6,123 employees added (54,792 × $300) est — four to five percent of the university's entire Education & General budget. Nobody should sign that.

But the real choice is not "expensive cloud or nothing." Three facts, all verified this week, change the picture:

  1. Tools that add no license cost already exist. UVU's own guidance pages list Copilot Chat as included for students and staff under the current Microsoft license, and Google's Gemini for Education base tier costs education institutions nothing. Whether Gemini is actually switched on for UVU's account is something UVU's Digital Transformation office must confirm. fact license inclusion · tenant scope
  2. The biggest signed education deals run ~$19–30 per user per year. Cal State pays ~$19.20/user·yr across 675,000 people; Colorado ~$20; Maine ~$23 per billed FTE. Against those, the $25/month fear is off by 10–16× — though mid-size schools have been quoted far more, which is why week one requests UVU's own numbers. fact
  3. Owning the floor is now cheap. The best free-to-download models score six points behind the best paid ones on the 0–100 capability score explained above, run on a $2,139 Mac mini, and improve at no new license cost — each upgrade still earns a qualification pass. Faculty/staff serving hardware is roughly $64k upfront (~$80k over three years with power and care); operating staff and adoption support are the recurring costs, and they're priced in the open throughout this guide. fact est
The strategy is three layers: keep the free cloud floor, own a local fleet for daily work and data-sensitive teaching, and rent frontier models only for the research that genuinely needs them.
LAYER 0
$0

The no-added-license-cost starting option — already theirs

Copilot Chat appears to be included in UVU's Microsoft license, and Gemini for Education's base tier costs education institutions nothing; UVU must confirm whether Gemini is switched on for its account. Together they can put a competent assistant in front of most people at no added license cost, whether or not this plan proceeds. This is the no-added-license-cost starting option: it removes the reason to wait, and nothing built later depends on it. fact the license terms · unknown whether Gemini is enabled

LAYER 1
OWN

Basic on-campus service — a Mac fleet UVU owns

Mac minis and Studios running free-to-download models, behind UVU's existing campus login. Today's pick is a model with 27 billion learned settings — written “27B” — about the largest that fits comfortably on one Mac mini. Work on this route stays on campus; if the campus service is unavailable, a request for sensitive work stops rather than being sent outside — the written rules that decide which data must stay on campus are in §08. There is no charge per person. Model upgrades are free downloads, and each one must pass testing before students see it. It is built as a self-contained service — its own front door, its own operations, its own budget — so it stands whether or not any existing UVU system is available to it. If Digital Transformation wants it, the fleet can also connect to UVU's existing artificial-intelligence entry point: a connection, not a dependency. est

LAYER 2
RENT

The pay-as-used research pool — for work the campus machines cannot do

Intense mathematics, physics and deep engineering get access to the strongest paid models (Claude, GPT‑5.6, Gemini) through a fund with a spending limit on it: roughly $22,000 a year covers 200 heavy researchers, allowing each about 20 million pieces of text. Programs that cost nothing — Anthropic scientist seats, research credit grants of up to $50,000, and the US National AI Research Resource — are claimed first, and this line is only what is left to buy. est

The one deadline that matters

Apple replaced its entire Mac line on August 25. The new machines ship September 22; nobody has production benchmarks yet. The move: order four minis as a technical validation — not a service launch — to arrive with the first wave. Two weeks of hard measurement — how fast it replies, how many people it serves at once, whether one person's work can reach another's, and whether it stays stable over days — while quotes and shared-resource applications run in parallel, and every later step passes its gate before people touch it. fact dates · est plan

The words this plan uses, in plain English

Every specialist term on this site is defined here, and each is glossed again where it first appears. If a word on any page is not in this list and not explained beside it, that is a fault in the writing — tell us.

  • Conversation at once (elsewhere “stream”) — one person waiting for one answer. Capacity is counted in these, never in people: 36,629 people do not all type at the same moment. The capacity rule.
  • Model — the program that writes the answers. Free-to-download models can be run on your own machines; paid-only models exist only as a rented service.
  • Model score — the Artificial Analysis Intelligence Index: an independent lab runs every major model through the same reasoning, mathematics and coding tests and combines the results into one 0–100 number. Higher is better.
  • Parameters, or learned settings — how big a model is. “27B” means 27 billion of them. Bigger usually answers better and needs more memory.
  • The strongest paid models (elsewhere “the frontier”) — the best commercial services of the moment: Claude, GPT‑5.6, Gemini.
  • Router — the piece of software that decides which model should answer a given request, and whether that request is allowed to leave campus.
  • Entry point, or gateway — UVU's existing front door for artificial-intelligence services, run by Digital Transformation. This plan can connect to it and does not need it.
  • FTE, full-time equivalent — one person's full working year. “1.25 FTE” means one and a quarter full-time posts, however many people that is.
  • MFA, multi-factor authentication — signing in with a password and a second proof, such as a phone prompt.
  • Privacy-separation tests (elsewhere the “isolation canary battery”) — planting a unique secret in one person's session and then trying, from another, to reach it.
  • Approved outside destinations (elsewhere “egress allowlist”) — the short, written list of places the service is permitted to send anything at all.
  • Dx — UVU's Digital Transformation office. SCET — the Scott M. Smith College of Engineering and Technology. HERFP — Utah's Higher Education Research Fund Program. IRB — the university committee that approves research involving people. CISO — the university's chief information security officer. RACI — a chart naming, for each task, who does it, who is accountable, who is consulted and who is told.
  • FERPA — the US federal law protecting student records. ADA — the Americans with Disabilities Act. WCAG — the international standard for accessible web content. USHE — the Utah System of Higher Education.
02

Why this fall

Three clocks are striking at once

Downloaded models have become useful enough for daily campus work. Two years ago, running a useful model on a desktop was a hobby. Today the best free-to-download models — Kimi K3 and GLM‑5.3 — score 60 on the Artificial Analysis index against 66 for the best paid-only model, and the lag behind the strongest paid models has fallen from about twelve months to three or four. The strongest paid models still do some work better. fact The chart below is the whole hardware argument: the score of the best model that fits in a 64GB Mac mini, measured on the same index.

closed frontier today = 66 9 17 30 52 Sep 2024 Sep 2025 Apr 2026 Aug 2026 AA Intelligence Index v4.1.1
Best model runnable in 64GB, re-scored on today's index version so the points are comparable: Qwen2.5‑72B (9), Qwen3‑Next‑80B (17), Gemma 4 31B (30), Qwen3.8‑27B (52). Trend ≈ +20.6 points a year. If that trend continues, a mini bought this month would reach today's best paid-model quality by fall 2027 at no license cost. It is a forecast, not a guarantee, and every upgrade still needs testing and staff time. fact scores · est projection

The hardware just reset. Apple announced new Mac minis (M6 / M5 Pro) and Mac Studios (M5 Max / M5 Ultra) on August 25, first arrivals September 22, with a 512GB flagship following in late October. Buying last-generation hardware right now would be the one unforced error — the previous models are no longer sold new, and the survey in §04 and Appendix O found only single refurbished units and reseller backorders; timing the validation to the first shipment wave costs nothing extra. fact

Utah is funding exactly this. The 2026 Legislature put $15M one-time into a shared higher-ed AI research data center (administered under the University of Utah's budget item), created a research fund that reserves 15–25% for institutions like UVU, and UVU's own research office lists a $5M "Utah AI Moonshot" and a $45M Higher Education Research Fund with full proposals due October 6. A new president took office August 10. The October 6 round is available only if UVU filed the required notice of intent by August 28; if it did not, the next round is the target. fact the dates · unknown whether the notice was filed

03

How to read the scores in this guide

Whenever this guide gives a model a number like 52, 60, or 66, it is the Artificial Analysis Intelligence Index: an independent lab runs every major model through the same set of reasoning, math, and coding tests and combines the results into one 0–100 score, updated as models are released. "Open" means the model is free to download and run on your own machine; "closed" means it exists only as a paid service you rent. A higher score is better; a six-point gap is small — roughly the difference between this year's and last year's best. fact source · est "small" is our reading

What cloud AI really costs

The fear, the verified reality, and the trap that remains
Verified education pricing, checked live 2026‑09‑02
OptionPer user / yearAll-campus (54,792)Evidence
The fear: list-price seats at $25/mo$300$16.4Mhypothetical est
Microsoft 365 Copilot (academic list)$216$11.8MMicrosoft fact
Google AI Pro for Education (list)$240$13.2MGoogle fact
Cal State × OpenAI — renewal~$19.20$1.05MCSU's own $1.60/user·mo figure fact
U. Colorado × OpenAI — signed~$20~$1.1M$2M ÷ 100k fact
U. Maine System × OpenAI — signed$22.64/billed FTE~$1.24M$1.4M / 2yr ÷ 30,912 FTE pricing basis fact
Mid-size reality check: UC Davis / U. Iowa internal rates$144–156$7.9–8.5Mpublished campus rates fact
Copilot Chat (UVU's A-license) & Gemini for Education$0$0included / free tier fact
Hosted open models via API (gpt‑oss‑120b class)~$0.20–0.75~$10–40klive price tables fact usage est
Deal rates are other institutions' negotiated outcomes at their scale — benchmarks, not UVU quotes. The $19-30 rates came from 100,000-675,000-seat systems; mid-size schools publish $144-156. Anthropic and OpenAI education pricing is negotiated per campus; neither publishes a list price. unknown Week-one action in §09: request binding UVU quotes at four scopes so this table becomes like-for-like.

So cloud is not unaffordable. The honest problem is different:

Cloud seats launch fast and scale on demand — but the bill returns every year, and the record shows universities pay for seats people never use. Cal State had activated at least 250,000 of its 500,000 allocated seats in year one, and 0.7% of students finished the training. fact

Owning trades that recurring bill for operating work: hardware is mostly one-time money (the kind universities actually have), capacity grows as free model upgrades land, conversations stay on campus under UVU's own policies, and the machines double as teaching objects for the AI programs UVU already runs — but somebody has to run them, and that staffing cost appears beside every hardware number in this guide. Rent where elasticity matters. Own where usage is daily and data is sensitive. Decide with UVU's own quotes and telemetry, not anyone's brochure — including this one.

04

The local floor: what to buy and what it does

Measured speeds, honest capacity, one-time money

The serving unit is unglamorous and right-sized: the new Mac mini M5 Pro with 48 GB of shared memory, a 512 GB disk and 10-gigabit networking — $2,399 retail, $2,229 education — likely purchasable on Utah's existing cooperative contract, with Procurement confirming the route (§09). The plan prices the 10-gigabit machine because its own rack kit assumes 10-gigabit switching; the same mini without it is $2,299 retail / $2,139 education and ships three weeks sooner. Published tests on this exact chip class measured the fast-lane model family reading prompts at 2,100–3,700 tokens per second (a token is roughly a word piece) and writing replies at 97–105. The higher-quality 27B model runs slower per word, so the router sends each job to the lane built for it (§11). Every number here gets re-proven on the delivered minis before any purchase beyond the four-unit validation. fact prices & published measurements · est service planning

Measured by Utlyze on its own machine — not quoted from someone's blog

Utlyze ran the exact software and model this plan proposes, on a different and larger machine than the one the plan would buy: our own 512 GB-class Mac Studio, deliberately while it carried other work, so these are conservative floors. (“t/s” is tokens a second, and a token is roughly a word piece, so 27 t/s is comfortably faster than reading speed.) One conversation: 27 t/s. Four at once: 4 × 10.7 t/s. At eight, the queue behaved as designed — everyone still got 10.7 t/s, in two waves. A 4,563-token document was read in at 389 t/s. The privacy-separation tests — a unique planted secret per user, probes that try to reach another user's, and follow-up questioning — found zero leaks. These tests run before every future update. The four-computer test must confirm all of this on the planned Mac minis, which are smaller machines. fact — measured 2026‑09‑02 by Utlyze; method in the appendix pack

The three serving tiers, as this plan would configure them (Apple education prices, displayed 2026‑09‑03)
MachineMemory / storagePrice (education)Model it servesConversations at once, mixed workRole
Mac mini M5 Pro48 GB / 512 GB, 10-gigabit$2,229Qwen3.8‑27B class (score 52, handles pictures, Apache‑2.0)4The workhorse. Buy many.
Mac Studio M5 Max128 GB / 512 GB$4,639120B-class (Mistral Small 4 / gpt‑oss‑120b)4Heavyweight node for Engineering & Science
Mac Studio M5 Ultra256 GB / 1 TB$9,869GLM‑5.3‑Flash (320B, 18B active; score 57; MIT) — handles pictures and video4Later, if demand earns it.
Model mirrorone node moved from 512 GB to 2 TB+$720holds every approved model on disk, 715–721 GBSo a rebuilt machine copies over the campus network instead of pulling 720 GB from the internet again
Mac Studio M5 Ultra — research node512 GB / 1 TBprice unknownGLM‑5.3 (score 60) — the strongest model that runs with no cloud at all2, in a research lane of its ownConditional. Apple showed it on September 3 as “coming late October”, with no price.
Every price is the exact total Apple's own configurator displayed for that exact build on 2026‑09‑03 fact. Storage sizes are this plan's arithmetic, not Apple's advice: a serving machine has to hold macOS, its models, retained logs, the serving software's automatic cache (a tenth of whatever disk is fitted) and a 20–30% free-space reserve, which puts a mini at about 119–172 GB, a 128 GB Studio at 207–275 GB and a 256 GB Ultra at 370–465 GB est. Speeds elsewhere on this page are same-chip measurements; the new chassis ship September 22 and have no public benchmarks yet — the four-computer test produces them. What “conversations at once” means.

Two definitions this plan uses on every page

1. What a computer can carry — the capacity rule. One machine does not have one capacity; it has a different one for each kind of work. This plan therefore publishes one planning figure and the table behind it, and never mixes them:

Conversations at once, per serving computer, at comfortable speed
Kind of workMac mini128 GB StudioWhy
Plain chat and tutoring88Limited by how fast the machine writes the answer
Coding help48A whole repository is far more to read than a chat message
Questions about a document it has not read24Limited by how fast it can read the file; 4 per mini once indexed
A coding agent working on its own12One agent occupies a machine by itself
Research writing24Uses a model tested to hold the stated answer quality
Mixed-work planning limit44One average across the mixture above. This, and only this, is what the budget model buys against, and what every “conversations at once” figure on this site means unless it says otherwise.

2. What a fleet is made of — the fleet definition. A machine count on this site always breaks down into five named parts, and never into anything else:

  • Active minis — Mac minis assigned to a college and carrying conversations.
  • Studios — 128 GB Mac Studios, one per heavy college, carrying the larger everyday model.
  • Ultra — at most one 256 GB Mac Studio Ultra, shared campus-wide through the router.
  • Spares — about one machine per seven, held unconfigured and switched off until another fails.
  • Shared equipment — the rack kit (one per six machines: shelf, 10-gigabit switching, battery backup) and one node's storage upgraded to hold the model mirror. These are money, not extra computers.

So “25 computers” at the $120,000 scenario means 19 active minis + 2 Studios + 1 Ultra + 3 spare minis, with the rack kit and the model mirror priced beside them. A DGX Spark, where one is bought, is a teaching computer and is counted in the cost and never in the capacity. est capacity figures; fact the prices behind the equipment. Detail: Appendix R.

The newest lineup beside the generation it replaced

A fair question from the first read-through: the plan is priced on Apple's newest machines, so what did the generation before them cost, what does it cost today, and how much faster should the new one really be? Here is the answer, priced both ways, with the grades. The short version: the previous machines are no longer sold new; single refurbished units exist at prices near or above what they cost new; and the new ones should carry about 10% more conversations per box (40% more for the Ultra) — credits we set well below Apple's headline claims. fact prices, availability, published measurements · est speed credits and like-for-like deltas

The same three serving roles, priced on both generations (survey of September 3, 2026)
RolePrevious generationWhat it cost then · what it costs todayNewest lineup (announced Aug 25, 2026)Price now (retail · education)Price change, like for likeExpected speedOur call
Workhorse mini, 48GB Mac mini M4 Pro · memory speed 273 GB/s fact Family launched Oct 2024 at $1,399 ($1,299 education) fact; this configuration's launch price unknown. Today: not sold new; Apple Refurbished $1,949 with 10GbE, listed one unit at a time fact Mac mini M5 Pro · 307 GB/s fact; ships Sept 22 $2,299 · $2,139 fact; with 10GbE like the refurbished unit, $2,399 · $2,229 est New (education, 10GbE) vs refurbished: +$280, about +14% est +10% conversations per box est — memory speed +12%; same-build tests +11% reading, +21% writing fact Buy new. Supported supply, warranty, contract price, and a modest speed gain. Old units only against a written quote for the exact quantity.
Heavyweight Studio, 128GB Mac Studio M4 Max · 546 GB/s fact Launched Mar 2025 at $3,699 fact. Today: Apple Refurbished $4,159, single units fact — more than it cost new Mac Studio M5 Max · 614 GB/s fact; ships Sept 22 $5,099 · ~$4,639 est (512GB drive; $5,399 at 1TB fact) New (education) vs refurbished: +$480, about +12% est +10% per box est — memory speed +12%; same-build tests +12% reading, +25% writing fact Buy new for planned nodes; refurbished only if the exact quantity and drive size are confirmed.
Ultra, 256GB Mac Studio M3 Ultra · 819 GB/s fact Top chip launched Mar 2025 at about $7,099 est. Today: Apple Refurbished $8,149 (lower 28/60 chip, 2TB), single units fact; CDW-G $11,640, backordered fact Mac Studio M5 Ultra (top 36/80 chip) · 1,200 GB/s fact; ships Sept 22; 512GB version late October ~$10,799 · ~$9,900 est — 256GB needs the top chip; $11,299 at 2TB fact; education price not yet posted unknown New vs refurbished at 2TB: about +39% est +40% per box est — memory speed +47% fact Buy new only when models need more than 128GB, and get an institutional quote first. The 512GB version: wait for its price and a delivered-unit test.
Previous-generation prices are Apple Certified Refurbished listings seen September 3, 2026: one price, no education discount, one unit per listing — whether a matched set of 4 or 26 exists is unknown until reserved. Apple's own store and education store no longer list the previous models; CDW-G and SHI show them backordered. Speed credits are deliberately below Apple's headline claims ("up to 4× faster prompt processing"); they come from memory-speed gains and same-build tests, and the delivered units re-prove them before any purchase past the four-unit validation. Full tables, measurements, power ratings, and sources: Appendix O.
The fleet priced both ways — the workhorse mini, like for like (10GbE on both)
FleetPrevious, refurbishedNewest, educationNewest, retailWriting speed carried, previous → newest
4-mini validation$7,796$8,916 (+14%)$9,596 (+23%)480 → 528 tokens/s (+10%)
26-mini faculty & staff fleet$50,674$57,954 (+14%)$62,374 (+23%)3,120 → 3,432 tokens/s (+10%)
Hardware per unit of capacity (15 tokens/s)$244$253 (+4%)$273 (+12%)
Previous column is a list-price calculation at $1,949 a unit — it assumes the refurbished units could be found in quantity, which is unproven. Writing speed uses the measured 120 tokens/s per previous-generation mini (eight conversations at once) and the +10% planning credit for the new one. fact refurbished price and measurement · est everything else

What it means. At education pricing the newest fleet costs about 14% more than refurbished previous-generation units and should carry about 10% more — roughly 4% more per conversation carried — and it is the only one UVU can order in quantity, with a warranty, on contract. That is why the plan is priced on the newest lineup, and why we call buying the old generation the one unforced error. One caution for Facilities: Apple's maximum power ratings rose (mini 140 → 155 W; Studio 145 or 270 → 480 W). Those are electrical ceilings for circuit planning, not expected draw — the one measured AI load on the previous mini was 46 W. The budget simulator now has a generation switch and the configurator prices every configuration both ways, so no comparison is hidden. est

What each machine can actually run — the matrix, not one number

A fair challenge from the read-through: why not put the most intelligent open models on 512GB Studios, and did we test every model against every machine, capacity included, with two models on one box? We had not. Appendix Q now does it — ten open models × six machines, each cell with memory fit, speed, conversations at once, cost per conversation, and an intelligence-per-dollar index. The short version is below; scores are the same 0–100 index explained in §02. fact measured or published · est derived, basis in Appendix Q

Conversations at once, at a comfortable ≥10 words-per-second pace, by model and machine — planning estimates for delivered hardware
Model (score)Memory neededMini 48GB · $2,139Mini 64GB · $2,499Studio 128GB · ~$4,639Ultra 256GB · ~$9,900Ultra 512GB · price not posted
Qwen3.5-9B (22) — small fast lane7 GB88888
Qwen3.8-27B (52) — the daily driver18 GB8 (1 at 20+)8 (1 at 20+)888
Qwen3.6-35B-A3B (32) — sparse fast lane21 GB88888
gpt-oss-120B (24) · Mistral Small 4 (20)64–69 GB8 (6–7 at 20+)88
Qwen3.8-Flash-Next (56) — license on hold107 GB3 (memory-limited)88
GLM-5.3-Flash (57) — multimodal, MIT182 GB8 (4 at 20+)8 (4 at 20+)
GLM-5.3 (60) — the open ceiling431 GB2 (none at 20+); 18 words/s each
Kimi K3 (60) · DeepSeek V4 Pro (53)850–930 GBDo not fit any single Mac — frontier pool only
"8" means at least eight under the measured batch tests; no claim above eight is made. Each count reserves a 32,000-token conversation memory per seat plus a 15% operating reserve for macOS. The plan's engine keeps four per box as the blended admission cap for mixed campus traffic — conservative for chat, fair for coding, optimistic for cold documents and agents (§06) — until the delivered units are measured. est throughout; measured anchors and the method in Appendix Q

Same money, both ways

Five 48 GB minis cost about the same as the priced 256 GB Ultra. Five minis carry about 40 conversations of a score-52 model; one 512 GB node carries two conversations of a score-60 model. Multiplying quality by conversations gives 2,080 against 120 — that is a rough Utlyze comparison, not an industry measure. The same-money comparison against the 512 GB machine cannot be completed at all until Apple posts its price. The eight points matter for advanced coding (Terminal-Bench 88.2 vs 73.0; DeepSWE 66.9 vs 42.2) and complex professional documents (a 217-point lead on a work-products test); for tutoring chat, summaries, and ordinary drafts no matched test shows a difference. So the campus default stays minis and the frontier lane is the metered cloud pool — until a 512GB node earns its place. fact task results · est capacity

When a 512GB node earns its place — decided in advance

Path 1, recommended: a conditional shared research node. Kept in the plan as an unpriced option behind the router; bought only when the hardware pool reaches about $100,000, it takes no more than 15% of hardware spend, Apple has posted the price (late October), and a delivered-unit test passes. Path 2: a department or grant buys one once the price is posted (earlier access, low utilization, no redundancy). Path 3: no 512GB node — the same money buys minis, 128GB nodes, and cloud. What breaks first on the big node: reading long documents (a 131,000-token prompt took about 23 minutes on the previous generation) and a single point of failure. est

Two models on one machine — yes, and it is why 64GB exists

The serving software keeps several models loaded at once (llama-server's router mode, oMLX, LM Studio, vLLM-MLX), sharing the box's memory speed. A 48GB mini holds the 27B daily model plus a 9B fast model comfortably; the 64GB mini is the one that holds the 27B plus the sparse 35B fast model with production margin — that, not speed, is what its extra $360 buys. A 128GB Studio holds a 120B-class model plus the 27B; a 256GB Studio holds GLM-5.3-Flash plus the 27B. A 512GB node running GLM-5.3 at the preferred quantization has no room for a companion. Swapping from disk: the 27B in about three seconds, a 120B in eleven to thirteen, GLM-5.3 in about seventy-five — so small models are pinned and the big ones stay resident or stay off. fact software support and swap measurements · est pairings

What the mixed-work planning limit of four turns into

The same mini carries about 8 tutoring chats, 4 coding-help sessions, 2 sessions on a document it has not read (4 once that document is indexed), 1 coding agent, or 2 research-writing sessions at comfortable speed — a 22-fold spread in requests an hour hidden inside one number. Documents break first: a 30-page upload read cold by the 27B model takes one to three minutes before the first word, so uploads are indexed once and then served from the fast lane. A coding agent uses about 1,200 times the text of a chat, so it gets its own queue and a usage limit per person; public-class work may go to an approved cloud service under the written routing rule, never automatically or silently. The budget model buys against four, the mixed-work planning limit — the full table is in the two definitions above and in Appendix R. est

~$12,700

Four-computer test: equipment only. Four minis, the shared rack shelf, 10-gigabit switching, battery backup, and one node's storage upgraded to hold the model mirror. Staff time to run the validation (about a quarter to half of one full-time person) is real money and is priced visibly in the simulator — never hidden in a hardware number. est

~$66,900

Faculty and staff: equipment only, sized for exam week. 26 minis ($57,954) plus shared kit ($8,250) and the model mirror ($720) — $66,924 to buy, and about $81,600 over three years once AppleCare and power are added: about $13 per employee over three years, for the machines alone. That fleet covers the exam-week stress case (~104 simultaneous conversations); an ordinary day needs a fraction of it. Operating staff (evidence band: one to one-and-a-half full-time people) is the larger cost and is shown beside it, always. est

~$250,000+

Everyone: the broader program estimate. Students included, heavyweight machines in Engineering and Science, the research pool and the full adoption program funded, and headroom sized for exam week. This figure includes staff and training, unlike the two above. The simulator explores every point between — including the staffing line. est

Want to try a different mix of machines? The hardware configurator starts from the computers instead of the dollars: pick how many of each, and it shows people served on an ordinary day and in exam week, the model sizes it can run, cost to buy and to keep for three years, staff, and power — and prices that same configuration on the previous generation beside the newest.

Reality check kept in view: peak-hour demand for a 5,000-person population models at ~7 simultaneous conversations (base case); for the whole campus, ~65–70. Demand, not headcount, sizes the fleet — which is why the pilot's telemetry matters more than any projection in this guide. est

05

The budget simulator

Because the honest answer to "what does it cost" is a dial, not a number

There is no single right budget. There is a curve: every dollar buys coverage in a particular school, and every cut has a name. We built the curve into an instrument. Type any budget — or set each school's own budget, since money inside a school stays inside the school — and it solves the per-school supply-and-demand equations live: who is covered, where the first bottleneck forms, what falls below the line, and what the same coverage would cost to rent forever.

Interactive

Open the budget simulator

Slide the one-time equipment and launch money from a $12,700 four-computer test to a quarter-million-dollar campus build; the yearly staff cost and the three-year total are shown beside it throughout. Toggle students in or out, stress-test exam-week peaks, override per-school budgets, and print the scenario you like as a one-page sheet.

Launch the simulator →  Or plan on the campus map →

The map view puts the same engine on real overhead imagery of the Orem campus: each college is a circle sized by its modeled demand, colored by what the budget covers — click to include or exclude colleges and watch the money re-flow.

Prefer paper? Budget scenarios at a glance freezes the simulator onto one printable sheet: the coverage curve and four worked budgets.

06

Who will actually use it

Estimated UVU demand, built from other campuses' measured results — not vendor optimism

These are estimates of UVU demand, not UVU measurements: no UVU usage data exists yet, and the pilot's own figures will replace every number here. They are anchored to the best published whole-campus results in the country — a full-campus service at Cedarville (82% of eligible people activated, 75% used it in a given week, 49 messages per active person per week), Virginia Tech's controlled pilot (78% weekly, median 11 messages a week), and 2026 national surveys (32% of students and 25% of faculty now use AI daily). The busiest hour of the day carries about 8% of that day's traffic. fact the other campuses' figures · est everything applied to UVU

Projected base-case demand by school — students + staff eligible (est)
SchoolPeople (phase 1)Peak streams (base)Why it skews
Engineering & Technology (SCET)8,36514.2CS/engineering: 62% regular use nationally; Burns's own college
Woodbury Business6,80810.6Business: 51% regular use; case-work heavy
Humanities & Social Sciences6,6698.5Big population, moderate skew; writing support
School of Education5,8877.5Teacher-prep AI literacy mandate rising statewide
College of Science3,4795.4Math/science tutoring + research pool users
Health & Public Service3,0373.7Strong privacy constraints favor local serving
School of the Arts2,3842.9Lowest text-AI use; vision models change this
Phase-1 total36,62952.7generated by the same engine the simulator runs
People = 30,506 degree-seeking students (Fall 2023 college shares on Fall 2025 headcount) + 6,123 employees apportioned by share. The 18,163 concurrent-enrollment high-schoolers are excluded from phase one on purpose — minors, different consent rules; including them later raises campus base demand to roughly 65–70 streams. Skews triangulated across Berkeley/SERU (n=95,513), Lumina-Gallup 2026 (business 70% / tech 68% / humanities 43% weekly-plus), College Board faculty 2026 (71–88% by field), St. Olaf, Middlebury, and Anthropic usage data. UVU's own degree production confirms the ranking: business (2,345 awards/yr) and computer/information sciences (2,048) are the largest specific fields (IPEDS).

The planning insight: for the phase-one population, base-case peak load is about 53 simultaneous conversations — roughly 65–70 if concurrent-enrollment students join later. The financial risk is not under-buying — it is over-buying before measuring. est

The growth story, from the federal database — with a twist that matters. UVU's headline growth is real (32,670 → 48,669 since 2010; Utah's largest and fastest-growing big institution). But the government data shows 76–87% of the last decade's net growth came from concurrent-enrollment high-schoolers; the phase-one population — degree-seeking students plus staff — grew only ~0.6–1.3% a year. So year-three capacity planning is driven by adoption deepening, not headcount: the sourced band is 1.3–2.4× year-one demand, base case 1.9×. Utah is also more resilient than the national "enrollment cliff" (high-school graduates projected −5.6% by 2041 vs −10.3% nationally). Design integration and facilities for the 2.4× envelope; buy compute on the measured triggers. fact IPEDS/USHE/WICHE · est band

We stress-tested this model until it broke, and it broke in instructive places. The base rate is anchored to measured whole-campus telemetry, not surveys: Cedarville's full-year deployment works out to 2.15 requests per eligible person per day — 1.43 peak streams per 1,000 people against our 1.44 — and Berkeley's course bot and Michigan's token volumes land inside the same envelope. Two cases demand far more capacity: a 60-second reasoning reply doubles the streams needed (response time is the single biggest lever), and one 300-student class starting an assignment at 2pm needs ~143 slots to keep waits near a minute. So the design shows the queue, spills only public-class overflow to approved cloud routes, and carries numeric reorder triggers: buy six more minis when the weekly 95th-percentile load — the level exceeded only 5% of the time — holds above 70% of capacity for two weeks; open the overflow valve when any 15-minute wait passes 30 seconds. Automated/agent traffic gets its own queue class entirely — the one European campus that published full telemetry saw 99% of volume come from scripts, not chats. est model · fact anchors

One more honest correction, from the engineering review (Appendix R): the model's fixed 30-second service time is right for chat and wrong for a finals-week mix of chat, coding, documents, and research writing — the weighted service time is about 51 seconds, so the same arrivals occupy about 93 conversations instead of 54, a 70% understatement. That sits inside the 9× stress envelope the fleet is already sized for, so the numbers stand; the reason is now stated plainly. What fits inside no envelope is coding agents admitted freely: one agent turn uses about 580,000 prompt tokens, roughly 1,200 chats' worth, and if 5% of users started agents in the peak hour a 26-mini fleet would run at 133% of capacity with 16-minute waits. Agents therefore run in their own queue with a metered allowance and overflow to the frontier pool before utilization reaches 85%. est · fact agent telemetry (GitHub production trace, 3.2 million users)

07

Getting professors to actually use it

The hardest problem — and the one UVU is already closest to solving

Machines are the easy half. The national record is brutal about the hard half: Cal State bought 500,000 seats and had activated at least 250,000 by spring 2026; voluntary training completion was 0.7% of students and 16% of faculty — and activation is not the same as use. UK government trials left ~20% of Copilot licenses untouched. Seats do not create usage. fact

What does? The programs that worked share five moves, and the evidence is consistent:

  • Pay for a finished artifact, not attendance. CSU Bakersfield's $500-on-completion program hit 86% completion with a waitlist. Hawai'i paid 50 faculty $1,000 each to redesign an assignment. Small money, real deliverables. fact
  • Gate access with a short training, then support weekly. Virginia Tech's trained pilot ran 78% weekly-active. Manchester reported 90% adoption in 30 days with cohort support. fact
  • Teach inside the discipline. Faculty ignore generic prompt classes; they show up for "AI for your anatomy course." (Ithaka S+R, 19 institutions.) fact
  • Design adjunct-first. In fall 2025 UVU counted 1,310 part-time and 802 full-time instructional faculty fact. That is 62% part-time est. It counts people, not full-time jobs. Any program built around salaried faculty misses most of the people who teach. Asynchronous modules, evening clinics, paid completion. UVU Common Data Set 2025–26, §I-1 (fall 2025); the narrower count is in Appendix J.
  • Address integrity fears head-on. In every faculty survey we found, academic dishonesty outranks job loss as the top concern by a wide margin — but no single survey pins one share, so this guide no longer quotes one. The integrity evidence itself is calmer than the fear: no clean post-2022 jump in proven cases where institutions published totals, and AI detectors too unstable and uneven to prove anything (Appendix V). Lead with assessment redesign and authorship policy, never with "AI is inevitable." fact survey ranking · unknown any single share

Which departments first — evidence, not guesswork

A dedicated receptivity study of all seven colleges (courses taught, named faculty champions, policies published, grants won, Academy participation) lands on a clear sequence: wave one: Engineering & Technology (AI degrees, an AI lab in the new Smith building, the NVIDIA partnership, published department AI protocols) and Woodbury Business (AI courses in the 2026-27 catalog, the Institute's director on its faculty, three ready-made syllabus stances). Wave two: Education (a quiet powerhouse — AI in its accreditation mission, three faculty on UVU's UNESCO AI-and-education committee, statewide teacher-prep relevance) and Humanities & Social Sciences (largest writing load, biggest assessment strain). Science follows with research and lab pilots; Health & Public Service gets a separate safety lane (synthetic cases before anything clinical); Arts gets a rights-aware multimodal pilot once consent and licensing controls exist. fact signals · est sequence

The behavioral engine — engineered adoption, honestly

The adoption program is built on two published playbooks — Influencer's six sources of influence and Nudge's choice architecture — and then audited against thirty cases, from hospital hand-washing campaigns to Canvas rollouts, where mass adoption was actually achieved cheaply (Appendix N). That audit sharpened it into four vital behaviors: (1) run one real, bounded task, check the result, and apply, edit, or reject it within seven days — a reasoned rejection counts; (2) before the first graded assignment, publish an assignment-level AI rule with a disclosure example and a checking requirement, keeping a no-AI route wherever AI is not the learning outcome; (3) reuse and improve the workflow at its next natural occurrence within thirty days — grading, course copying, recommendation letters — rather than on an artificial weekly cue; (4) each peer catalyst personally helps three to five colleagues and shares one useful result and one rejected one. The main moment is not the first week of term, which is too late and too busy, but course-shell creation, four to six weeks before, where a removable Canvas block offers an active choice — permitted, limited, or prohibited — one example, and one support link. Defaults are engineered but honest — visibility is opt-out, participation is always opt-in; social-norm messages run only when true, current, and drawn from groups of twenty or more; catalysts are chosen by confidential peer nomination using the two questions the evidence favors ("who spreads useful teaching practices?" and "whom do you consult when a tool may not be ready?") — including respected skeptics and adjuncts — not by enthusiasm. Success is measured as verified adopters — applied use within seven days, one course rule published, a repeat within thirty — never as activations, and never in a way that could reach an evaluation file. The first wave costs about $20,000 (ten catalysts paid $1,000 each for a reusable artifact plus two clinics; twenty $500 artifact awards for adjuncts and faculty) and is expected to produce 110–150 verified adopters at roughly $133–$182 each; it scales to a $60,000 second wave only if cost per adopter, adjunct parity, and trust all hold. No honest plan promises a majority of instructors in one term at these budgets — reaching the 1,057 that a majority requires would take roughly $140,000–$190,000 directly, less if peer diffusion is real, which the first wave measures. est design · fact book frameworks & cited cases

The people already exist. Before any nomination round, the public record already names faculty in every one of the seven colleges who have built something with guardrails and said so out loud — a revision coach that refuses to draft, virtual patients for nursing, a marketing "boss" that argues back. The roster, with sources, is Appendix J. Peer nomination starts from real people, not hope. fact

Two safeguards this plan would require before launch — assessment assurance, and learning outcomes

Assessment assurance (Appendix V). One standard replaces the three templates: every material assessment carries an assignment-level Secure or Open label, the allowed and prohibited AI functions, a disclosure rule, and an individual proof point where the learning outcome demands it — set at the course-shell moment, with a report of unlabeled high-stakes tasks. Secure checks stay sparse: a few oral, written, coding, or practical proofs where UVU must show individual skill; everything else open, applied, staged, and explicit about allowed help — blanket redesign would burden honest students, adjuncts, and online learners most. Detector output alone is never evidence and cannot by itself cause a grade change, referral, interview, or sanction; a fair follow-up protocol gives the student the rule and the evidence, admits drafts and assistive-tool explanations, asks questions aligned to the original task, gives written reasons, and preserves appeal. Each paid catalyst's deliverable becomes one redesigned signature assessment, run and measured for workload and student trust — a code defense in Engineering, a changed-fact decision task in Business, staged research plus close reading in Humanities, microteaching in Education, a lab defense in Science, a synthetic practical in Health, a provenance-rich portfolio with live critique in Arts. Success counts substantiated cases, overturns, validity, workload, access, and trust — never detector flags. fact detector evidence · est standard

Learning outcomes (Appendix W). Access to a chatbot is not an intervention: the strongest trials that helped used course-aligned material, structured practice, feedback, and active student work; the strongest warning trial found unrestricted AI raised assisted practice scores and lowered the next unaided exam, and refusing premature answers removed that harm without adding learning by itself. No credible study shows a campus-wide deployment moved retention, completion, or DFW rates — so "verified adopters" stay an adoption measure, never a learning measure. The router becomes purpose-first (learn, produce, check, decide); learning-first tutoring — attempts, progressive hints, self-explanation, retrieval, transfer — is the course-practice default, with a labeled productivity mode where producing the product is the outcome and an attempt gate removable without shame or disclosure. A pre-registered, section-randomized pilot with delayed rollout, masked grading, and delayed unaided tests on problems the tutor never rehearsed is funded as part of the service. The contract: proceed to a larger trial only with no credible harm on delayed unaided learning; expand after the first credible year only if the delayed-learning estimate is positive with its lower 95% bound above −0.10 SD, transfer is not negative, DFW is no more than 2 points worse, and no subgroup shows repeated harm near 0.20 SD; reshape if assisted scores rise while delayed or transfer scores do not; stop a mode if its upper confidence bound sits below no effect. No adoption target overrides these rules. fact trial findings · est thresholds

UVU's head start

This is the part most universities would envy: the Kahlert Institute already runs an AI Academy (120 completers), curriculum grants, twice-monthly Power Hours, and a confidential one-on-one faculty coaching pilot. Board-reported employee AI use rose from 61% to 76% in a year. The plan here is not a new program — it is scaling what exists, paying for reusable artifacts and peer help rather than attendance, and measuring behavior instead of sign-ups. The 120 Academy completers are the candidate pool for nominations, not an automatic champion list. fact

The coach bot — phase two, built only when the support log proves the need

The idea: a UVU-tuned assistant inside the approved environment that asks a professor's discipline and course and walks them through a ten-minute build of something they will use this week, then hands hard cases to the Institute's human coaches. The audit's verdict was blunt and we accepted it: no precedent shows a bot alone moves faculty adoption unknown, and making it a launch dependency would turn a cheap behavior program into a software project. So the launch ships three faculty-reviewed templates, the existing Power Hours, and human help beside the task; the bot is built only when support logs show a repeated problem that templates and people cannot solve cheaply, with a clear test when it does ship — beat the self-service path on thirty-day repeat use, or be retired. Under the telemetry charter, its prompts are shielded from supervisors and barred from evaluations, with the charter's narrow legal exceptions stated in it, not hidden. The templates ship on day one regardless: allowed-use levels, disclosure and citation language, and an explicit rule that detector output alone is never misconduct evidence. est

Budget shape: adoption money is its own budget, never carved from hardware numbers. A one-college test costs $20–40k (30 faculty × $500 + two paid department leads); a cross-campus program runs $150–250k. At scale budgets the simulator holds an adoption line (default 18%) — because the evidence says the machines fail without it. est

The rights that make faculty willing — written in, not implied

Everything here is voluntary and faculty-governed: instructors decide whether and how AI enters their courses; students always keep an equivalent non-AI path with no penalty. Required adjunct work is paid at a stated hourly equivalent with no effect on rehire, and the intellectual-property terms for paid artifacts are handed to participants before enrollment. Usage telemetry lives under a published charter: aggregate reporting only, small groups suppressed, and a hard rule that no metric touches evaluation, tenure, discipline, or rehire. If pilot results are headed for publication, the IRB determination comes first. These are launch gates, and Faculty Senate, department chairs, adjunct representatives, and student government sit inside the governance from day one. est — commitments the plan binds itself to

08

How it's built

Boring, auditable parts — and one engineering problem taken dead seriously

Every component is open, inspectable, and swappable; the stack fits UVU's existing systems like it was designed for them — because it was:

Front door

LibreChat (free, open-source, MIT license) supports UVU's existing Microsoft login (Entra) with two-step verification, per-user workspaces, and per-department budgets — capabilities of the product; wiring them into UVU's tenant is a named validation-phase task with Dx. Students never see an API key (the technical password that connects to a model). Two conditions from the reviews: its accessibility must be proven on the exact pinned build with real assistive technology, and UVU's existing Wilson assistant may be the better student front door — the fleet can serve behind Wilson first, and no competing student front door launches until UVU decides Wilson's role (Appendices X, Y). fact product · est UVU wiring

Traffic control

LiteLLM routes each request to the right machine and the right model, enforces quotas, fails over automatically within the routing table's data-class rules, and exposes live capacity — the "how busy are we" gauge that feeds the simulator's real-world twin. fact product · est config

Serving

llama.cpp pinned builds on each Mac — the most battle-tested engine on Apple hardware — with popular models held resident (no cold-start roulette) and cold models swapped on demand. fact

Fleet care

Enrolled in UVU's existing Jamf Pro (already mandatory for university Macs), monitored with the same dashboards Digital Transformation already runs, updated in staged rings with tested, timed rollback paths for the app, the models, and the database. fact Jamf · est ops design

The one hard problem: isolation

Each of the major Mac serving engines has had at least one documented cache bug where text from one user could surface for another (llama.cpp #27148 — open; MLX‑LM #965 — fixed; Ollama's MLX path #17599 — open; LiteLLM response cache #29955 — open). This system therefore ships with every shared cache disabled, one active generation per user, and a standing "canary" test — synthetic users with unique markers hammer the system across concurrency, restarts, retries, and failover before every update, and any cross-user leak fails the release. The battery's first production run already happened (§04). Privacy here is an engineering discipline with a regression test, not a promise. fact named issues · est controls

Where each kind of data may go — the routing rule the router enforces
UVU data class (official scheme)Local Mac routeLabeled cloud routesOn overflow / outage
Public — general coursework, published materialYes (default)Yes — destination shown before sendMay spill to free layer or approved cloud
Sensitive — student information, most institutional dataOnly routeNeverFails closed — waits or declines, never reroutes
Restricted — the most protected categoriesExcluded from the pilot entirelyNeverNot applicable — this system does not accept it until Dx approves a design for it
Class names follow UVU's official data-classification scheme (Restricted / Sensitive / Public — student information is Sensitive). This resolves the tension between "local privacy" and "cloud burst" explicitly: bursting is a convenience for Public-class traffic only, restricted to a short written list of approved outside destinations, and what happens during an outage is part of the test battery rather than an assumption. est — design commitment mapped to fact class scheme

Built in after the engineering, safety, and law reviews

Lanes by kind of work

Short chat and coding turns ride the fast lane; uploads are indexed once and answered by retrieval (2,000–6,000 relevant tokens per question, never the whole file); exact course prefixes are cached and a student's follow-ups stay on the box holding that cache; a large prompt the machine has not seen before goes to the fast, lighter model rather than the slow, heavy one. Coding agents get their own queue, a usage limit per person, and are kept on the machine holding their cache (“session affinity”); public-class work may go to the paid research pool under the written routing rule before a machine reaches 85% of its capacity, never automatically or silently (Appendix R). est

Crisis handoff, before broad student use

A separate safety detector — rules plus a tested classifier, never keywords alone; an honest message ("I'm an AI tutor, not a crisis service — if anyone is in immediate danger call 911 or UVU Police; for a suicide or emotional crisis call or text 988"); the smallest useful packet to an acknowledged human queue, not an email; a restricted case record; graduated handling for low-confidence signals; no discipline on classifier output alone. The trained person, not the model, decides Title IX, Clery, abuse-reporting, and police questions (Appendix S). est design · fact duties

The first pilot leaves the dangerous features out

No tools, no web fetching, no shared document retrieval, no automatic cloud spill; outbound network denied from inference and parsing workers; uploads quarantined — allowed formats only, real type verified, size capped, active content disarmed, parsed in a no-egress sandbox. Each feature returns only behind its own control and test gate. est

Labelled, accessible, on the record

"UVU AI assistant — not a human" shown before and throughout every chat, the safe harbor under Utah's AI disclosure law; counseling, clinical, legal, and financial advice kept out of the general assistant. Accessibility (WCAG 2.1 AA) is a launch gate tested with assistive technology on the real chat, uploads, streaming answers, and generated documents — the federal deadline for UVU is April 26, 2027. A records map before launch says which chats are education records, which retained data are open-records material, and what retention applies (Appendix T). fact law · est design

A finished control plane — the settings and small databases that run the service

Redundant routers that hold no state of their own with tested failover; identity outage fails closed for new sessions with a status page outside the login path; approved model weights mirrored on every compatible node plus a checksum-verified mirror; a same-class spare or a documented degraded route for the one heavy node; backups of configuration, prompts, manifests, and required state with a proven restore; recovery targets — a failed mini in five minutes, public-data fallback to the metered route in thirty, the restricted route fails closed and rebuilds in four hours; a change freeze from seven days before finals through the grade deadline (Appendix S). est

Model admission record

Every model enters through a record: repository and immutable revision, download source, cryptographic hashes, exact license and notices, model card, conversion recipe, scan and test receipts, named approvers, a read-only local copy and an independent mirror. Licenses are pinned by exact checkpoint, not by family — the daily driver and both 128GB workhorses are Apache-2.0; Qwen3.8-Flash-Next is on hold under its community license; GLM-5.3 and Kimi K3 sit behind counsel review (Appendix T). est

One trust contract, everywhere

The same source states, uncertainty states, refusals, report route, and human route across Wilson, LibreChat, and every routed model. Every sourceable claim opens the exact supporting passage; every answer shows whether it came from approved UVU material, an uploaded file, the open web, or model memory; an evidence-sufficiency gate lets the assistant say "not found" or "sources conflict" instead of filling the gap; no raw confidence percentages until UVU validates them per model, task, and course. Errors become repair cases with a receipt, an owner, a status, and a visible correction notice. Tutoring is student-first and hint-first, enforced outside the model prompt; calculators and code tests are used wherever the task permits. Results are published by task and route, never as one blended "accuracy" score (Appendix X). est design · fact calibration studies

Accessible text chat first; the rest behind its own gate

The front door is a conditional candidate until the exact pinned build passes an assistive-technology audit (VoiceOver, NVDA, JAWS, switch access, magnification) on complete conversations; streaming is treated as a visual effect — short states and complete messages are announced, never every token; focus stays under the student's control. Uploads, math, diagrams, PDF export, and voice each pass a separate gate later. Interface language, content language, and reply language are separate student choices, and every deployed model passes a paired UVU English/Spanish benchmark before Spanish is promised. Paid disabled and bilingual student testers under an IRB determination; a funded accessibility owner; a release-specific conformance report and a barrier-response route; phone, weak-Wi-Fi, and interrupted-upload tests; generated documents as accessible HTML first (Appendix Y). fact standards and dates · est priorities

Logging, precisely: chat history lives in the user's own workspace on UVU-controlled systems (that's what makes conversations resumable), with user deletion and a published retention rule — stated exactly: "Chat content logging is off by default. UVU retains only the data listed in its records notice, except when a user saves a conversation or a scoped legal or records hold requires preservation." A saved chat tied to a student is a FERPA education record; retained chats and metadata are records under Utah's open-records law — classified and, where required, produced, never promised "confidential" (Appendix T). Operational analytics see metadata only — tokens, latency, model, department — never prompt text; capacity planning and department accounting run entirely on those counters. Security-event logs required by Policy 447 capture system events, not conversation content, and the forensic procedure for a suspected leak is a written runbook with CISO sign-off. The full data-flow diagram, retention table, and runbooks are the first deliverables of the validation phase — drafted by us, owned and approved by Dx before any pilot user logs in. The control plane (the small databases and router state behind the curtain) is named, priced, backed up, and owned by the service's operator — with Dx holding approval, audit, and takeover rights — in that same package; no invisible single points of failure, and no borrowed Dx staff time. est

Stands on its own; fits the house: this service is designed to exist independently — its own hardware, front door, model operations, staffing, and budget — because everything UVU already runs (the employee AI Gateway launched in July, Ask Wilson in 44 courses, the shared state compute) is spoken for by the people using it today. Nothing here borrows their capacity. What it needs from UVU is small and specific: a room with power and a network drop, sign-on federation with the campus Entra tenant, a data-class ruling, and a sponsor. Three harness tiers ride the same router: a chat front door for everyone (LibreChat), a sandboxed coding agent for builders (OpenCode or Cline class), and a governed research search for scholars (Onyx class) — compared product by product in Appendix L, with the whole market of 171 harnesses — adoption, growth, funding, stability, and the dead-or-declining list — in Appendix M. The surfaces it plugs into are already mapped from public sources — identity, Canvas, portal, ticketing, security tooling, network — and the questions only UVU can answer are down to five (Appendix K). Where UVU wants the two worlds to meet, the connection points are ready — the fleet can serve as a local engine behind the Gateway, and course assistants like Wilson can call it — as optional integrations, never as dependencies. Dual sponsorship still fits the governance pattern UVU's documents describe: the college owns the academic mission; Dx owns identity, security review, and the Policy 445/447 package (data flows, classifications, runbooks), which arrives drafted for Dx and the CISO to own, change, and approve before any pilot user logs in. fact Gateway/Wilson · est arrangement

09

The path through the university

Procurement, approvals, and money — mapped to UVU's actual rules
  • No new bid expected. Apple minis and Studios are on NASPO ValuePoint contract MA 23003 / Utah addendum PA4282 — the competition requirement is already satisfied through June 2027; UVU Procurement confirms the SKUs and process (their call, made easy). fact
  • Utah's trial-use statute (§63G‑6a‑802.3) was built for this: minimum quantity, defined test measures, no obligation to scale. Last-published signature levels: dean/executive at $25k+, VP at $50k+ — a $12k validation sits comfortably low (current schedule to be confirmed; the live PDF is down). fact verify current schedule
  • Week-one, in parallel, at zero cost: request binding seat quotes at four scopes (500 / 2,112 instructors / 6,123 employees / all-campus) so the cloud comparison becomes like-for-like; apply for a Redtail early-access allocation (Utah's new 264-GPU statewide AI system — UVU faculty are eligible today) for research workloads; ask USHE whether a system-wide seat deal is in motion (Burns sits on the statewide AI task force); and stand up the governance charter (§07) with Faculty Senate at the table. fact eligibility · est plan
  • Reviews to plan for, not fear: academic tech committee (ATSC), IT Oversight, security (Policies 445/447), and accessibility as a tested release gate — the web accessibility standard (WCAG 2.1 AA) with disabled-user testing, a funded remediation owner, and an equally effective alternative for every required task (Policy 452). UVU's campus AI policy (draft 441) is still in the approval pipeline; the pilot charter follows the enacted policies and adopts 441 when it lands. Timeline estimate for a funded on-premise pilot with no student-record data: 8–16 weeks to users, owned by those reviewers, not promised by us. Campus-wide: 6–18 months, phased. est
  • Two gates added by our own gap-hunt: a one-page Facilities + Network + CISO signoff (named room, measured power/cooling headroom — 26 minis draw ~4kW and ~1.15 cooling tons — and the 10GbE build-to-order option specced, +$100/box) before any purchase order; and counsel-approved interim service terms (acceptable use, a "report this response" route, an incident taxonomy that carries UVU's existing 24-hour reporting duties into the service, and a narrow rule that the assistant never makes official decisions — it explains, cites, and hands off) before the first user. est
  • One scheduling courtesy: Dx cuts the campus portal over to Pathify on November 16 — the pilot goes live before or after that window, not during it. fact
  • Three rules that arrived in 2024–2026 (Appendix T): Utah's AI disclosure law — the service labels itself "UVU AI assistant — not a human" before and throughout every chat, the statute's safe harbor, and keeps counseling, clinical, legal, and financial advice out of the general assistant; since May 6, 2026 a state-funded university may ask Utah's Office of AI Policy for a joint interpretation of the one open question (is a free campus assistant a "consumer transaction"?). The federal ADA Title II web rule — WCAG 2.1 AA for the real chat, uploads, streaming answers, and generated documents by April 26, 2027, tested with assistive technology, not a page scanner. Records — a records map before launch: which chats are FERPA education records, which retained data are open-records material, which retention series applies (advising records: five years after separation), and how a legal hold stops deletion. fact law and dates · est design
  • Ride the state's programs; count none as funding yet: the Utah Board of Higher Education's statewide AI direction (December 2025) and its AI Task Force (May 2026 — Dr. Burns is a named member); the free statewide AI Workforce Credential (available since July 1, 2026) goes into faculty and student onboarding instead of a competing certificate; the 2026 budget's $15 million shared AI research data center for Utah's public universities and $3 million ongoing AI workforce accelerator — UVU's share is unknown until an allocation exists; the NSF State and Regional AI Infrastructure Hubs solicitation (deadline November 4, 2026; $4–12 million over five years; one award per state) as a Utah consortium, not a solo request. fact programs and dates · unknown UVU's share
  • Who pays, and who runs it (Appendix U): $0 from general student fees — Utah's fee rule bars them for instruction, academic support, and administration; central IT funds the shared base (fleet, identity, security, monitoring, refresh reserve, baseline research pool, the accountable owner); colleges and grants buy marginal nodes and reserved capacity; showback first, chargeback only for reserved or unusually heavy use; a college joins only after a signed three-year commitment covers its hardware and half its marginal staff time. Staffing: a professional owner with a paid student operations team beneath — more coverage at the smallest band, about a quarter off staffing cost at the larger bands (the simulator's staffing switch shows both). Buy the first fleet outright unless a written education-lease quote beats ownership; plan a year-three review and a year-four refresh, with three-year resale values near half of purchase price. est · fact fee rule, wage anchors, financing terms
  • Buy in gated waves, never 26 at once (Appendix Z): four minis after the delivered-unit test (load the model; 1k/8k/32k-token tests; four streams at the 10-token-per-second floor; batch 1/2/4/8, cold-load, peak-memory, and 24-hour soak; identity, logging, isolation, accessibility, and recovery checks); up to eight more only after half the pilot returns at their next natural task within 30 days, measured demand uses at least 60% of capacity for four teaching weeks, and support stays inside the staffed plan; up to the planned 26 after spring results and a refreshed local-versus-cloud comparison. Waiting for a rumored next chip means about 15 months — 60 validation device-months and 390 fleet device-months forgone — while a machine 25% faster would repay no more than 9.6 months of waiting inside a 48-month life; so delay only for a failed must-have, an official Apple announcement within 90 days that clears a stated bar (the required memory, or 35% more measured throughput for no more than 20% more cost), a dated grant award within 60 days, or a blocking accessibility or security result. The 512 GB line stays unpriced until Apple posts it and two approved research workloads need it. A dated 24-month watchlist — Apple events, model checkpoints, grant deadlines, state milestones, the April 26, 2027 accessibility deadline — closes the appendix. fact cadence and arithmetic · est gates
  • The full engineering design lives in its own document: the architecture dossier — component contracts and a substitution ledger, the data-class routing with per-contingency behavior, the model-provenance and release-integrity rules, 24 pre-decided contingencies ("if the CISO rejects a component… if quotes come back cheap… if one department eats 80% of usage…"), five ready decision memos, and the eight weekly tripwires with owners. est
  • Money paths beyond the college budget: the $5M Utah AI Moonshot and $45M HERFP — full proposals due Oct 6, but only for teams whose notice-of-intent was filed by Aug 28, so the first question for Burns is whether UVU filed one (if not, target the next cycle); the state's $15M shared AI-compute appropriation; NSF NAIRR compute; the Kahlert relationship for a named lab; and the free research programs in §11. fact

The 14-week sequence — every step behind its gate

Weeks 0–2: dual sponsorship (college + Dx), governance charter with Faculty Senate, order 4 minis on the co-op contract for the Sept 22 shipment wave — accepted only after one delivered unit passes the acceptance test (Appendix Z) — and fire the free parallel actions: four-scope seat quotes, Redtail application, USHE system-deal question, IRB determination request. Weeks 2–6: technical validation. Stand up the stack on the real hardware, benchmark the exact models, run the isolation battery and a soak, wire up Microsoft Entra sign-on — UVU's own campus login, with multi-factor authentication — and deliver the written security package (data flows, how long anything is kept, written procedures for operators, and a chart naming who does each task and who is accountable for it) for Dx and CISO approval. Weeks 6–14 (opens only when those gates pass): a voluntary instructor cohort — proportionally adjunct, paid, IP terms in hand — plus an opt-in student cohort in SCET, telemetry under the charter. Week 14: the build/rent/hybrid decision made with UVU's own curves and the returned quotes in the same spreadsheet — with a Moonshot/HERFP application already in flight. est

10

The risks, weighted and named

A failure-modes analysis — because a plan that hides its risks isn't one
FMEA — likelihood × impact × detection difficulty (each 1–5; higher = worse) (est)
RiskLIDScoreSafeguard required before launch
A crisis disclosure lands in a tutoring chat and no trained person sees it35345Separate safety detector, honest crisis message (988 / 911 / UVU Police), an acknowledged human queue rather than an email, restricted case record, graduated handling of low-confidence signals — built before broad student use (Appendix S)
Assisted scores rise while unaided learning falls — the strongest warning trial's exact pattern35345Learning-first tutoring by default (attempts, hints, self-explanation); delayed unaided tests on unrehearsed problems in every pilot course; the learning-outcomes contract's expand / reshape / stop rules; adoption never overrides them (Appendix W)
The front door fails the assistive-technology audit, or Spanish quality is worse and nobody notices34336Front door is a conditional candidate; exact-build audit with real assistive technology; accessible text chat first; paired English/Spanish benchmark per deployed model; paid disabled and bilingual testers; an alternative front door held ready (Appendix Y)
Protected data pasted into the wrong tier — in one 2026 clinical study, 23.5% of health students who used AI on placement admitted disclosing a patient identifier; education has the same shape with minors (HB 273), business with licensed case content34336The local Sensitive lane exists precisely so protected work has a safe home; department onboarding states permitted use per course (the appendices map all 51 units); Public-cloud tiers carry paste-guards and no-PHI defaults; placement work follows the partner's rules
Adoption stalls; hardware idles44232Adoption fund with a $20k first wave (catalysts paid for a reusable artifact plus peer help; artifact awards for adjuncts and faculty; clinics); a "verified adopter" measure — applied use within seven days, repeat within thirty — instead of activations; scale gates on cost per adopter, adjunct parity, and trust; the coach bot deferred until support logs show a repeated need
Cross-user privacy leak25330Caches off, per-user isolation, canary battery on every release, content-free logging
A minor's data or a patient record enters the service before the rules are settled25330Concurrent-enrollment students stay out until eleven named gates pass — parental permission alone is not one; placement patient data fails closed; synthetic clinical cases only; the FERPA-or-HIPAA status of UVU's own health records resolved with counsel (Appendix S)
First-mover risk: no public campus Apple-serving precedent found33327Pilot-gated scale, hosted lane running beside it, published exit thresholds, fallback documented (§12)
New M5 hardware underperforms same-chip projections23424Nothing beyond 4 boxes ordered until 2-week validation on real hardware
Coding agents admitted freely swamp the shared queue — one agent turn ≈ 580,000 prompt tokens; 5% of users starting agents in a peak hour puts a 26-mini fleet at 133% of capacity34224Agents run in their own queue with a metered allowance, session affinity, and overflow to the frontier pool before 85% utilization; chat is protected from long work (§06, Appendix R)
Long documents read cold take minutes to the first word; users conclude "it's slow"43224Uploads indexed once and served by retrieval from the fast lane; exact course prefixes cached; follow-ups stay on the box holding the cache; cold large prompts never hit the dense model (Appendix R)
Accessibility deadline missed — the ADA Title II web rule requires WCAG 2.1 AA by April 26, 202734224Accessibility is a launch gate: assistive-technology testing of the real chat, uploads, streaming answers, and generated documents; tagged PDFs; a funded remediation owner (§09, Appendix T)
A student is accused on an AI-detector flag34224Binding rule: detector output alone is never evidence and cannot trigger a grade change, referral, interview, or sanction; fair follow-up protocol; Secure/Open labels set before term (Appendix V)
Harmful output or misuse incident in the early weeks33218Interim service terms before first user; "report this response" route; incident taxonomy with urgent lanes; narrow no-official-decisions rule; pre-written response memo
A model's license is misread — the family is not the checkpoint33218Licenses pinned by exact checkpoint in the model admission record; Qwen3.8-Flash-Next on hold under its community license; GLM-5.3 and Kimi K3 behind counsel; the daily driver and both 128GB workhorses are Apache-2.0 (Appendix T)
The one Ultra node, the router, or identity fails during finals33218Redundant stateless routers; identity outage fails closed with a status page outside the login path; a same-class spare or a documented degraded route for the heavy lane; recovery targets (a failed mini: 5 minutes; public-data fallback: 30 minutes; restricted route: 4 hours); a change freeze from seven days before finals (Appendix S)
Stress peaks swamp capacity (compound envelope ≈9× base)42216Visible queue + Public-class overflow valve + numeric reorder triggers (§06)
Ongoing ownership unfunded after year one34112Operator staffing priced inside the plan (~0.25–0.5 FTE pilot), no borrowed Dx time; Utlyze support offer; Dx holds takeover rights if it ever wants them
Open-model releases slow (policy/geopolitics)23212Today's models already exceed the use case; fleet still serves them; frontier pool unaffected
Approval sequencing slips32212Dual sponsorship, drafted security package at first review, bounded-validation framing
Software-stack churn (fast-moving projects)32212Pinned versions, staged rings, rollback manifests; every component swappable
Wave order reads as favoritism between colleges32212The scoring formula is published (Appendix A) and arguable; the planning map lets any college be included or excluded live; the free layer reaches every college on day one; waves 2–3 carry calendar dates, not "later"
A "free" tier the campus leans on changes terms, rotates its models, or trains on the traffic4218Free tiers are never on the critical path: sandboxes and Public-class burst only, behind the data-class rules; contributor/training tiers barred for UVU data; every outside route has a pre-qualified substitute (Appendix L)
Arts, music, and language programs need speech and media the core stack doesn't serve4218Named openly: the local stack is text + vision; routes R5/R6 attach separate approved speech and media services, rights charter first — no pretending one box does everything
Model licenses shift under a key model2228Daily driver is Apache-2.0; immutable approved checkpoint retained; two alternates pre-qualified

Said plainly: in every search we ran, no public example surfaced of a university running campus-wide AI on Apple hardware — the schools that self-host (Münster, UCSD, Florida, Indiana, Purdue) all run Linux/NVIDIA on central research computing. UVU would likely be first, which is exactly why the plan buys four boxes before it buys twenty-six, measures everything, keeps the hosted lane running beside it for honest comparison, and publishes exit thresholds up front. If the validation says the Macs lose, the same gateway, sign-on, governance, and adoption work carries straight onto whichever backend wins — nothing is wasted but the price of four excellent computers UVU keeps anyway. fact search result · est framing

11

The models

A portfolio, not a bet — different strengths for different work
Recommended portfolio, scored on live leaderboards 2026‑09‑02
ModelRoleScoreLicenseRuns on
Qwen3.8‑27BDaily driver — chat, tutoring, coding, vision52Apache‑2.0mini 48GB
Qwen3.6‑35B‑A3BFast lane — classrooms, batch, overflow, cold documents; pairs with the 27B on a 64GB mini32Apache‑2.0mini 48GB
Mistral Small 4 (119B)Heavyweight workhorse, stable (gpt‑oss‑120B, score 24, Apache‑2.0, is the alternate)20Apache‑2.0Studio 128GB
Qwen3.8‑Flash‑NextQuality ceiling for 128GB — on hold: only three conversations fit, and its license needs counsel (Appendix T)56Qwen Community License — separate license for hosted servicesStudio 128GB
GLM‑5.3‑Flash (320B/18B)The 256GB shared node — multimodal, premium coding lane, eight conversations at once57MITStudio 256GB+
GLM‑5.3 (753B/40B)The open ceiling — conditional research node only: two conversations at once at about 18 words/s (§04)60Custom GLM license — counsel reviewUltra 512GB (late Oct; price not posted)
Kimi K3 · DeepSeek V4 ProConsidered — do not fit any single Mac (850–930GB at 4-bit); frontier pool only. K2 Horizon (Sept 3) is a watch item60 · 53Custom · MITCloud only
Muse Spark 1.3 (Meta)Frontier-pool candidate — coding and agent route, standard no-training tier only; never the contributor tier61–62Closed — API only (weights "on the roadmap")Meta cloud, metered
Different subjects genuinely favor different models — which is the point of a router: tutoring, code, vision, and long documents each hit their best tool automatically. Non-Apache licenses get legal review before production, by exact checkpoint — "the Qwen3.8 family is Apache-2.0" was wrong; only the 27B checkpoint is. Scores from Artificial Analysis, September 3, 2026; capacity per machine in §04 and Appendix Q. Changing models needs no new hardware — but it still needs a license review, testing, setup and staff time.

For the mathematicians and physicists — Burns's hardest requirement — the frontier pool routes heavy research to the strongest current models: Claude and GPT‑5.6-class models now solve International Math Olympiad sets essentially perfectly, score ~94% on graduate-level science exams, and produced at least one passing referee grade on 7 of 10 unpublished research problems in a controlled trial this summer. Metered access for 200 heavy researchers models at ~$22k/year — and the free lanes come first: Anthropic's scientist program (free seats + up to $50k API credits per project), OpenAI's researcher seats, NSF NAIRR compute. fact scores & programs · est pool cost

12

The alternatives, priced honestly

The NVIDIA question answered straight — then every option on one page

Any hardware review will ask "why not NVIDIA?" Here is the honest comparison with NVIDIA's own campus-friendly box, the DGX Spark, using verified prices and published measurements:

Verified 2026‑09‑02 — price, memory, bandwidth
SystemPriceMemoryBandwidth$ / GB
NVIDIA DGX Spark (4TB)$4,699128GB273 GB/s$36.71
Mac mini M5 Pro 48GB$2,29948GB307 GB/s$47.90
Mac Studio M5 Max 128GB$5,099128GB614 GB/s$39.84
Mac Studio M5 Ultra 256GB$10,799256GB1,200 GB/s$42.18
Memory bandwidth is a large part of how fast a reply is written, though the processor, the software, the model's size and the number of people using it all matter too: the $2,299 mini moves memory faster than the $4,699 Spark; the roughly price-matched Studio moves it 2.25× faster. A prior matched test found the previous-generation Mac Studio 27–82% faster than Spark at generation, while Spark led at long-prompt ingestion.

Where Spark genuinely wins — and earns a place in this plan: it runs CUDA, the language of every datacenter GPU a UVU graduate will ever touch. For the ML courses, fine-tuning labs, and NVIDIA-certification work the Smith College already teaches (UVU has an active NVIDIA partnership), one or two Sparks as named lab machines are the right purchase — for teaching, not for serving. For serving inside a Jamf-managed campus, the published evidence favors the Macs on management fit, power, and single-user speed, while NVIDIA/Linux keeps the more mature high-concurrency serving stack — which is precisely what the validation phase measures head-to-head before any fleet money moves. fact capabilities · est verdict

Every option on one page — where each honestly wins

The decision matrix — seven ways to give UVU AI, priced on the same evidence — but note that they cover different time periods, so put a common three-year total beside them before comparing
OptionCost basis (population · horizon)Data controlOps burdenWins when
Local Mac floor (this plan's Layer 1)~$80k machines 3-yr + ~1–1.5 FTE/yr · 6,123 staff · stress-sizedBest — never leaves campusReal: fleet + model opsData rules bite, usage is daily, teaching value counts, staffing is funded
Hosted open-model APIs (zero-retention routes)~$2.5k/yr (5k users) to ~$23k/yr (45k) + gateway opsContractual only; data leavesLight-moderateVariable load, Public-class data, fastest start — runs as our comparison lane
Negotiated frontier seats (ChatGPT Edu class)$118–184k/yr · 6,123 staff if mega-deal rates; $880k+/yr at quoted mid-size ratesVendor contractLightestThe quote comes back near $19–30/user·yr and polish beats control
Utah shared resources (Redtail) + AI GatewayUnknown — quote requested week 1State-governedApplication + queue realitiesResearch bursts, training, classes — complement, not yet a daily-chat backend
DIY 128GB PCs (AMD Strix Halo class, ~$3.7k)~30% less hardware $ than Studios · same horizonSame as localHeavy — Linux/driver fleet outside Jamf128GB under $4k matters more than management fit and warranty
Free tiers (OpenRouter free models, AI Studio, OpenCode Zen "contributor" access to Meta Muse Spark 1.3)$0 · capped at 50–1,000 requests per day per account, ≈1.4% of phase-1 daily demand with the three most generous recurring offers combined under one institutional account; multiplying accounts is barred by the providers' termsWeakest — consumer terms; contributor tiers train on the trafficNone — and no admin plane, SLA, or continuitySandboxes, model evaluation, Public-class burst — never the floor (Appendix L)
NVIDIA DGX Spark teaching computer (1–2 units)$4.7k each · one-timeLocalMedium (separate OS world)Teaching NVIDIA's own software and training small models — always, in a named laboratory role
Sourced in the appendix pack (devil's-advocate and hardware-sweep reports, with the searches that failed as well as succeeded). The commitment behind this table: if UVU's returned quotes or the validation data favor a different mix, the recommendation moves — and the response is already pre-decided (a binding all-in quote at or below ~$30/user·yr that passes the security, accessibility, retention, and exit gates flips the faculty tier to seats, and the fleet contracts to a four-Mac data-sensitive enclave; the gateway, governance, and adoption work carry over unchanged). That's what makes this a decision aid instead of a sales document. est
13

About this work

Method, sources, and what Utlyze offers

How this was made. Twenty-seven independent research passes ran against the live web across September 2–3, 2026 — model leaderboards, Apple's store, signed contracts, Utah statutes, UVU's own policy manual, published campus telemetry — plus direct pulls from the federal IPEDS database (enrollment trajectories for UVU and eight Utah peers; UVU degree production by field). Every research thread is logged as its own tracked issue. Claims carry grades: fact verified at a primary source that day; est derived, with the math shown; unknown honestly unresolved. Then the work attacked itself, twice: an adversarial pass re-verified the eighteen most load-bearing numbers (ten confirmed, eight sharpened, zero broken); a paid-skeptic pass built the best case against this plan (its strongest points now live inside §09 and §12); three independent adversarial reviews — a blocking CIO, a skeptical CFO, a wary Faculty Senate — attacked the draft; a second audit round then graded the fixes and re-checked every number, catching an engine bug and several inconsistencies, which were corrected before publication. A fresh-eyes review of the published site on September 3 found more, and those corrections are in this version; the calculators and the written figures were reconciled against one another as part of it. The demand model was stress-tested to its breaking points (§06), and the proposed stack was physically benchmarked on our own hardware, isolation battery included (§04). Known remaining gaps are stated where they live, not hidden. Failed searches are logged, not papered over. This is how we believe AI-era consulting should be done — and the method is itself a demonstration of what UVU's own people can build with these tools.

From Utlyze

About this work, and an offer

Utlyze prepared this research and plan because we believe Utah's largest university deserves a first-rate answer to the most important infrastructure question of the decade — and because the person asking it is asking exactly the right questions.

If UVU wants help making it real — architecture, the isolation test battery, the coach bot, the pilot build, training, or ongoing operations — we would be honored to help, and we will work inside whatever budget the university actually has. If UVU builds it without us, this guide is still yours, and we will still cheer.

— The Utlyze team · utlyze.com

Key sources

Live leaderboards: Artificial Analysis, LMArena, Epoch AI lag analyses. Hardware: Apple newsroom (Aug 25, 2026), Apple education store configurations, NVIDIA DGX Spark documentation, llama.cpp/oMLX published benchmarks. Contracts & pricing: CSU–OpenAI signed agreement, CSU renewal reporting, U. Maine System, U. Colorado, Microsoft education pricing, Google education pricing. UVU & Utah: UVU institutional data, Kahlert Applied AI Institute, UVU Policies 445/447/452, NASPO Apple contract, Utah Code §63G‑6a, UVU HERFP / AI Moonshot notice. Adoption evidence: Virginia Tech pilot report, Cedarville telemetry, Tyton 2026, Ithaka S+R, CSU rollout reporting. Precedents: Münster uniGPT, UCSD TritonGPT, U‑M ITS reports. Teaching, trust, accessibility, and timing reviews (Appendices V–Z): randomized trials and meta-analyses on AI tutoring 2023–2026 (Harvard, Turkey, Nigeria, Khan Academy evaluations), the Challenge Success and ICAI integrity data, independent AI-detector tests, the AI Assessment Scale, W3C WCAG 2.1 and DOJ guidance, LibreChat's public issue tracker, Apple's release history and newsroom, the Artificial Analysis index history, Wikimedia Commons image licenses. Engineering, safety, law, and money reviews (Appendices Q–U): Artificial Analysis model pages and methodology, the oMLX and llama.cpp benchmark databases, MLX-LM quantization benchmarks, GitHub's production coding-agent telemetry (arXiv 2608.00101), Utah Code Titles 13, 53E, 53H, 63G and 78B, the DOJ ADA Title II rule and its 2026 interim rule, Department of Education FERPA/PPRA guidance, HHS HIPAA guidance, USHE and Governor's Office announcements, Apple Financial Services terms, USHE R516, the NSF 26-513 solicitation, EIA electricity data. The complete source log — several hundred dated citations with the searches that failed as well as succeeded — is available on request.