Read the five lines first. Every appendix opens with the same card: the question it answers, the answer, the numbers the answer turns on with their fact est unknown grades, what the plan does about it, and what is still unknown. Read those five lines, and open the tables under them only where you want to check the work.
College sequencing & receptivity
Appendix A in five lines
- QuestionWhich colleges should go first?
- AnswerSmith Engineering and Technology and Woodbury Business go first; Education and Humanities and Social Sciences follow in wave two.
- Deciding numbersscores of 12 for Smith, Woodbury, and Educationest10 for Humanities and Social Sciences and Health and Public Serviceest5 for Artsest
- What the plan doesStart Smith and Woodbury, add Education and Humanities and Social Sciences in wave two, then Science with a research lane; Health and Public Service runs on a separate safety lane, and Arts is a later pilot once rights controls exist.
- Still unknownNo public college-level use survey exists, so every score here reads visible behavior rather than measured demand — and a high score does not by itself buy an earlier wave.
No public UVU survey reports AI use by college unknown — so receptivity here measures visible behavior: AI courses in the 2026–27 catalog, named faculty leads, college policies, Kahlert curriculum grants, training completions, and applied projects. That evidence, college by college — cross-checked against what deans and faculty have said in public (Appendix J) — produces the scores below. The scoring rule is printed with the table so it can be argued with: priority = 2×receptivity + demand + story value − complexity est.
A score is not a place in the queue. The score says how ready a college looks. The wave says when the service can actually be built for it, and only two colleges can be stood up at once. Education scores level with Engineering and Business and still opens wave two; Health & Public Service scores above Science and still runs on a lane of its own, because its safety controls have to exist first. The sequence the plan follows is the one in §07 of the guide and on the campus map: wave one Engineering & Technology and Woodbury Business, wave two Education and Humanities & Social Sciences, Science following with the shared research service, Health & Public Service on the safety lane, Arts later once rights controls exist. est sequence
| Score rank | College | Receptivity | Demand | Complexity | Story | Score | Wave | Why, in one line |
|---|---|---|---|---|---|---|---|---|
| 1 | Smith Engineering & Technology | HIGH | 5 | 4 | 5 | 12 | 1 | Deepest AI curriculum (MS-AAI, AI certificate, IS Applied-AI emphasis), the new Smith building's AI lab, the NVIDIA agreement, named faculty leads. |
| 2 | Woodbury School of Business | HIGH | 5 | 3 | 4 | 12 | 1 | Highest national student use (70% weekly+); live AI courses in foresight, automation, strategy; the institute's director teaches here. |
| 3 | School of Education | HIGH | 5 | 4 | 5 | 12 | 2 | AI Academy completions, UNESCO AI-education committee seats, and Utah's K-12 AI rollout (HB 273 + statewide Gemini) make teacher prep time-sensitive. Starts with no live P-12 student data. |
| 4 | Humanities & Social Sciences | HIGH | 4 | 3 | 3 | 10 | 2 | The only college with published AI guiding principles and a faculty resource guide; AI courses in English, Communication, Political Science; a Kahlert grant. Largest writing load. |
| 5 | Health & Public Service | HIGH, controlled | 5 | 5 | 4 | 10 | separate safety lane | Raised from MED after the public-voice sweep (Appendix J): named nursing champions have already built virtual-patient bots; a $35M AI-equipped campus is funded. Clinical work cannot precede partner, privacy, and safety controls — simulation first. |
| 6 | College of Science | MED | 4 | 4 | 4 | 8 | follows wave 2 · shared research service | Real applied work (ML in CHEM 3300, the student-built DIFFRAX device) but concentrated in pockets; needs reproducibility and lab-safety design first. Its heaviest demand is the metered research pool. |
| 7 | School of the Arts | MED | 3 | 4 | 2 | 5 | later · rights-aware pilot | Multimodal value is real, but rights controls — likeness consent, provenance, licensing — must exist before a creative pilot. |
How to read "story": the Utah-facing narrative each college can carry. Education ranks first there — "UVU prepares Utah teachers to use AI safely before they enter classrooms" lands statewide while HB 273 and the state's Gemini rollout are live. Engineering's NVIDIA-agreement story and Woodbury's employer-evidence story follow. Sources: UVU 2026–27 catalog, Kahlert Institute pages, UNESCO Chair rosters, Ethics Awareness Week programs, college policy documents, Utah HB 273, the June 2026 Google–USBE announcement — full URL log in the working papers fact (behaviors) / est (scores).
Discipline usage — the national data
Appendix B in five lines
- QuestionWhich fields use artificial intelligence most often?
- AnswerBusiness, technology, and engineering students use it most often, but every field still needs support.
- Deciding numbersBusiness 70%factTechnology 68%factEngineering 65%fact
- What the plan doesThe program starts Engineering and Business, then adds Humanities and Social Sciences with writing and source-check work.
- Still unknownThe 94,060-person California State University survey publishes no open field table, so the two largest datasets are logged as checked but unusable.
These rates are not interchangeable — "use" ranges from daily coursework habit to once-a-semester. Read the measure column before comparing rows. The pattern that survives across all of them: business, technology, and engineering lead student frequency; faculty use is above 70% in every field; writing-heavy disciplines feel the most assessment strain.
| Study & population | What it measured | Rates by field | Grade |
|---|---|---|---|
| Lumina–Gallup, 3,801 U.S. students, Oct 2025 (pub. Mar 2026) | Daily + weekly AI use in coursework | Business 70% · Technology 68% · Engineering 65% · Vocational 58% · Healthcare 56% · Social sciences 53% · Natural sciences 49% · Humanities 43% | fact |
| College Board, 3,482 U.S. faculty, summer 2025 (pub. Feb 2026) | At least one AI use in the faculty role | Business & comms 88% · Health sciences 86% · CS 84% · Engineering 83% · Social sciences 81% · Math 74% · English 73% · Arts & music 71% | fact |
| Berkeley/SERU, 95,513 undergrads at 20 public research universities, spring 2024 | Regular use (monthly+) | Computer science 62% · Mathematics 53% · Business 51% · Arts 24% | fact |
| Middlebury, 634 students across 43 majors, Dec 2024–Feb 2025 | Any academic use during the semester | Natural sciences 91% · Social sciences 85% · Humanities 75% · Arts 73% · Languages 57% · Literature 49% | fact |
| Anthropic telemetry, 574,740 academic conversations, Apr 2025 | Share of conversations (intensity, not adoption) | CS 38.6% of conversations vs 5.4% of U.S. degrees — a 7× intensity ratio; business under-indexes (8.9% vs 18.6% of degrees) | fact |
| Michigan Engineering, n=182, Jun 2025 | Ever used GenAI | Engineering students 96% | fact |
| Grenoble (JMIR), n=388 health students, Apr 2026 | Used GenAI during a clinical placement | Pharmacy 64% · Nursing 54% · Medicine 48% · Physiotherapy 46% · Midwifery 26% | fact |
| Kogod Business School, 483 students over 3 waves, Aug 2026 | Used AI for academic work, prior 6 months | Business students >80%; 29% use it 11+ times weekly | fact |
What this means for the plan: capacity follows the frequency leaders — Engineering & Technology and Woodbury Business take wave 1 and its hardware — but service design cannot skip the "low" fields. Humanities' 43% weekly is still nearly half the college, and its faculty carry the heaviest redesign load. That is why wave 2 brings Humanities & Social Sciences in with citation-verification and writing-process tooling rather than raw capacity, and why how often a field uses AI does not by itself set its place in the queue (Appendix A). Sources linked by name; the two large gated datasets (CSU's 94,060-respondent survey, Digital Education Council) publish no open field tables unknown and are logged as checked-but-unusable.
What each college needs built
Appendix C in five lines
- QuestionWhat different artificial intelligence service does each college need?
- AnswerOne shared service needs different tools, controls, and success tests for each college.
- Deciding numbersseven different servicesestplanned target of zero incidents with personally identifiable informationestplanned target of zero incidents with protected health informationest
- What the plan doesChange the tools by college, start Health with made-up cases, and wait on Arts until consent and license rules exist.
- Still unknownNo live build or result is reported for any college service; these checks are NOT_RUN.
"What is in a school stays with that school" is a budget rule and a design rule. Money and use cases stay local; security, identity, and data rules stay central. This table is the design contract per college est, with the binding compliance anchors fact.
| College | Highest-value uses | Build differences | Hard controls | Success means |
|---|---|---|---|---|
| Engineering & Tech | Code explanation & debugging, cyber labs, technical-image help, simulation, project docs | Code execution sandboxes, model choice, technical vision, long context | Learning sandboxes separated from production; secret scanning; human review for aviation/safety-critical outputs | Better project rubrics, testable code, faster feedback, AI credential completions |
| Woodbury Business | Spreadsheet automation, SQL & financial modeling, case analysis, market research | Tables/charts, structured data, cited web research, document generation | Licensed case material stays out of general models (Harvard Business Publishing requires written permission for GenAI uploads); client-data audit trails | Verified model accuracy, employer-scored case decisions, honest disclosure |
| Education | Standards-aligned lesson plans, rubrics, differentiation, practicum rehearsal, licensure prep | Grounded retrieval, reading-level control, tutor mode, teacher-controlled student experiences | Utah HB 273: approved tools only, educator judgment preserved, no AI in high-stakes decisions; no student records or PII in unapproved tools (FERPA) | Candidates can explain safe vs unsafe use; less prep time; zero PII incidents |
| Humanities & Soc Sci | Writing feedback, citation verification, multilingual tutoring, qualitative coding, deepfake analysis | Source-linked retrieval, corpus search, citation checking, transcription | Preserve student voice and process evidence; never accept invented citations; CHSS's own published AI principles govern; no detector-based accusations | Citation accuracy up, documented writing process, better source evaluation |
| Science | Math/stats tutoring, notebooks, literature synthesis, spectrum/microscope image interpretation | Math + vision + reproducible notebooks, uncertainty reporting | No unsupervised lab-safety decisions; reproducible calculations; faculty review; synthetic or approved research data only | Concept mastery, reproducible outputs that survive faculty checking |
| Health & Public Service | Synthetic clinical cases, anatomy tutoring, licensure prep, documentation practice, emergency simulations | High-reliability grounded answers, strict refusals, full audit, human sign-off | Synthetic/de-identified first (HHS Safe Harbor); clinical partners' rules govern placements; separate high-assurance environment before any real clinical workflow | Licensure and simulation results up; every consequential output reviewed; zero PHI incidents |
| Arts | Ideation, previsualization, storyboards, variants, portfolio critique, production planning | Multimodal generation/editing, asset history, provenance metadata | Likeness and voice consent recorded; source licenses tracked; human authorship preserved (Copyright Office: prompts alone create no copyright) | Faster iteration at equal or better jury quality; portfolios still show the student's judgment |
Department by department
Appendix D in five lines
- QuestionWhat route and safety rule does each department need?
- AnswerAll 51 catalog units need access, but the model route and safety rule change with the work.
- Deciding numbers51 child unitsfact1.8× load for computer, information technology, and cyber workest0.8× text load for culinary and trades workest
- What the plan doesKeep budgets by college and set each department’s route, tools, and stop rules.
- Still unknownThe type of the Reserve Officers’ Training Corps unit and direct use rates for several fields are UNKNOWN.
The college is the budget unit; the department is where work actually differs. This appendix maps every unit in UVU's 2026–27 catalog — four colleges and three schools, 51 child units fact — to a demand tier, its highest-value uses, a model route, and the constraint that must be honored. Tiers rest on the strongest available evidence per field (a mechanical-engineering cohort study, a 91%-use mathematics sample, clinical-placement rates by health profession, the 38.6% computer-science share of academic AI conversations) and are graded est where triangulated. Notably: no family lands in a LOW tier — every department shows enough recurring task overlap to need access, guidance, and a course-level policy.
The ten routes
| Route | Meaning |
|---|---|
| R1 — local general | Fast local models for routine scale; the long-document local model for writing and multilingual work; the vision-capable local model when images help. |
| R2 — coding hybrid | Strong local code models first; frontier escalation for hard multi-file work. |
| R3 — quantitative/science hybrid | Local for tutoring and checking; frontier for hard proofs and advanced research. |
| R4 — controlled health/social | Approved local retrieval first; secure frontier escalation; no clinical or counseling decisions, ever. |
| R5 — multilingual/speech | Local text translation; a separate approved speech service for audio, transcription, pronunciation. |
| R6 — creative multimodal | Local analysis; approved frontier image/audio/video services with rights controls. |
| R7 — safety-critical technical | Source-grounded local support; frontier for review only; human authority always controls. |
| R8 — business analytics | Local routine analysis; frontier for complex models and current external facts. |
| R9 — writing/research | Local long-document model with retrieval; frontier for difficult synthesis. |
| R10 — education | Local lesson and material creation; multimodal or accessibility escalation when needed. |
Which models fit which family — and when frontier is genuinely needed
| Department family | Capabilities that matter | Local fit | Frontier trigger | Route |
|---|---|---|---|---|
| CS/IT/cyber | Code, logs, repositories, tools, long interaction | Qwen3.8 or GLM-5.3; Qwen3.6 for routine help | Multi-repository work, difficult debugging, advanced defensive security | HYBRID est |
| Mechanical/civil/electrical engineering | Math, code, diagrams, documents, CAD-adjacent vision | Qwen3.8 or GLM-5.3-Flash | Novel designs, difficult simulation, high-stakes review | HYBRID est |
| Math/statistics | Stepwise reasoning, code, symbolic checking | Qwen3.8/Qwen3.6 plus a calculator or symbolic system | Advanced proofs, exploratory research, persistent hard reasoning | HYBRID; GPT-5.6 preferred for the hardest tier est |
| Physics/chemistry/biology | Math, diagrams, literature, code, long papers | Qwen3.8 or GLM-5.3-Flash for learning and preprocessing | Novel synthesis, advanced research, complex chemistry or biology | HYBRID/FRONTIER est |
| Nursing/health | Retrieval, documents, case simulation, privacy | Approved local model with curated retrieval | Complex evidence review in an approved secure tenancy | CONTROLLED HYBRID est |
| Accounting/finance | Math, spreadsheets, tables, current documents | Qwen3.8/GLM-5.3 for routine analysis | Complex models, regulations, current markets, high-stakes review | HYBRID est |
| Marketing/management | Drafts, presentations, research, analysis | Qwen3.6, Mistral Small 4, or Qwen3.8 | Large multimodal campaigns or deep market synthesis | LOCAL-FIRST est |
| English/writing | Long documents, revision, voice, feedback | Mistral Small 4 or Qwen3.8 | Large-source synthesis or high-stakes publication | LOCAL-FIRST HYBRID est |
| History/philosophy | Long documents, source comparison, argument | Mistral Small 4/Qwen3.8 with retrieval | Large archives, difficult synthesis, current political facts | HYBRID est |
| Languages | Multilingual text, translation, speech | Mistral Small 4 for text; pilot actual UVU languages | Real-time speech, transcription, pronunciation, 70+ language translation | HYBRID; Gemini speech stack est |
| Psychology/sociology | Long documents, qualitative coding, statistics | Local secure model with retrieval | Sensitive-data research or complex literature synthesis | CONTROLLED HYBRID est |
| Education | Lesson creation, differentiation, accessibility, documents | Qwen3.6/Mistral/Qwen3.8 are sufficient for routine work | Large multimodal curriculum packages or speech access | LOCAL-FIRST est |
| Art/design/media | Vision, layout, images, video, iterative review | Qwen3.8 or GLM-5.3-Flash for analysis | Native generation, audio/video, polished multimodal artifacts | HYBRID; Gemini media services where approved est |
| Music/dance/theatre | Audio, video, scripts, translation, rehearsal | Local text models for scripts and planning | Speech, transcription, music/audio/video processing | HYBRID; separate media stack est |
| Aviation/public safety/trades | Manuals, current rules, scenarios, vision | Local retrieval system for instruction and search | Complex synthesis only; never autonomous operational control | LOCAL-FIRST HYBRID est |
Model names reflect the current local portfolio (the Qwen, Mistral, and GLM classes the guide's §11 details) and current frontier services. Benchmark caution carried from the working papers: providers publish their own numbers under different scaffolds — no clean cross-vendor ranking exists, and a declared million-token context is a ceiling, not a promise fact. The upgrade rule stands: models are the cartridge, not the console — routes persist while names change quarterly.
Relative compute load per family
For capacity planning: relative text-token load per active user against a mixed-campus average of 1.0×, with practical working-context ranges est. Vision and speech are flagged separately — token counts understate their compute.
| Family | Load | Working context | Why |
|---|---|---|---|
| CS/IT/cyber | 1.8× | 32K–128K | Repositories, logs, repeated debugging, tests, and many-turn sessions |
| Mechanical/civil/electrical engineering | 1.5× | 16K–64K | Multi-step problems, standards, code, diagrams, and reports |
| Math/statistics | 1.4× | 8K–32K | Prompts are often short, but reasoning and checking outputs are long |
| Physics/chemistry/biology | 1.5× | 32K–128K | Papers, lab context, equations, data, diagrams, and scientific reasoning |
| Nursing/health | 1.2× | 16K–64K | Cases, study material, documentation, and retrieved clinical guidance |
| Accounting/finance | 1.4× | 16K–64K | Tables, spreadsheet logic, regulations, explanations, and revisions |
| Marketing/management | 1.3× | 32K–64K | Research packets, campaigns, cases, plans, and presentation iteration |
| English/writing | 1.5× | 32K–128K | Full drafts, readings, multiple revisions, and detailed feedback |
| History/philosophy | 1.4× | 32K–128K | Long readings, primary sources, argument maps, and comparison |
| Languages | 1.2× text | 8K–32K | Many short translation and dialogue turns; speech cost is separate |
| Psychology/sociology | 1.3× | 32K–64K | Literature, qualitative records, survey design, and statistics |
| Education | 1.4× | 32K–64K | Lesson sets, differentiation, rubrics, examples, and feedback |
| Art/design/digital media | 1.2× text; about 1.8× compute-equivalent | 16K–64K | Image and video input make token counts a poor compute proxy |
| Music/dance/theatre | 1.1× text | 8K–32K | Scripts and plans are moderate; audio/video processing is separate |
| Aviation | 1.0× | 16K–64K | Episodic manual, regulation, scenario, and technical-writing work |
| Culinary/trades | 0.8× text | 8K–32K | Short instructional sessions with occasional image-heavy bursts |
| Public administration/public safety | 1.2× | 16K–64K | Policy documents, reports, scenarios, and current-source retrieval |
The master table — all 51 units
| UVU unit | College/school | Demand tier & basis | Primary uses | Route | Governing constraints |
|---|---|---|---|---|---|
| Allied Health | Health & Public Svc | HIGH est; health-program evidence, 2026 | Study aids, terminology, evidence search, case simulation, documentation practice | R4 | No patient identifiers; approved evidence; clinician/instructor signoff |
| Criminal Justice | Health & Public Svc | MED est; social-science and writing task fit | Case comparison, policy, reports, scenarios, statistics | R4/R9 | Current law; sensitive records; bias and civil-rights review |
| Emergency Services | Health & Public Svc | MED est; health plus safety-critical task fit | Scenario practice, protocols, reports, public education | R7 | No operational reliance; current approved protocols; no incident identifiers |
| Health Sciences | Health & Public Svc | HIGH est; 83.2% health-education use, 2026 | Literature, study support, data, public-health materials | R4 | Privacy, evidence quality, no diagnosis |
| Nursing | Health & Public Svc | HIGH est; 12 uses/student/semester, 2025 | Exam prep, cases, communication, documentation practice | R4 | No PHI; clinical judgment remains human |
| Physician Assistant | Health & Public Svc | HIGH est; clinical evidence base | Cases, literature, differential-reasoning practice, patient education drafts | R4 | No unsupervised diagnosis or treatment; source and faculty check |
| Public Administration Graduate Program | Health & Public Svc | MED est; writing/policy/business task fit | Policy memos, budgeting, public records, stakeholder analysis | R8/R9 | Current law and policy; records privacy; citation check |
| Public Health | Health & Public Svc | HIGH est; health, data, and communication task overlap | Epidemiology support, data, literature, outreach | R4/R3 | Population bias, privacy, current evidence, no clinical claims without review |
| Utah Fire and Rescue Academy | Health & Public Svc | MED est; vocational evidence, 2026 | Protocol search, scenarios, study aids, report practice | R7 | Approved manuals only; no live operational authority |
| Communication | Humanities & Soc Sci | HIGH est; business/communication faculty use 88% | Drafting, audience adaptation, media analysis, research, presentations | R1/R9 | Attribution, source verification, likeness and media rights |
| English | Humanities & Soc Sci | HIGH est; writing study, 2026 | Brainstorming, revision, feedback, close reading, rhetoric | R9 | Preserve student voice and process evidence; verify citations |
| History and Political Science | Humanities & Soc Sci | MED est; history direct evidence, 2025 | Primary-source comparison, timelines, argument, policy/current affairs | R9 | Primary-source fidelity, current facts, bias, provenance |
| Integrated Studies | Humanities & Soc Sci | MED est; mixed task composition | Cross-field synthesis, planning, writing, reflection | R9 | Apply each contributing course’s rules; avoid false cross-field certainty |
| Languages and Cultures | Humanities & Soc Sci | MED est; 54.2% regular ChatGPT use, 2025 | Translation, conversation, grammar, pronunciation, cultural comparison | R5 | Dialect/cultural review; audio privacy; voice consent |
| Marriage and Family Therapy Graduate Program | Humanities & Soc Sci | HIGH est; health plus psychology task fit | Role-play, theory review, documentation practice, literature | R4 | No client data; no therapy; supervisor and ethics review |
| Philosophy and Humanities | Humanities & Soc Sci | MED est; 54.5% philosophy-reading use, 2024 | Argument maps, counterarguments, reading support, writing feedback | R9 | Text fidelity, authorship, explicit uncertainty |
| Psychology and Counseling | Humanities & Soc Sci | MED est; 54% psychology use, 2026 | Concept tutoring, study design, statistics, literature, role-play | R4 | No client/participant data; no diagnosis or counseling |
| Social and Behavioral Sciences | Humanities & Soc Sci | MED est; social-science proxies; sociology direct rate unknown | Qualitative coding, surveys, theory, statistics, literature | R4/R9 | IRB and participant privacy; bias; source verification |
| Biology | Science | MED est; 50% small-sample use, 2026 | Literature, bioinformatics, diagrams, lab preparation, tutoring | R3 | Lab safety, dual-use review, source checking |
| Chemistry | Science | MED est; grouped science evidence | Calculations, spectra explanation, literature, lab preparation | R3 | Chemical safety and dual-use controls; verify units and methods |
| Earth Science | Science | MED est; science plus visual/data task fit | Maps, field notes, geospatial code, diagrams, literature | R3/R6 | Map/data provenance; field and hazard decisions remain human |
| Exercise Science and Outdoor Recreation | Science | MED est; health and science task fit | Anatomy, program-design practice, data, risk scenarios, public education | R4 | No personal health prescriptions; safety and privacy checks |
| Mathematical and Quantitative Reasoning | Science | HIGH est; 91% math use, 2026 | Stepwise tutoring, alternative explanations, practice, checking | R3 | Protect independent skill formation; verify every calculation |
| Mathematics | Science | HIGH est; same direct math evidence | Proof critique, modeling, statistics, code, research preparation | R3 | Formal validation; frontier required for hard research; AI is not proof |
| Physics | Science | MED est; n=804 physics evidence, 2025 | Problem explanations, simulation, lab analysis, literature, diagrams | R3 | Verify equations, assumptions, units, and experimental interpretation |
| Elementary Education | Education | HIGH est; preservice-teacher evidence, 2025 | Lesson plans, differentiation, examples, family communication | R10 | No minor data; developmental suitability; district and course rules |
| Graduate Education | Education | HIGH est; education output-creation share 74.4% | Curriculum, research, leadership, policy, assessment design | R10/R9 | Research citations, student data, and policy must be verified |
| Secondary and Special Education | Education | HIGH est; teacher-education direct evidence | Subject lessons, accommodations, rubrics, simulations | R10 | No IEP or minor identifiers; accessibility and disability-bias review |
| Student Leadership and Success Studies | Education | HIGH est; education and social-science task fit | Coaching practice, study aids, student communication, program materials | R10/R4 | No student records; do not replace human advising |
| Art and Design | Arts | MED est; design-student evidence, 2024 | Ideation, critique, visual analysis, layout, portfolios | R6 | Copyright, provenance, style appropriation, disclosure |
| Dance | Arts | MED est; direct rate unknown | Movement analysis, rehearsal planning, grants, promotion, video review | R6 | Performer consent, likeness, choreography rights |
| Music | Arts | MED est; direct rate unknown | Theory tutoring, transcription, rehearsal, program notes, promotion | R5/R6 | Composition and recording rights; voice/performer consent |
| Theatrical Arts for Stage and Screen | Arts | MED est; direct rate unknown | Scripts, dramaturgy, design, production planning, captions | R6/R9 | Script rights, performer likeness, disclosure and provenance |
| Applied Engineering and Transportation Technologies | Engineering & Tech | HIGH est; engineering plus vocational evidence | Diagnostics, manuals, calculations, reports, code, instruction | R7/R2 | Approved manuals; physical safety; no automated equipment control |
| Architecture and Engineering Design | Engineering & Tech | HIGH est; engineering and visual task fit | Drawings, code checks, CAD assistance, specifications, presentations | R6/R7 | Licensed professional review; building codes; design provenance |
| Aviation Science | Engineering & Tech | MED est; aviation course evidence, 2024 | Manuals, regulation study, scenarios, weather explanation, writing | R7 | Current FAA/approved material; never operational authority |
| Computer Science | Engineering & Tech | HIGH est; 38.6% Claude share, 2025 | Coding, debugging, architecture, tutoring, testing | R2 | Course-specific authorship rules; secrets and sandbox controls |
| Construction Technologies | Engineering & Tech | HIGH est; engineering/trades task fit | Estimating, plans, codes, scheduling, safety instruction | R7/R8 | Current code and manufacturer sources; field signoff |
| Culinary Arts Institute | Engineering & Tech | MED est; n=313 culinary evidence, 2024 | Recipe scaling, costing, substitutions, menu writing, instruction | R1/R6 | Food safety, allergy verification, cultural attribution |
| Digital Media | Engineering & Tech | HIGH est; code plus multimodal task fit | Image/video analysis, storyboards, prototypes, code, accessibility | R6/R2 | Copyright, likeness, training-data provenance, disclosure |
| Electrical and Computer Engineering | Engineering & Tech | HIGH est; Mines daily-user cluster | Circuits, embedded code, signals, diagrams, reports | R2/R3/R7 | Validate calculations and hardware behavior; lab safety |
| Information Systems and Technology | Engineering & Tech | HIGH est; CS/IT/cyber direct evidence | Code, databases, cloud, networking, business systems, cyber labs | R2 | Authorized systems only; no secrets; sandbox offensive-security work |
| Mechanical and Civil Engineering | Engineering & Tech | HIGH est; ME direct evidence, 2024 | Calculations, simulation, code, drawings, standards, reports | R3/R7 | Engineer review; current standards; physical and public safety |
| Technology Management and Mechatronics | Engineering & Tech | HIGH est; engineering, code, and business overlap | Controls, automation, code, operations, project planning | R2/R7/R8 | Physical-system safety; isolated testing; management decisions remain human |
| Accounting | Woodbury | HIGH est; direct classroom evidence, 2024 | Reconciliation, explanations, spreadsheets, memos, standards research | R8 | Current standards; audit trail; confidential financial data |
| Air Force and Army ROTC | Woodbury | MED est; organizational type partly unknown | Public doctrine study, writing, history, leadership scenarios | R7/R9 | No controlled, operational, or personal data; authority check |
| Business Graduate Studies | Woodbury | HIGH est; 78% program integration, 2024 | Cases, strategy, finance, operations, research, presentations | R8 | Confidential employer/client data; current sources; decision audit |
| Finance and Economics | Woodbury | HIGH est; broad business 70% weekly use | Modeling, data, policy, markets, explanations, research | R8/R3 | Current market and regulatory data; no unchecked financial advice |
| Marketing | Woodbury | HIGH est; 27% of business syllabi added AI tasks, 2026 | Research, segmentation, campaigns, content, tests, media | R1/R6/R8 | Brand control, copyright, privacy, bias, disclosure |
| Organizational Leadership | Woodbury | HIGH est; management and writing task fit | Cases, coaching practice, communication, change plans | R1/R8 | Personnel privacy; no automated employment decisions |
| Strategic Management and Operations | Woodbury | HIGH est; business and quantitative task fit | Process maps, forecasts, cases, supply chains, presentations | R8 | Current data, confidential operations, human decision ownership |
Unit names and college placement are fact from the 2026–27 catalog (each unit links to its catalog page); tier-basis links go to the underlying study. Two structural notes from the catalog sweep: Dental Hygiene and Respiratory Therapy nest under Allied Health; the ROTC unit's organizational type is partly unknown. Sequencing among colleges stays in Appendix A — this table is about what to build for whom, not who goes first.
What UVU already has
Appendix E in five lines
- QuestionWhat does Utah Valley University already have, and what should it not buy twice?
- AnswerThe public record shows an employee gateway, paid seats, a course helper, device management, computer nodes, and an active program office.
- Deciding numbers154 Copilot licensesfact51 ChatGPT licensesfact120 Artificial Intelligence Academy completersfact
- What the plan doesMeasure current seat use and reuse existing systems before adding another product or program.
- Still unknownCurrent seat use and prices, Canvas settings, internal contracts, and exact hardware are UNKNOWN.
Before this plan recommends a single purchase, it inventories everything UVU already owns, licenses, runs, or can reach — from board minutes, budget documents, IT service pages, procurement postings, and vendor announcements. The plan's stance toward all of it: assume none of it is available. Every existing system, seat, compute node, and staff hour is already allocated to the people using it today, so the service is designed to stand entirely on its own. This inventory exists so we know what we are building on top of — the integration surfaces (identity, LMS, device management, network, data rules) — and what we must never needlessly duplicate. It is a map of the house, not a list of things to borrow. This is a public-web inventory fact — it cannot see inside UVU's contract ledger, which is exactly why the ask-once packet at the end exists.
The headline assets
- An employee AI Gateway is already live — launched July 1, 2026 after a 360-person pilot: ChatGPT, Claude, Gemini, and local open-source models behind one door, about $5 in tokens per employee per week. Students not yet included. Our service keeps its own front door for its own users and can offer the Gateway a local-engine hook — a connection point, not a dependency, and never a competing campus-wide portal pitch.
- Paid seats already exist and are the experiment — 154 Copilot and 51 ChatGPT licenses deployed January 2026. Their usage data should decide seat expansion, not vendor pitches.
- Wilson, UVU's own course assistant — 44 courses and ~2,000 students by February 2026, funded ($163k donated services + $96k Academic Affairs). Our service is not a course chatbot and does not replace it; a Wilson-style assistant can call our engine if UVU chooses.
- The management layer is done — Jamf Pro already auto-enrolls every university Apple device. The Mac fleet this plan proposes needs zero new management software.
- Compute today is CPU-shaped — three UVU-owned Notchpeak nodes (~224 cores, 680 GB RAM, no disclosed GPU) plus guest routes to shared Utah GPU systems (80 H200s on One-U, preemptible; the Redtail supercomputer exists but no UVU allocation is public). Shared, preemptible access is not the same as owned capacity — and none of it is assumed available to this plan, which is exactly why the plan owns its floor.
- The program office exists and is funded — the Kahlert Applied AI Institute ($5.2M gift), AI Academy (120 completers), twice-monthly Power Hours, a faculty coaching pilot, curriculum grants, and a paid AI apprenticeship. New work routes through it.
Full inventory
| Grade | Asset | Scope and current evidence | What it means for the plan |
|---|---|---|---|
| FACT | Microsoft 365 | Office 365 is available to current students and employees. Basic Microsoft Copilot is included for all students and staff. A current UVU page says most employees have M365 A5 and receive Power BI Pro. UNKNOWN: the full A1/A3/A5 mix, especially for students. Software for Students, AI Toolkit, and Power BI licensing, undated; accessed 2026-09-03. | Do not buy a second general office, collaboration, or basic chatbot suite without a proven need. |
| FACT | Paid Copilot and ChatGPT seats | January 2026 employee deployment: 154 Copilot licenses and 51 ChatGPT licenses. Copilot had grown 175% since July 2024. UNKNOWN: SKUs, assigned/active seats, utilization, prices, owners, or renewals. Board of Trustees minutes, 2026-01-29, pp. 7–8. | The paid-seat model is already proven. Measure these seats before expanding or buying another product. |
| FACT; secondary detail | UVU AI Gateway | Free to all employees; no install required. Reported launch: 2026-07-01 after a 360-user pilot. Reported models: ChatGPT, Claude, Gemini, and local open-source models. Reported features include multi-model comparison, projects, knowledge bases, prompt chains, and agent export. Reported allowance: about $5 in tokens per employee per week, resetting Mondays. Students were not yet included. UVU employee resources, accessed 2026-09-03; TechBuzz report using UVU executive statements, published 2026-07-14. | Treat the Gateway as the front door. Build governance, retrieval, local models, and routing behind it instead of creating a rival portal. |
| FACT | Wilson AI / Ask Wilson | UVU-developed website and Canvas assistant. Fall 2025 board snapshot: 26 courses and more than 1,800 students. February 2026 CIO interview: 44 courses and 2,000 students. Current UVU page says “select” courses. UNKNOWN: the 2026-09-03 count, models, hosting, retention, and active-use rate. Board minutes, 2026-01-29; CIO interview, 2026-02-11; Wilson AI, accessed 2026-09-03. | Expand and measure Wilson before buying a second general course chatbot. |
| FACT | Wilson development funding | Earlier development received $163,000 in donated services and $96,000 from Academic Affairs for FY2025. The pilot had 14 classes in the October 2024 budget presentation. Digital Transformation PBA presentation, 2024-10-30, pp. 4 and 14. | Wilson is an existing funded product, not a concept-stage purchase. |
| FACT; EST for tier entitlement | Canvas by Instructure | Canvas is UVU’s only approved and supported LMS. Instructure says every current Canvas customer is now Canvas Core, with five included IgniteAI capabilities. EST: UVU therefore has Core entitlement. UNKNOWN: UVU feature enablement and any Plus/Next purchase. UVU Canvas technology, accessed 2026-09-03; Instructure Q2 2026 tier announcement, April 2026. | Test and enable included Core functions before buying substitutes. |
| FACT | Copyleaks | UVU’s current OTL page lists Copyleaks among the supported course technologies. An indexed UVU guide said its plagiarism and AI detection tool was integrated with Canvas and available to all instructors. Faculty Senate recorded false-positive concerns. UNKNOWN: current term, cost, and detector configuration. OTL, accessed 2026-09-03; indexed integration guide, published circa 2025, now returning 404; Faculty Senate agenda, 2025-04-29. | Do not buy a duplicate detector until Copyleaks’s accuracy, use, and renewal need are reviewed. |
| FACT | Canvas-associated tools | Teams, Kaltura, Qualtrics, Honorlock, and OneDrive are approved. Honorlock is UVU’s official online proctor after the move from Proctorio. Approved publisher LTIs include Cengage, Macmillan, McGraw-Hill, Pearson, VitalSource, and WileyPLUS. UNKNOWN: AI add-ons and contract terms. Canvas technologies and employee resources, accessed 2026-09-03. | Use the supported Canvas ecosystem. Do not assume publisher AI products are included merely because their LTIs are approved. |
| FACT | Zoom posture | Zoom is not approved or supported for employee or student use; Teams is UVU’s standard. An isolated purchase cannot be ruled out. Zoom at UVU, accessed 2026-09-03. | Do not plan around Zoom AI Companion. Use Teams unless UVU changes its standard. |
| FACT | Adobe Creative Cloud | Available to every active student, employee, and faculty member. UNKNOWN: Adobe Firefly entitlement, credit pool, admin controls, or paid premium features. Adobe Creative Cloud Suite, accessed 2026-09-03. | Do not buy a duplicate creative suite. Confirm Firefly credits before buying a separate image service. |
| FACT / UNKNOWN | Grammarly | The free version is installed in open labs. UNKNOWN: any paid institutional Grammarly contract. EUTS software inventory, accessed 2026-09-03. | Do not describe Grammarly as a paid campus license without internal proof. |
| FACT | Lucidchart and Lucidspark | Free education upgrade for students and employees. Software for Students, accessed 2026-09-03. | Reuse for diagrams and collaborative ideation before adding another whiteboard product. |
| FACT | LinkedIn Learning | Free access is publicly stated for students and employees. Board data records 161 AI-course uses in Spring 2025 and 301 in Fall 2025. UNKNOWN: current contract term, cost, and whether counts represent unique people. Online Student Hub, accessed 2026-09-03; Board minutes, 2026-01-29, p. 11. | Use the existing catalog before buying generic AI-awareness training. |
| FACT / UNKNOWN | Pluralsight | At least some institutional licenses exist. The Art & Design page says unlimited access for all current students, faculty, and staff; the central learning page describes narrower Dx and Academic Affairs coverage. Art & Design labs and Dx Learning, accessed 2026-09-03. | Confirm scope once. Do not budget new technical-course licenses until the conflict is resolved. |
| FACT / UNKNOWN | Google services and Gemini | UVU Status lists Google Suite as operational. Broad `MY.UVU.EDU` Google Workspace access closed on 2023-08-11, while curriculum-specific `UVU.EDU` use may remain. Gemini is proven through the employee Gateway, not as a separate campus-wide license. UVU Status, checked 2026-09-03; Google Transition, closure 2023-08-11. | Treat Gateway Gemini as the proven route. Confirm any remaining Workspace or Gemini Education contract before relying on it. |
| FACT | Microsoft storage | OneDrive and SharePoint are official storage for all faculty, staff, and students. Box and Dropbox have been discontinued for most users. Planned limits effective 2026-12-31 are 100 GB per person and 500 GB per Teams/SharePoint site. File Storage, accessed 2026-09-03. | Build governed AI retrieval against Microsoft stores. Do not center new AI work on Box or Dropbox. Account for the coming quotas. |
| FACT | Civitas Learning | UVU’s predictive student-success platform: Course Insights, Persistence Insights, Completion Insights, and Inspire. Access is available by request to relevant faculty, staff, administrators, advisors, and support teams. Civitas access article, published 2023-10-19; current page checked 2026-09-03. | Reuse existing risk and persistence analytics rather than creating a second student-risk model without need. |
| FACT / UNKNOWN | BI and data platforms | Current public evidence supports Power BI, Power BI Pro for most A5 employees, Tableau dashboards, Business Objects, MyDataHub, a university data warehouse, a data lake, and renewed Collibra. Canvas, Jira, finance-budget, and EvolveFM data had been moved into the data lake by October 2024. UNKNOWN: lake vendor and current architecture. Power BI, Tableau dashboard example, and Institutional Reports, accessed 2026-09-03; Dx PBA, 2024-10-30. | Start from the governed data stack. Do not assume Databricks or Snowflake is needed. |
| UNKNOWN | Databricks and Snowflake | Exact official-domain searches found course, résumé, and employment references but no institutional deployment. Public-web search completed 2026-09-03; no qualifying UVU deployment source found. | Do not list either platform as an owned asset or proposed need until the internal architecture is checked. |
| FACT | Enterprise software estate and funding context | Dx reported 211 maintained titles and a $5,998,301 software-expense snapshot in October 2024. It requested $600,000 ongoing for one year of M365 after five prepaid CARES-funded years; that figure is a request, not verified final spend. Final FY2025 allocations included $150,000 for AI Institute startup and a bundled $200,000 for background checks, LMS, AwardCo, and LinkedIn Learning increases. Dx PBA, 2024-10-30; FY2025 final allocations, FY2025. | The public list is only a small part of UVU’s software estate. Use the internal catalog and contract ledger for purchase checks. |
| FACT | Jamf Pro | MDM for university-owned Macs, iPhones, iPads, and Apple TVs. New university Apple devices auto-enroll; enrollment is mandatory. UNKNOWN: device count and coverage rate. Jamf Pro, accessed 2026-09-03. | No separate Apple MDM purchase is needed. |
| FACT | Labs and MyApps | Physical labs provide Adobe, Microsoft, SPSS, creative software, and other tools. MyApps gives all students virtual access to Adobe, ArcGIS, Avid, SPSS, NVivo, Power BI Desktop, RStudio, SAS, Stata, VS Code, Mathematica, and other software. EUTS software inventory, accessed 2026-09-03. | Pilot software and modest models through existing labs/virtual delivery before buying a new general lab. |
| FACT / UNKNOWN | Apple labs | The Fall 2026 Art & Design page lists six Apple labs; EUTS lists four rooms and a different fourth-room number. Models, station counts, Apple silicon, RAM, and GPUs are not public. Art & Design labs and EUTS software, accessed 2026-09-03. | The Mac fleet exists, but a hardware census is needed before sizing local AI work. |
| FACT; EST totals | Three UVU-owned Notchpeak CPU nodes | One 96-core/380 GB node and two 64-core/150 GB nodes. UVU account holders receive priority submission through SLURM. EST: 224 cores and 680 GB RAM total. No GPU is disclosed on these nodes. UVU High Performance Computing Resources, undated; accessed 2026-09-03. | Use this capacity for CPU preprocessing, simulation, inference, and teaching. Do not count it as GPU capacity. |
| FACT / EST | Shared CHPC GPU access | CHPC supports UVU accounts and has general GPU resources that do not require allocations. One-U Responsible AI has ten nodes with eight H200s each; all CHPC users may submit guest/preemptible jobs. EST: active UVU CHPC users can use guest capacity. Priority eligibility for an external UVU group is unknown. CHPC accounts, updated 2026-08-26; CHPC GPU guide, updated 2026-06-04; One-U AI resource notice, January 2026 issue. | Benchmark queue time and usable GPU hours before buying hardware. Shared/preemptible access is not the same as a dedicated cluster. |
| FACT statewide; UNKNOWN for UVU | Redtail supercomputer | Statewide system: 33 nodes, 264 NVIDIA H200 SXM5 GPUs, 3,696 CPU cores, 66 TB RAM, and about 1 PB usable scratch. Early-access proposals were open to USHE faculty and staff. No public UVU allocation was found. Redtail, updated 2026-07-08; launch report, 2026-07-23. | Confirm whether UVU received access before including Redtail in a capacity plan. |
| FACT / EST / UNKNOWN | Smith Engineering Building | The nearly 200,000-square-foot building opened 2026-01-22. Post-opening evidence confirms a large drone lab with cameras, workstations, and drones. Pre-opening documents name AI, VR, computer-science, manufacturing, and prototyping spaces. A brochure planned a 30-person interactive AI classroom. Installed GPU and workstation details are unknown. Ribbon cutting, 2026-01-22; building preview and named-spaces brochure, 2025-04-14/undated. | Use the building’s real labs, but do not claim a GPU cluster until the as-built inventory is supplied. |
| FACT service / UNKNOWN hardware | SCaFL and Business Resource Center AI lab | Current page advertises AI inference training, simulation, big-data/data-science work, fabrication, and a UTOPIA “10GB” connection. Independent use requires consultation, certification, and fees. Its former server-detail page now returns 404. SCaFL, accessed 2026-09-03. | Reuse the lab for pilots after confirming server specifications and availability. |
| FACT | NVIDIA partnership | Three-year voluntary collaboration with no exchange of funds. Includes DLI ambassador training, teaching kits, workshops, materials, advanced tools, and GPU-accelerated cloud workstations used for DLI activity. UNKNOWN: persistent credits, GPU hours, cloud tenant, research allocation, or donated hardware. NVIDIA announcement, 2025-03-10; UVU partnership page, accessed 2026-09-03. | Use the training and workshop rights. Do not book “NVIDIA cloud credits” as compute capacity until quantified. |
| FACT | Kahlert Applied AI Institute funding | $5.2 million Kahlert Foundation gift announced 2025-10-07. Earlier final PBA allocation supplied $150,000 in startup funds. UVU-issued gift announcement, 2025-10-07; FY2025 allocations. | Use the funded institute as the central program office instead of creating a parallel AI office. |
| FACT | AI Academy and training | Board snapshot: 120 AI Academy completers, 22 Advanced AI participants, and more than 3,000 AI learning opportunities. AI in Action counts were 349, 204, 484, and 1,875 for Fall 2024 through Fall 2025. Board minutes, 2026-01-29, pp. 0 and 11. | The base training layer exists. Focus new spend on role-specific practice and measured outcomes. |
| FACT | Power Hours and faculty coaching | Power Hours run twice monthly for faculty and staff. The faculty coaching pilot is free and confidential, uses three sessions per semester, and is limited to a small founding cohort. AI Institute and faculty coaching, accessed 2026-09-03. | Scale the existing coaching model if results support it; do not buy generic coaching first. |
| FACT | 2026 Faculty Summer Institute AI track | May 5 build lab, two-week sprint, May 29 showcase, and $1,500 stipend after deliverables. Faculty built reusable course agents and AI-ready assignments. Applied AI Fluency in the Classroom, 2026 program. | Reuse its templates, deliverables, and trained faculty. |
| FACT lower bound / UNKNOWN total | Curriculum Integration Grants | At least two distinct 2026 awards are public. One Cao/Tang project is explicitly $5,000; Devin Gilbert lists a separate award with no public amount. Total awards, dollars, courses, and results are unknown. Cao/Tang profile and Gilbert CV, accessed 2026-09-03. | Treat two projects as the proven floor, not the full grant count. |
| FACT | Applied AI Apprenticeship | Paid, registered, full-time, 12 months; current partners SchoolAI and AskElephant. First cohort began August 2025. Completers earn industry certifications and a stackable business-analysis certificate. UNKNOWN: cohort size, completions, placements, and Clarion’s current role. AI Apprenticeship, accessed 2026-09-03. | Expand the existing employer pathway rather than creating a separate apprenticeship structure. |
| FACT | Academic AI programs | Current offerings include a 30-credit MS in Applied AI, an 18-credit AI Graduate Certificate, two 18-credit undergraduate AI certificates, and a 120-credit Information Systems BS with an Applied AI emphasis. MS, graduate certificate, Applied AI certificate, Applied AI in Organizations, and BS emphasis, 2026–27 catalog. | New credentials should fill a measured gap, not duplicate this stack. |
| FACT | UNESCO Chair and AI committee | UNESCO Chair on AI and Environmental Stewardship for Sustainable Futures, established in 2025, Chair ID 2025US3042. Current UVU AI and Education committee has 12 named members. UNESCO registry, 2026-01-12; UVU committee, accessed 2026-09-03. | Use this as an ethics, education, sustainability, and international network—not as evidence of funded compute. |
| FACT announcement / UNKNOWN delivery | USHE AI Task Force and statewide credential | Barclay Burns is one of 11 listed task-force members. USHE announced a free AI Workforce Credential for more than 50,000 eligible 2025–27 graduates, beginning 2026-07-01. A June 2026 state briefing still described an August RFP. No public launch, platform, enrollment, or UVU delivery record was found. USHE announcement, 2026-05-01; legislative briefing, 2026-06-17. | UVU students may be eligible, but do not count UVU as the platform or credential operator until confirmed. |
| FACT / UNKNOWN | AI and data governance | Policy 445 governs institutional data. Policies 446, 447, and 452 cover privacy, security, and accessibility. Policy 441 on AI was proposed in 2025 but is absent from the approved manual as of the cutoff. Policy 445, effective 2024-03-28; Policy Manual, checked 2026-09-03; Policy 441 executive summary, 2025-09-25. | Build on the approved data controls. Confirm the binding AI standard before adding another governance layer. |
Canvas: what the LMS tier already includes
Instructure moved every current customer to Canvas Core, which includes five AI functions at no extra cost — but inclusion is not enablement, and the free preview of the premium tier ended June 30, 2026. First move: turn on and test what's already paid for.
| Canvas capability | Tier | Availability & control | UVU state |
|---|---|---|---|
| IgniteAI Summaries for Discussions | Core | Generally available 2026-03-21; disabled by default; admin opt-in. | UNKNOWN: no public UVU enablement proof. |
| IgniteAI Search for Courses, formerly Smart Search | Core | Institution opt-in; semantic course-content search; admins may delegate teacher control. | UNKNOWN: no public UVU enablement proof. |
| IgniteAI Translations for Discussions, Inbox, and Announcements | Core | Included for all tiers; feature control remains with institution/educator. | UNKNOWN: no public UVU enablement proof. |
| Question Authoring Assistance for Quizzes | Core | Included without an upgrade. | UNKNOWN: no public UVU enablement proof. |
| Accessibility Remediation for Courses | Core | Included in the Course Accessibility Checker. | UNKNOWN: no public UVU enablement proof. |
| Discussion Insights, Rubric Generator, and Grading Assistance | Plus or Next | Paid higher tier. | UNKNOWN: no public proof UVU has Plus or Next. |
| IgniteAI Agent, Ask Your Data, and Study Tools | Next | Paid highest tier. | UNKNOWN: no public proof UVU has Next. |
The gap ledger — "already have, therefore…"
Read the "therefore" column as a duplication guard, not a borrowing plan: it says what we must not pitch or buy twice. The service still brings its own front door, engine, operations, and budget.
| Already have | Therefore the plan does — or does not — need |
|---|---|
| FACT: M365 and basic Copilot for all students/staff; AI Toolkit, accessed 2026-09-03. | No second broad productivity suite. Start with adoption, role fit, and data boundaries. |
| FACT: 154 paid Copilot and 51 ChatGPT seats; Board minutes, 2026-01-29. | No automatic seat expansion. First obtain assignment, use, renewal, and outcome data. |
| FACT: Employee AI Gateway; UVU employee page, accessed 2026-09-03. | No competing campus-wide portal pitch. The service keeps its own front door for its own users and offers the Gateway a local-engine hook. |
| FACT: Wilson AI reached 44 courses/2,000 students; CIO interview, 2026-02-11. | No generic course-chatbot purchase until Wilson’s measured limits are known. |
| EST: Canvas Core includes five AI functions; Instructure, April 2026. | Enable and test summaries, search, translation, quiz assistance, and accessibility remediation before buying duplicates. |
| FACT: Copyleaks is supported in the Canvas environment; OTL, accessed 2026-09-03. | No second detector without an accuracy or contract reason. Detection results should not be treated as sole proof of misconduct. |
| FACT: Teams is standard; Zoom is not approved; Canvas technology, accessed 2026-09-03. | No Zoom AI Companion plan. Use Teams and its included capabilities. |
| FACT: Adobe Creative Cloud is campus-wide; Adobe KB, accessed 2026-09-03. | No replacement creative suite. Confirm Firefly credits and controls first. |
| FACT: Power BI, Tableau, Civitas, Business Objects, MyDataHub, a warehouse/data lake, and Collibra exist; Dx PBA, 2024-10-30. | No assumed Databricks/Snowflake purchase. Map the current stack and missing workload first. |
| FACT: OneDrive/SharePoint are official; Box/Dropbox are mostly discontinued; File Storage, accessed 2026-09-03. | Put governed retrieval on Microsoft stores, not on a new Box/Dropbox knowledge layer. |
| FACT: Jamf manages university Apple devices; Jamf, accessed 2026-09-03. | No Apple MDM purchase. Use Jamf inventory to find AI-capable Macs. |
| FACT: Physical labs and MyApps already deliver software; EUTS, accessed 2026-09-03. | Pilot on the existing fleet before buying a general AI lab. |
| FACT: Three dedicated CPU nodes plus general CHPC access; UVU HPC, accessed 2026-09-03. | Use CPU nodes for preprocessing and suitable work; measure shared GPU queues before GPU capital spending. |
| EST: UVU users can reach preemptible One-U H200 capacity; CHPC, January 2026. | The plan still may need dedicated GPU capacity if queueing, privacy, or uptime requirements fail; shared access is not ownership. |
| UNKNOWN: Redtail allocation; Redtail, updated 2026-07-08. | Confirm UVU’s allocation before treating Redtail as available capacity. |
| FACT: NVIDIA training and cloud-workstation access for DLI activity; NVIDIA, 2025-03-10. | Use activated DLI benefits. Do not count unspecified “cloud credits” in capacity planning. |
| FACT: AI Academy, Advanced AI, AI in Action, LinkedIn Learning, Power Hours, coaching, and faculty institutes exist; Board minutes, 2026-01-29. | No new generic AI literacy catalog. Fund role-based practice, coaching capacity, and outcome measurement. |
| FACT: $5.2 million gift, startup funds, grants, and apprenticeship structure exist; gift release, 2025-10-07. | No parallel AI program office. Route new work through the institute and disclose how current funds are committed. |
| FACT: Data governance, privacy, security, and accessibility controls exist; Policy Manual, checked 2026-09-03. | No duplicate governance framework. The missing item is a confirmed binding AI policy and operating decision record. |
| FACT: UNESCO and USHE relationships exist; UNESCO, 2026-01-12; USHE, 2026-05-01. | Use them for standards, reach, and coordination. Do not count them as software, funding, or compute without a specific award. |
The full data ledger — deferred diligence, not a precondition
The twelve items below were the original public-record gaps. A second public sweep (Appendix K) closed most of what matters for building and reduced the real ask to five interface questions. This ledger is kept for later commercial diligence — request it as an export when useful, never as a gate on the pilot.
- Software and spend: Export UVU’s current AI-relevant license ledger with product, SKU, vendor and reseller, contract/PO, owner, funding unit, term, renewal, purchased seats, assigned seats, 30/90-day active users, and annual cost. Include OpenAI, Anthropic, Microsoft, Google, Adobe, Instructure, Copyleaks, Grammarly, Turnitin, Zoom, Qualtrics, Kaltura, Honorlock, LinkedIn Learning, Pluralsight, Tableau, Civitas, ServiceNow, Salesforce, AWS, Azure, GCP, Databricks, and Snowflake.
- Microsoft: Confirm the current student and employee A1/A3/A5 mix; the exact 154 Copilot SKU; the 51-seat ChatGPT tier; Power BI Pro/Premium capacity; Copilot Studio/Azure OpenAI access; and the outcome of the October 2024 $600,000 M365 request.
- Gateway: Provide its current model/provider list, hosting diagram, API contracts, weekly-budget rule, total and per-model spend, paid-tenant rules, active users, student-launch decision, retention, provider-training terms, logging, data classes, and approved connectors.
- Wilson AI: Confirm the 2026-09-03 course and section count, unique students, active users, current name, models/providers, retrieval and Kaltura design, hosting, cost, retention, provider-training terms, assessment results, and whether the Faculty Senate oversight group was formed.
- Canvas: Confirm Core/Plus/Next contract status, term and cost; current feature flags for all IgniteAI functions; data-processing terms; Wilson, Copyleaks, Honorlock, Kaltura, and publisher LTI status; and whether Khanmigo or another tutoring pilot exists.
- Compute: Confirm all three UVU Notchpeak nodes are online and provide CPU model, hostnames, usable cores/RAM, storage, partition/QoS, priority terms, and FY2026 use. Confirm UVU eligibility and limits for general GPUs, One-U H200 priority, and any Redtail allocation.
- Campus hardware: Supply the current Jamf and IT asset export for Macs, GPU workstations, servers, Citrix GPU profiles, Smith Building AI/VR/drone/robotics rooms, SCaFL servers, DGM/XR equipment, and any surviving departmental GPU servers. Include model, count, room, owner, access rule, and operational state.
- NVIDIA: Provide the executed or redacted MOU, exact term, activated deliverables, current ambassadors, completed certifications/workshops, learner counts, cloud tenant, instance type, GPU hours, dollar credits, expiry, and any donated equipment.
- People and funding: Give one dated roster for AI Academy, Advanced AI, Power Hours, coaching, Summer Institute, curriculum grants, and apprenticeships. Include unique people, completions, outcomes, every grant/project/amount, apprenticeship employers and participants, credentials, placements, and the committed versus uncommitted Kahlert budget.
- Data and agreements: Confirm the warehouse/data-lake architecture and current roles of Collibra, Power BI, Tableau, Business Objects, MyDataHub, Civitas, Databricks, and Snowflake. List all executed AI MOUs, DPAs, DUAs, and data-sharing agreements with data classes, processors, retention, model-training use, IP, and expiry.
- State and UNESCO: Confirm whether the AI Workforce Credential actually launched, who owns the platform, UVU’s delivery role, enrollment/completions, and RFP status. Provide the current UNESCO Chair and committee charter, funding, partner agreements, and AI/data projects.
- Public-payment reconciliation: Produce UVU vendor payments for FY2024–FY2026 for the named vendors, including payments through VLCM, CDW, SHI, Carahsoft, cooperative contracts, marketplaces, and other resellers. Reconcile them to contracts and the license ledger.
Search coverage note: procurement search fields showed no events for the major AI vendors — which proves only that the searchable fields are empty, not that no purchases exist (reseller, cooperative, and low-value purchases don't appear there). The state transparency portal's vendor search loads through a client-side form that automated public access could not operate; the reconciliation request above covers it fact.
The free-and-discounted catalog
Appendix F in five lines
- QuestionWhat can Utah Valley University claim before it buys more?
- AnswerUse the $0 programs first, get a paid-seat quote in week 1, and let the small pilot set the hardware need.
- Deciding numbersMicrosoft Copilot Chat $0factpilot Mac $2,139factseat benchmark $1.60/mofact
- What the plan doesClaim the $0 rows, seek research awards for later growth, and let measured pilot use set the owned hardware size.
- Still unknownA binding university seat quote is NOT_RUN, and several program amounts, access rules, and renewal terms are UNKNOWN.
| Program | What UVU gets | Cost | The catch | Grade |
|---|---|---|---|---|
| Microsoft Copilot Chat | Enterprise-protected chat for everyone on the existing M365 A-license — already deployed at UVU | $0 | No campus control of models or data locality; per-user paid tier is the upsell path | fact |
| Google Gemini for Education | Gemini with enterprise data protection free on an edu tenant; Utah's K-12 system already adopted statewide (June 2026) | $0 | Requires tenant setup and policy review; free tier terms can change at renewal | fact |
| OpenAI researcher access | Frontier-model seats for named researchers (institutional program, ~5 per institution) | $0 | Named individuals only — not a campus service | fact |
| Anthropic science program | Scientist seats plus API credits for approved research projects — published ceiling up to $50k credits per project | $0 | Application-gated, project-scoped, renewal not guaranteed | fact |
| NAIRR pilot | National AI Research Resource compute allocations for research projects | $0 | Competitive applications; batch research compute, not a campus assistant | fact |
| NVIDIA–UVU agreement | Training, certification, tools, and cloud resources under the existing 3-year no-funds MOU | $0 | No hardware included; workforce/curriculum scope | fact |
| Apple education pricing on NASPO ValuePoint | Mini M5 Pro 48GB at $2,139 (vs $2,299 retail) on the competed Utah PA4282 cooperative contract — no new bid required through June 2027 | −7% | Education price verified for the pilot config; volume quotes may improve it (ask) | fact |
| Utah AI Moonshot | $5M state program; UVU's research office lists it with October deadlines | grant | Notice of intent was due Aug 28 — first question for Dr. Burns: was one filed? | fact |
| HERFP research fund | $45M pool; HB 373 reserves 15–25% of research money for regional institutions like UVU | grant | Competitive; proposal work required | fact |
| HB2 AI compute ($15M) | Statewide AI data-center appropriation | partner | Sits under University of Utah Item 81 — UVU access is a partnership conversation, not an entitlement | fact |
| Hosted open-weight APIs | Burst overflow on open models (gpt-oss-120b class) at commodity metered rates — roughly $0.17 per user-year at campus scale in our demand model | metered | Zero-data-retention terms must be verified in writing per provider | est |
| Negotiated edu seats (benchmark) | Cal State's renewal: ~$19.26/user·yr at 675k users; Colorado ~$20; Maine ~$22.64 per billed FTE | $1.60/mo | Benchmarks, not quotes — UVU-scale pricing must be quoted in week 1 (the plan's F1 branch takes the deal if it beats the floor) | fact |
Education and research programs — the wider sweep
A second pass on September 3 checked every education-specific program a campus could claim, including several missed above. Personal benefits (a student's own Azure credit, a verified student's free GitHub Copilot) are real and worth publicizing; none is campus capacity.
| Program | Current benefit | Limit or catch | Campus value |
|---|---|---|---|
| Google Gemini for Education | No added charge through qualifying Education Fundamentals accounts; institution controls and enterprise protections. Program | Product limits remain; not equivalent to an unrestricted Gemini API pool. | High. Use as a primary entitlement. |
| Microsoft Copilot Chat | No added charge for Microsoft 365 A1/A3/A5 faculty, staff and higher-ed students 13+. Current education page | Full Microsoft 365 Copilot is $18/user/month; agents and Graph integration differ. | High. Use existing tenant controls. |
| GitHub Copilot Student | Free to verified students; teachers and some maintainers can receive Copilot Pro. Eligibility | Coding assistant, not a general campus inference API. Eligibility checked regularly; automatic model selection applies. | High for coding courses. |
| OpenAI Academic Researchers | Up to five managed seats for 12 months; business protections and Pro-level limits. FAQ, updated Aug. 2026 | Selective; waitlist; no API credit. | Research teams only. |
| OpenAI Researcher Access | Up to $1,000 API credits for 12 months; quarterly review. Program FAQ | Responsible-AI research, not operations. | Apply project by project. |
| OpenAI Codex for Students | Verified U.S./Canada university students receive $100 in Codex credits. Terms, updated 2026-09-02 | Personal ChatGPT/Codex credit, not general API credit. | Useful for student coding. |
| Anthropic Claude for Education | Institution plan with Learning Mode, SSO, training and faculty research API credits. Program | Paid; public seat price and API-credit amount unknown. | Procurement candidate, not free entitlement. |
| Anthropic AI for Science | Newer 2026-08-27 announcement offers up to $50,000 per project. Announcement | Selective scientific research. Older help material still says $20K. | Strong grant target. |
| Anthropic External Researcher Access | Typically $1,000 API credit for qualifying safety/alignment research. Program | Selective; not operational capacity. | Apply when research fits. |
| Perplexity Education Pro | Verified students/educators: $10/month; institutional Enterprise Pro: $300/user/year. Pricing, updated 2026-09-02 | Discounted, not free. Personal plan lacks institutional controls. | Optional personal benefit. |
| Notion Education | Free Plus workspace for individual students/teachers; verified student organizations can receive shared workspaces. Education plan | Notion AI is only an undisclosed limited trial on Free/Plus/Education. | Productivity benefit, not AI capacity. |
| NVIDIA Teaching Kits | Free teaching material; approved courses can receive codes worth up to $90/course/student for labs. Teaching Kits | Course-specific approval; not an unrestricted hosted-NIM grant. | Good for GPU/AI courses. |
| NVIDIA Inception/DGX credits | Inception is free; qualified startups have seen offers up to $100K DGX Cloud credit. | Startup program, not a university program. No public general renewal. | Only for qualifying spinouts. |
| AWS Educate/Academy | Free learning content and Academy labs. Research-credit proposals accept faculty and students. Research FAQ | Student research awards capped at $5,000; faculty cap not public. Credits generally expire in one year. Educate no longer provides the old blanket cloud credits. | Course labs and funded research. |
| AWS general Free Tier | $100 at signup plus up to $100 earned. | New customers; free-plan account lasts at most six months. | Onboarding only. |
| Azure for Students | $100 for 12 months, no card, renewable annually while eligible; Azure OpenAI is within the service catalog. Offer | One account/customer; education, teaching and noncommercial research only; nontransferable. | One of the strongest legitimate personal cloud benefits. |
| Google Cloud teaching credits | Up to $100 per teaching staff member and $50 per student; generally usable for 12 months after course start. Faculty program | Course-specific, nontransferable. Google’s general $300 trial cannot fund Gemini Developer API charges after March 2026. | Good course sandbox. |
| Google Cloud research credits | Faculty researchers and PhD candidates may apply. | Amount unknown; project award, not campus entitlement; expires after about one year. | Apply per project. |
| Hugging Face Academia Hub | SSO/admin/audit plan with $2 monthly compute credit per seat. Program | Paid: starts at $10/seat/month, minimum 250 annual seats. | Possible managed teaching/research hub. |
| Hugging Face Classrooms | Free teaching workspaces and resources. Classrooms | Compute/API allowance unknown. | Course organization, not capacity promise. |
| NAIRR | Selective access to federal and industry compute, data, models and training. NSF NAIRR | Allocation-based; not an automatic university entitlement. | Continue pursuing for research workloads. |
Order of operations: claim the $0 rows first (they kill the "we can't afford anything" objection and generate the adoption data every later decision needs), quote the seat benchmarks in week 1, and let the measured pilot decide how much owned floor to build. The research-grant rows fund scale, not the pilot — the pilot is deliberately small enough to need no grant.
The contingency playbook
Appendix G in five lines
- QuestionWhat does the plan do when a main price, speed, or use test fails?
- AnswerEach failure has a set response: switch what is bought, stop adding computers, or freeze growth.
- Deciding numbers≤$30/user·yrestunder 70% of projected speedestunder 10% active after 30 days and under 20% return at month 3est
- What the plan doesUse seats if they pass the price and safety gates; stop growth and compare other paths if speed or use misses.
- Still unknownThe hardware-speed and user-return checks are NOT_RUN.
The architecture dossier carries a full branch matrix: eight bureaucratic (B1–B8), seven technical (T1–T7), four financial (F1–F4), five demand (D1–D5), and four external (E1–E4) branches, each with a trigger, a pre-decided response, and an owner. Eight weekly tripwires watch the triggers. The three highest-leverage pre-decisions, in full:
| Branch | Trigger | Pre-decided response |
|---|---|---|
| F1 — the quote flip | A binding, all-in seat quote at ≤$30/user·yr that passes security, accessibility, retention, and exit gates | Take it. Seats become the faculty front door for Public-tier work; the owned fleet contracts to a 4-Mac Sensitive-data enclave plus evaluation and teaching. Gateway, governance, and adoption work carry over unchanged. The plan is not loyal to hardware — it is loyal to the floor price. |
| T1 — the benchmark miss | Shipped M5 minis measure under 70% of projected throughput on two independent runs | Stop at 4 boxes. Run a same-corpus bake-off: Studio tier vs Linux/NVIDIA vs hosted APIs vs seats. Lowest passing 3-year cost wins. No sunk-cost defense — $12.7k is the price of knowing. |
| D2 — the adoption floor | Under 10% 30-day active and under 20% return rate at month 3 | Freeze scaling. Twenty user interviews, two funded department workflows, re-measure at month 6. If still under, end the general service and keep the enclave — the honest outcome beats a zombie service. |
Five decision memos are pre-written for the slower-burning cases: a major model release, a central-IT takeover offer, agent/automation traffic growth, retrieval workloads crushing prefill, and one department dominating usage. The full matrix travels with the architecture dossier rather than this page — it is an operating document, not a pitch.
Open items & the ask-once list
Appendix H in five lines
- QuestionWhat is still open, and who must answer it?
- AnswerThe pilot can start, but campus growth waits on room, service, insurance, access, records, money, sponsor, and buying decisions.
- Deciding numbersfour decisions only Dr. Burns can makeestfive technical questions for the information technology teamest$30 per user per yearest
- What the plan doesAsk each question once, name an owner, and close the needed items before buying, first use, or campus growth.
- Still unknownRoom approval is absent; insurance and the access report are UNKNOWN; records work is shallow; Dr. Burns’s answers remain open.
A plan that hides its unknowns isn't one. These are the items public evidence could not settle unknown, each with its next action and its owner-to-be. Nothing on this list blocks the validation pilot; several block campus scale.
| Item | State | Next action | Owner-to-be |
|---|---|---|---|
| Facilities, network, and continuity sign-off for the pilot room | Absent | One-page named-room sign-off before the purchase order | Facilities + Dx |
| Interim service terms (acceptable use + incident taxonomy) | Absent | Counsel-approved one-pager before the first user (Maine and UT-Austin templates exist) | Counsel + service owner |
| Insurance posture (property, cyber, E&O for AI advice) | Unknown | Written confirmation from risk management | UVU risk office |
| Asset accounting, tagging, and disposal path | Mostly policy-resolved | Expense code + custodian + retirement path named in the purchase packet | Controller |
| Chat-platform accessibility report (VPAT/ACR, WCAG 2.1 AA per Policy 452) | Unknown availability | Demand the exact report and version; test the deployed build regardless | Accessibility office |
| Records retention / GRAMA / legal holds vs user deletion rights | Shallow | Retention-table reconciliation | Records officer + counsel |
| Concurrent-enrollment minors | Excluded by design | Separate consent and safety design before any phase-2 inclusion | Program council |
| Utah AI Moonshot / HERFP notice-of-intent status | Unknown | Ask in week 1 — deadlines are October | Dr. Burns |
The four decisions only Dr. Burns can make — asked once, here, so no one asks twice (the five technical questions for UVU's IT side are in Appendix K):
- Was a Moonshot or research-fund notice of intent filed before the August 28 deadline?
- Is there a budget number — or should the simulator's dial stay the planning instrument?
- Which college sponsors the pilot, and does Dx co-sponsor identity and security review from day one?
- If a seat vendor beats $30 per user per year all-in (branch F1), should the flip be taken automatically or brought back for a decision?
The campus map data
Appendix I in five lines
- QuestionWhat place data and image rights make the campus map safe to use?
- AnswerUse the checked college homes and building pins with public federal aerial photos, not Utah Valley University’s map art.
- Deciding numbers53 building pinsestmedian offset near 19 metersfactone flagged pin ~880 m from its mapped buildingfact
- What the plan doesPut each college at its checked home, show split sites, and keep the federal photo credit on the map.
- Still unknownExact roles at four outlying sites and Payson’s construction schedule are UNKNOWN.
The planning map plots colleges on real overhead imagery. This appendix is its ground truth: every building pin on UVU's live campus map, the verified college→building mapping, every UVU location beyond the Orem core, and the licensing that makes the imagery reusable. Coordinates are UVU's own map pins rounded to four decimals — label and entrance points, not surveyed centroids (a footprint cross-check put the median offset near 19 meters) fact.
Where each college lives
Verified against deans' offices, advising pages, and department locations — not assumed from building names. Two findings worth pinning: the engineering core moved to the new Smith building in January 2026 (older maps still label the CS building as its home), and Health & Public Service is genuinely split — administration in the Losee Center, health teaching across the pedestrian bridge on West Campus, public safety in EN.
| College / school | Bldg | Role | Evidence |
|---|---|---|---|
| Smith College of Engineering & Technology | SE | PRIMARY; dean, Computer Science, Digital Media, Mechanical/Civil Engineering, Electrical/Computer Engineering, labs fact | SE building page, 2026 opening |
| Smith College of Engineering & Technology | CS | TRANSITIONAL_SECONDARY; some advising/support pages still use CS after the SE move est | CET advising, college directory |
| Smith College of Engineering & Technology | SA | SATELLITE; automotive and transportation technology fact | Transportation Technology |
| Woodbury School of Business | KB | PRIMARY; dean, advising, Accounting, Finance, Economics, Marketing, Management and Operations fact | Woodbury directory, advising |
| College of Humanities & Social Sciences | CB | PRIMARY; dean, advising and all principal department offices fact | CHSS directory, CHSS advising |
| School of Education | ME | PRIMARY; expressly identified as the school’s main hub for dean/faculty, advising and classrooms fact | Education locations |
| School of Education | NB | SATELLITE; autism classrooms, outreach and research fact | Education locations |
| School of Education | LC | SATELLITE; Student Leadership and Success Studies fact | Education locations |
| College of Science | SB | PRIMARY_ADMIN; dean and Biology fact | College overview, Biology |
| College of Science | PS | PRIMARY_TEACHING_LABS; Chemistry, Earth Science, Physics, advising and planetarium fact | Chemistry, Planetarium |
| College of Science | LA | SATELLITE; Mathematics fact | Mathematics |
| College of Science | RL | SATELLITE; Exercise Science fact | Exercise Science |
| College of Health & Public Service | LC | PRIMARY_ADMIN; dean’s office in LC 414 fact | CHPS locations |
| College of Health & Public Service | HP | WEST_HEALTH_HUB; remaining health teaching and labs, but older program lists are partly stale est | UVU locations, CHPS locations, current Dental location, current Respiratory location |
| College of Health & Public Service | EN | PUBLIC_SAFETY_CORE; Criminal Justice and Forensic Science activity fact | CHPS locations |
| College of Health & Public Service | ME | SPECIALIST; forensic laboratory fact | CHPS locations |
| College of Health & Public Service | H6 | SPECIALIST; Center for National Security Studies fact | CHPS locations |
| School of the Arts | NC | PRIMARY; dean, Music, Theatre, performance and rehearsal facilities fact | Arts directory, Noorda facilities |
| School of the Arts | GT | ART_DESIGN_CORE; Art & Design office, studios, labs and galleries fact | Art & Design, facilities |
| School of the Arts | LA | DANCE_SATELLITE; Dance office and rehearsal scheduling fact | Dance |
Every UVU location beyond the Orem core
| Location | Status | Coordinates | Role & programs |
|---|---|---|---|
| Orem West Campus | current | 40.2802, -111.7293 | Health Professions and Utah National Guard; HP used as anchor point UVU locations, HP marker, NG marker |
| Vineyard / Geneva | current fields future academic | 40.3071, -111.7424 | Current soccer camps, intramurals and recreation; planned 225+ acre health, wellness and innovation campus with future CHPS and CHSS space UVU Fields, Vineyard master plan |
| Provo Airport | current | 40.2178, -111.7178 | Aviation Sciences plus Utah Fire & Rescue Academy and emergency-service training Aviation visit page, Utah Fire & Rescue Academy, UVU locations |
| Lehi / Thanksgiving Point | current | 40.4283, -111.8959 | TG and TH; MBA, MS Cybersecurity, general education, Dental Hygiene, Respiratory Therapy, Paramedic and Police Academy UVU locations, Dental, Respiratory Therapy, Police Academy |
| Wasatch Campus, Heber | current | 40.5455, -111.4136 | Associate degrees, Elementary Education, graduate education, community education and WARM hospitality/resort-management pathway Wasatch Campus, academics |
| Canyon Park (Bldg L) | current | 40.3232, -111.6794 | Culinary Arts Institute, Canyon Park Building L UVU locations, UVU marker |
| Capitol Reef Field Station | current | 38.1852, -111.1795 | Field research, creative work and engaged learning across disciplines; capacity is 24 overnight or 40 day visitors, not enrollment Field Station, plan a trip |
| Museum of Art, Lakemount | current | 40.2650, -111.7007 | UVU Museum of Art at Lakemount; public museum and visual-art academic resource Museum visit page, UVU marker |
| Payson (land held) | planned not operating | — | UVU owns 38.7 acres northeast of the I-15 Main Street interchange; construction schedule remains undetermined UVU land announcement, Vision 2030 |
| Eagle Mountain | partner or instructional point | 40.3800, -111.9719 | Current official marker; exact 2026 program and ownership scope not established UVU marker, staff guide |
| Salem | partner or instructional point | 40.0610, -111.6701 | Current official marker; exact 2026 program and ownership scope not established UVU marker, staff guide |
| Santaquin | partner or instructional point | 39.9753, -111.7911 | Current official marker; exact 2026 program and ownership scope not established UVU marker, staff guide |
| Spanish Fork | partner or instructional point | 40.1107, -111.6623 | Current official marker; exact 2026 program and ownership scope not established UVU marker, staff guide |
No current official source publishes campus-by-campus enrollment — UVU's Fall 2025 total of 48,669 is institution-wide and is not assigned to any location here fact. The Lehi program pages override the older CHPS location list (dental hygiene and respiratory therapy now teach at Thanksgiving Point).
The full building roster — all 53 pins
| Code | Building | Pin (lat, lon) | Zone |
|---|---|---|---|
| AM | UCAS Gym | 40.2838, -111.7180 | North |
| AS | Utah County Academy of Science | 40.2832, -111.7175 | North |
| AX | Auxiliary Annex | 40.2738, -111.7318 | West (outlying) |
| BA | Browning Administration | 40.2771, -111.7141 | Core |
| BB | UCCU Ballpark | 40.2766, -111.7170 | South |
| BC | NUVI Basketball Center | 40.2778, -111.7169 | Core |
| BR | Business Resource Center | 40.2737, -111.7153 | South |
| C2 | Continuing Education 2 | 40.2778, -111.7051 | East |
| CB | Clarke Building | 40.2818, -111.7175 | North |
| CS | Computer Science | 40.2790, -111.7110 | Core |
| DX | Digital Transformation | 40.2785, -111.7328 | West (outlying) |
| EC | UCCU Events Center | 40.2787, -111.7169 | Core |
| EN | Environmental Technology | 40.2777, -111.7145 | Core |
| FA | Faculty Annex | 40.2775, -111.7107 | East |
| FC | Facilities Complex | 40.2803, -111.7052 | East |
| FG | Brandon D. Fugal Gateway Building | 40.2771, -111.7135 | Core |
| FL | Fulton Library | 40.2810, -111.7164 | Core |
| GT | Gunther Technology | 40.2782, -111.7110 | Core |
| H2 | H2; no current expanded name published | 40.2748, -111.7083 | East |
| H3 | H3; no current expanded name published | 40.2747, -111.7078 | East |
| H4 | Army ROTC contested pin | 40.2805, -111.7136 | Core |
| H6 | Center for National Security Studies | 40.2802, -111.7093 | East |
| H7 | 1112 South 400 West | 40.2770, -111.7052 | East |
| H8 | Forensic Science Facility | 40.2766, -111.7052 | East |
| H9 | 1140 South 400 West | 40.2763, -111.7052 | East |
| H10 | GEAR UP | 40.2781, -111.7057 | East |
| H11 | TRIO Upward Bound | 40.2784, -111.7058 | East |
| H12 | Facilities South | 40.2751, -111.7066 | East |
| H13 | 1052 South 400 West | 40.2783, -111.7052 | East |
| H14 | Student Alumni | 40.2751, -111.7071 | East |
| H15 | Executive Events | 40.2749, -111.7071 | East |
| HF | Hall of Flags | 40.2779, -111.7144 | Core |
| HP | Health Professions | 40.2802, -111.7293 | West Campus |
| KB | Scott C. Keller Building | 40.2765, -111.7125 | Core |
| LA | Liberal Arts | 40.2801, -111.7167 | Core |
| LC | Losee Center | 40.2784, -111.7125 | Core |
| ME | McKay Education | 40.2832, -111.7185 | North |
| NB | Melisa Nellesen Center for Autism | 40.2831, -111.7195 | North |
| NC | The Noorda Center | 40.2777, -111.7091 | East |
| NG | National Guard | 40.2805, -111.7284 | West Campus |
| PS | Pope Science | 40.2780, -111.7150 | Core |
| RL | Rebecca D. Lockhart Arena | 40.2789, -111.7155 | Core |
| SA | Sparks Automotive | 40.2776, -111.7118 | Core |
| SB | Science Building | 40.2783, -111.7159 | Core |
| SC | Sorensen Center | 40.2786, -111.7143 | Core |
| SE | Scott M. Smith Engineering Building | 40.2761, -111.7073 | East |
| SL | Student Life and Wellness | 40.2795, -111.7150 | Core |
| SP | School Community University Partnership | 40.2845, -111.7237 | North |
| WB | Woodbury Building | 40.2774, -111.7131 | Core |
| WE | Wee Care Center | 40.2766, -111.7058 | East |
| WH | School of the Arts Warehouse / Scene Shop | 40.2868, -111.7286 | West (outlying) |
| WS | Wolverine Service Center | 40.2820, -111.7225 | West Edge |
| YA | Young Alumni | 40.2830, -111.7205 | North |
Source: UVU's live interactive campus map, accessed September 3, 2026, cross-checked against the Fall 2025 printable map and a 2023 building directory. Known name drift preserved in the working papers (Gunther "Technology" vs "Trades"; the CS building's pre-Smith label). One pin — H4, Army ROTC — sits ~880 m from its mapped footprint and is flagged rather than silently corrected fact.
Imagery rights
The map's aerial layer is NAIP imagery (USDA Farm Service Agency) served through the U.S. Geological Survey's National Map — U.S. public domain, with the credit line printed on the map page. UVU's own campus-map artwork is not an open layer (personal, non-commercial terms): it is used here as a factual reference for codes, names, and pin locations, and its artwork is not reproduced. Building-footprint data, if ever added, would come from OpenStreetMap under ODbL with its own credit fact.
Who at UVU is already doing it — the public record
Appendix J in five lines
- QuestionWho at Utah Valley University has publicly shown what they think or built with artificial intelligence?
- AnswerThe public record gives the program someone to nominate in every college, but it also shows trust and part-time access problems.
- Deciding numbersemployee use rose from 61% to 76%fact75% expressed cautionfact62% of instructors are part-timeest
- What the plan doesUse peer picks from every college, pay part-time teachers for required work, and keep a path that does not require artificial intelligence.
- Still unknownUVU does not say whether concurrent-enrollment teachers sit inside the part-time count; adjunct access between terms and a formal student-government view are also UNKNOWN.
Adoption is the sponsor's first concern, so this appendix answers a plain question from public sources only: who at UVU has already said, in public, what they think and what they've built? Every row is a professional statement its author published — a news interview, a faculty profile, a conference program, a Senate minute — with the link beside it. Nothing here infers a private view; "silent" means the public record is silent, not that a person has no opinion. Sweep date: September 3, 2026.
What leadership has said
| Person | Public statement | Signal | Where / when |
|---|---|---|---|
| Jon Anderson, President since August 10, 2026 University | “Artificial intelligence is a tool we are actively integrating to shape the future of education.” | Champion at prior institution fact unknown | 2025-11-20, Pittsburgh Quarterly; then PennWest president · Statement; UVU profile |
| F. Wayne Vaught, Provost and Senior VP Academic Affairs | “UVU is taking a leadership role in applied artificial intelligence.” | Champion fact | 2025-03-20, UVU NVIDIA announcement · UVU |
| Christina Baum, VP Digital Transformation/CIO Digital Transformation | “Identifying those faculty members and getting them to be champions is critical.” | Champion with governance focus fact | 2026-02-11, EdTech interview · EdTech |
| Spencer Magleby, Dean Smith Engineering & Technology | “Make sure that it’s not just cool, but that it locks in with people.” | Cautious champion; human-centered fact | 2026-06-19, AI health hackathon · TechBuzz |
| Sue Jackson, Dean Health & Public Service | no attributable public statement found | Publicly silent unknown | Checked 2026-09-03 · Role; targeted name/AI searches found no attributable statement |
| Steven Clark, Dean Humanities & Social Sciences | no attributable public statement found | Publicly silent; college itself is active unknown | Checked 2026-09-03 · Role; no attributable statement found |
| Daniel Horns, Dean Science | no attributable public statement found | Publicly silent unknown | Checked 2026-09-03 · Role; a public AI-project repost added no attributable words |
| Krista Ruggles, Interim Dean Education | “Artificial intelligence is transforming elementary STEM education, yet evidence remains fragmented.” | Champion with evidence and privacy guardrails fact | 2025-10-30, coauthored review · Paper; role |
| Courtney Davis, Dean Arts | “We are proud to provide education to our students that will put them at the forefront of arts technology.” | Champion fact | 2026-02-20, AI Design Awards announcement · UVU story; syndicated full quote |
| Bob Allen, Dean Woodbury Business | no attributable public statement found | Publicly silent; college faculty are highly active unknown | Checked 2026-09-03 · Role; no substantive attributable statement found |
Continuity is strong: the institute was launched under the previous president with the stance that everyone, regardless of major, should engage with the technology; the current president championed AI integration at his prior institution. Four deans are publicly silent while their colleges are visibly active — an invitation to bring them in, not evidence against.
Faculty champions and skeptics, by college
No named faculty member was found publicly advocating a blanket AI ban. The common design choice among the people who have built something — a revision coach that refuses to draft, a marketing "boss" that challenges strategy, virtual patients, a course bot with disclosure and verification controls — is guardrails, verification, and human judgment. That is also the design of this plan, which is why these people are its natural first cohort.
| Name | College | Public stance and evidence | Source |
|---|---|---|---|
| Anne Arendt Technology Management; associate dean at statement | Smith | Champion; urged preparation across fields and focus on opportunities fact | Date UNKNOWN, circa 2024, UVU Review |
| Jenny Nehring; Troy Taysom Information Systems & Technology | Smith | Champions; presented “ChatGPT–AI: Embracing Change” to educators fact | 2023-06-13, Nehring profile, Taysom profile |
| Armen Ilikchyan Technology Management/Mechatronics; Applied AI program leader | Smith | Champion with guardrails; uses AI-simulated student feedback for course improvement fact | 2025-02-27, profile |
| George Rudolph Computer Science chair | Smith | Champion with guardrails; moved from treating AI work as near-cheating to teaching multi-agent orchestration after employer feedback fact | 2026-08-10, TechBuzz |
| Majid Memari Computer Science | Smith | Champion/cautious; built AI-fluency modules and a grounded course chatbot with disclosure, verification, privacy, and autonomy controls fact | 2026-08-10, TechBuzz |
| Xi Chen Computer Science | Smith | Cautious integrator; “Coding Twice” has students code independently before using AI fact | 2026-02-19, SIGCSE |
| Zac Taylor Applied Engineering & Transportation | Smith | Champion with caution; supports responsible use and broader access to technical knowledge fact | 2024-09-11, UVU Review |
| Tyson Riskas Information Systems/Applied AI | Smith | Champion/cautious; presented on assessing IS learning in the GenAI era and teacher preparation fact est | 2024-07-25, profile |
| Noah Myers Accounting | Woodbury | Champion with guardrails; uses agents for forensic accounting after students learn manual foundations fact | 2024-08-29 and 2026-08-10, KSL, TechBuzz |
| Diego Alvarado-Karste Marketing | Woodbury | Champion with guardrails; built an AI “boss” that challenges strategy rather than writing the finished work fact | 2026-08-10, TechBuzz |
| Yang Huo Strategic Management/Operations | Woodbury | Champion; completed GenAI and ChatGPT training to update curricula and workforce skills fact | 2025-04-11, profile |
| Kari Olsen Accounting | Woodbury | Cautious/evaluative; coauthored empirical work on ChatGPT performance on accounting assessments fact | 2023, profile |
| Qianwen “Rachel” Bi Finance/Personal Financial Planning chair | Woodbury | Champion/cautious; teaches practical AI while raising credit, privacy, ethics, and general-AI questions fact | 2023-06 and 2024-05, UVU Summit, graduation program |
| Jacob Burdis Product Management/Marketing adjunct | Woodbury | Champion in professional work; co-founded an AI-supported education company fact | Date UNKNOWN, profile |
| Maritza Sotomayor Economics; UNESCO Chair associate director | Woodbury | Collective ethical-AI role is public; individual classroom position was not found unknown | 2025, UNESCO report |
| Krista Ruggles Elementary Education/STEM | Education | Champion with guardrails; co-developed a multi-agent classroom simulation with teacher feedback fact | 2026, i-ETC program |
| Uzeyir “Adam” Ogurlu Elementary Education | Education | Cautious/champion; research finds openness to training alongside plagiarism, overreliance, reasoning, and authenticity concerns fact | 2023-12-08, RESSAT |
| Bing Han Elementary Education | Education | Named on the UNESCO AI and Education Committee; personal stance not separately attributable unknown | 2025, UNESCO report |
| Angie McKinnon Carter English | CHSS | Cautious adopter; built a revision coach that refuses to draft, retains human meetings, and permits opt-outs fact | 2026-08-10, TechBuzz |
| Yi Yin Sociology/Behavioral Science | CHSS | Champion with opt-outs; uses AI for course materials and comparative research while supporting informed refusal fact | 2024–26, profile, TechBuzz |
| Christa Albrecht-Crane English; Writing Program chair | CHSS | Strongest public skeptic; argues NotebookLM’s frictionless compression can conflict with writing pedagogy and educational values fact | 2025-10-24, SIGDOC paper |
| Tom Henry English | CHSS | Cautious/mixed; separates useful writing support from harmful substitution fact | 2025, Utah English Journal |
| Shane Smith Philosophy adjunct | CHSS | Pragmatic champion; opposes leaving students unprepared but warns that dependence can remove learning fact | 2024-09-11, UVU Review |
| Brian Whaley English and Literature chair | CHSS | Cautious/champion; supports experimentation and says policing-only responses waste effort fact | Circa 2024, UVU Review |
| Kelsey Hixson-Bowles English | CHSS | Champion/researcher; studies faculty and student perceptions to guide classroom practice fact | 2025–26, profile |
| Devin Gilbert Languages & Cultures | CHSS | Cautious integrator; teaches bounded use of machine translation and GenAI fact | 2024-11-20, profile |
| Acacia Overono Psychology | CHSS | Cautious/mixed; promotes disclosure and critical review of AI’s classroom and social effects fact | 2024–26, profile |
| Chris Weigel Philosophy; Center for the Study of Ethics | CHSS | Cautious; teaches AI ethics and publicly led “Wrestling With AI Policy at UVU” fact | 2025-10-02, Ethics Week |
| Eunmi Joung Mathematics | Science | Champion with guardrails; found that critiquing ChatGPT answers can deepen mathematical reasoning fact | 2025-09-19, study |
| Roxanne Brinkerhoff, Inyoung Lee, Ka Lun Wong, Ofa Ioane, Eunmi Joung Mathematics | Science | Study reports moderate openness, low confidence, changed teaching, and demand for responsible-use support fact est | 2026-03-15, EJMSTE |
| Britt Wyatt Biology | Science | Engaged/champion; presented on adapting practical science-learning tools for a GenAI environment fact est | 2025-02-28, profile |
| Jill Johnson Nursing | Health & Public Service | Champion after initial caution; used AI Academy learning to create five interactive virtual-patient chatbots fact | 2024–25, profile |
| Hsiu-Chin Chen Nursing | Health & Public Service | Champion; joined the 2026 AI Fluency Summer Institute to improve teaching and research fact | 2026-05-05, profile |
| John Fisher Emergency Services/Public Administration | Health & Public Service | Mixed; publicly covers AI’s benefits, competence, access, cost, ethics, and risks fact est | 2026-04-09, profile |
| Merilee Larsen Health Sciences chair | Health & Public Service | Cautious/champion; participated in health/public-safety ethics and UVU policy panels fact | 2025-10-02, Ethics Week |
| Brandon Truscott Entertainment Design | Arts | Mixed; presented “The Quick and the Dread: Threats and Benefits of AI in Teaching” fact est | 2024-08-14, profile |
| Shirin Abedinirad Art & Design, Sculpture/Ceramics | Arts | Publicly engaged through an “AI and Creativity” panel; exact position not recoverable unknown | 2025-09-30, Ethics Week |
| Cheung Chau Music | Arts | UNESCO executive-board participation is public; no individual AI teaching statement was found unknown | 2025–26, UNESCO |
About 37 faculty joined the May 2026 AI Summer Institute; six are publicly profiled. The behavioral program in the guide (§07) starts from peer nomination, not this list — but the list shows every college already has someone to nominate.
The friction people describe in public — and what the plan does about each
| Friction (public evidence) | What it calls for — and where the plan answers it | Grade / date |
|---|---|---|
| Copyleaks false positives reached Faculty Senate “CopyLeaks has been reporting false positives regarding AI generated content.” Minutes | Do not use detector output as sole proof; require human review and process evidence | fact 2025-04-29 |
| An ENGL 2010 syllabus scanned every paper while warning students to keep version history against false flags Syllabus | Consistent due process and clear evidence standards | fact Spring 2025 |
| UVU says assignment, course, department, and university AI rules may conflict Student guidance | A simple baseline policy, with course-specific additions | fact Accessed 2026-09-03 |
| Gateway is employee-only; free usage resets weekly at a $5 allowance; heavier use needs a paid tenant UVU employee page, TechBuzz | Student access plan; published quota behavior; department path for heavier use | fact 2026-07-14 |
| Wilson is in selected courses and carries a hallucination warning; no audited public error rate Wilson, EdTech | Publish accuracy, outcome, accessibility, and rollout tests | fact unknown 2025–26 |
| Employee AI use rose from 61% to 76%, yet 75% expressed caution; “very satisfied” fell from 23% to 15% Board packet | Trust-building and measured outcomes, not just activation counts | fact 2026-01-29 |
| Faculty’s largest training request was ethics/responsible use/academic integrity, 34%; staff asked for ethics/privacy/policy, 26% Board packet | Discipline-specific training plus usable policy examples | fact 2026-01-29 |
| Diego Alvarado-Karste abandoned a standalone build because institutional approval was unrealistic inside the two-week institute TechBuzz | A preapproved sandbox, reusable components, and fast review | fact 2026-08-10 |
| Policy 441 AI entered Stage 1 with a June 2026 Board target, but it is absent from the current approved manual Executive summary, current manual | Publish current stage, operative rules, and relation to student academic use | unknown 2025-09-25 to 2026-09-03 |
| Optional adjunct pedagogy/technology workshops are unpaid; appointments end each semester Policy 639 | Paid AI training and account continuity for active/returning adjuncts | fact est Current |
| No UVU-specific public accessibility complaint or AI-product accessibility evaluation was located UVU AI guidelines | Publish accessibility testing and feedback results; absence of complaints is not proof of accessibility | unknown Checked 2026-09-03 |
The most telling pair: employee AI use rose from 61% to 76% in a year while "very satisfied" fell from 23% to 15% and 75% expressed caution. Access is not the bottleneck; trust and fit are. That is the case for a service designed around guardrails, predictable quotas, a real non-AI path, and no detector-as-proof — the rights charter in §07.
Student voice
| What students said | The ask it implies | Grade |
|---|---|---|
| Computer Science student Jacob Barrus found AI useful for explanations but warned, “if you rely too heavily on AI, you don’t learn.” English student Salem Kimball worried about lost investment in learning. 2024-02-20, UVU Review | Permit help without replacing the student’s thinking | fact |
| Poster reported 27% Copyleaks, 4% Grammarly, and 10% Turnitin on the same claimed self-written paper. 2024-11-10, r/UVU | Do not treat one detector score as proof | fact |
| Commenters described supplying edit history, waiting for responses, and changing writing style to avoid flags. 2024-11-11, r/UVU | Prompt evidence review and a clear appeal path | fact |
| Anonymous respondents reported false flags and preferred human grading and transparent rules. More than one-third reportedly avoided AI for ethical or environmental reasons; sample details were not published. 2025-11-20, UVU Review | Human review, disclosure rules, and a valid non-AI path | fact unknown |
| Named student panels addressed AI and creativity and the future of AI. Public schedule proves participation; individual positions are UNKNOWN without transcripts. 2025-09-30 to 10-03, Ethics Week | Continue structured student forums and publish their recommendations | fact |
| Student preference survey: ChatGPT 71%, other tools 13%, Gemini 9%, Copilot 7%. Common uses included research/questions, concept explanation, ideas, and homework help. 2026-01-29, Board packet | A student service must support the tools they already use or give a clearly better equivalent | fact |
| Student-reporting synthesis raised employment, environment, mental-health, training-data, and artists’ rights concerns alongside educational value. 2026-02-18, UVU Review | Teach social and rights impacts, not only prompting | fact |
| CHSS and the Institute held “Student Voices on AI” and faculty/student-perception sessions. 2026-03-04 and 03-27, Yi Yin profile | Turn these forums into published design requirements and outcome measures | fact |
| No public UVUSA AI resolution, Copyleaks discussion, Policy 441 position, or recorded AI debate was found. Minutes checked through 2026-03-19, UVUSA minutes | Obtain and publish a formal student-government position before broad student rollout | unknown |
Adjunct facts
Most UVU instructors work part-time. The public record on their AI access is thin, and thin in a specific way:
| Evidence | What it means | Grade |
|---|---|---|
| Shane Smith, adjunct Philosophy professor, publicly discussed AI proof limits and estimated that an AI pre-grader could reduce about eight grading hours to three. UVU Review, 2024-09-11 | Adjunct workload is a direct adoption use case, but automated grading also raises due-process risk. | fact |
| UVU’s AI Task Force charge expressly includes training and support for full- and part-time faculty. Task Force charge | Adjunct inclusion exists at the intent level. | fact |
| The Board reports 120 AI Academy completers and 22 advanced participants, with no employment-class breakdown. Board packet | Whether any adjunct completed the Academy is UNKNOWN. | fact unknown |
| Public AI Academy pages do not state adjunct eligibility, compensation, selection rules, or account continuity. | The phrase “for faculty” is not enough to prove practical adjunct access. | unknown |
| The 2026 Faculty Summer Institute offered selected faculty a $1,500 completion stipend, but did not name or exclude adjuncts. Institute, AI brief | Eligibility by faculty type remains UNKNOWN. | fact unknown |
| Gateway is available to “all employees.” Policy 639 ends adjunct employment each semester. Employee page, Policy 639 | Active adjunct access is a reasonable EST; access before, after, or between appointments is UNKNOWN. | fact est |
| Required adjunct pedagogy orientations are paid hourly; optional pedagogy and technology workshops are unpaid. Policy 639 | Optional AI training competes directly with unpaid time. | fact |
| Published Faculty Senate eligibility is limited to salaried, benefits-eligible faculty; adjunct welfare is a Senate purpose, but adjuncts cannot serve as senators. Constitution | AI governance lacks a clearly published direct adjunct seat. | fact |
| No “UVU Adjunct Association” or AI-specific statement from the public UVU AAUP/AFT chapter was found. UVU AAUP/AFT | A public adjunct AI voice or advocacy channel was not located. | unknown |
| UVU’s own Common Data Set counts 1,310 part-time and 802 full-time instructional faculty in fall 2025. That works out to 62.0% part-time. UVU People & Culture publishes the same two counts. Common Data Set 2025–26, §I-1, People & Culture | The 62% share is confirmed at a UVU source. It counts people, not full-time jobs, and UVU does not say here whether every part-time teacher is an adjunct or whether concurrent-enrollment teachers are inside the count. | fact counts · est share |
| The narrower count. IPEDS reports 937 part-time and 809 full-time UVU instructional staff for fall 2024, which is 53.7% part-time. IPEDS counts only staff whose main job is teaching, and leaves out graduate assistants, so its list is shorter than the Common Data Set’s. IPEDS, UVU staff table | A reader who finds 53.7% somewhere else is looking at a different year and a narrower list of people, not a different answer. | fact counts · est share |
Service implication: paid, short, asynchronous training; accounts and course agents that survive the semester boundary; eligibility stated in plain language; an adjunct advisory seat before any policy or detector rule changes. The plan's artifact-gated stipend and adjunct-first design (§07) already assume this — the public record confirms why.
Does the public voice change the sequencing?
| College | Appendix A rating | Verdict after the public-voice sweep | Reason |
|---|---|---|---|
| Smith Engineering & Technology | HIGH | Stay HIGH | Widest technical roster, several named classroom uses, department-chair leadership, coding-tool evidence, applied programs, and explicit employer pull. |
| Woodbury Business | HIGH | Stay HIGH | Myers, Alvarado-Karste, Huo, Bi, Olsen, and others show practical use across accounting, marketing, finance, management, and assessment. |
| School of Education | HIGH | Stay HIGH | Interim dean Ruggles is a named AI researcher; Ogurlu supplies cautious evidence; AI Academy completions and a multi-department panel show breadth. Named classroom implementations remain fewer than Smith or Woodbury. |
| Humanities & Social Sciences | HIGH | Stay HIGH | This is the broadest public conversation: writing tools, source verification, sociology research, translation, disclosure, ethics, opt-outs, and the strongest skeptical voice. Caution is evidence of active engagement, not low receptivity. |
| College of Science | MED | Stay MED, strengthening | Mathematics research, biology teaching work, the science panel, and DIFFRAX show activity, but named public classroom deployment remains concentrated. |
| Health & Public Service | MED | Move to HIGH, controlled | Jill Johnson’s five virtual-patient bots, Chen’s Summer Institute work, Fisher’s adoption talk, and a panel spanning health, forensics, PA, and emergency services fill x24’s local-evidence gap. Start with synthetic cases and human review. |
| School of the Arts | MED | Stay MED, low confidence | Dean Davis supports arts technology; Truscott and Abedinirad show public engagement. However, there is still little public evidence of a repeatable classroom program, policy, or broadly used tool. |
One change carried into Appendix A: Health & Public Service moves to HIGH (controlled) — named nursing champions with built virtual-patient bots filled the local-evidence gap. The controls do not move: synthetic cases first, human review always.
What we plug into — and the five questions that remain
Appendix K in five lines
- QuestionWhat campus systems must the service join, and what must Utah Valley University tell us?
- AnswerPublic records confirm the main campus systems, but five questions that only the university can answer still block the connection.
- Deciding numbers360-person Gateway pilotfactabout $5 of tokens per employee per weekfact44 courses and 2,000 studentsfact
- What the plan doesBring its own computers, user page, and support, then use only approved sign-in, course, data, network, and support links.
- Still unknownThe five questions cover sign-in, Gateway and Wilson design, data rules, the network handoff, and operating owners; no login was attempted.
The service stands on its own; this appendix maps the house it stands in. Every surface below was established from public sources — UVU's own pages, job postings, policy documents, network registries, and a handful of passive checks any browser performs on first visit — so that the questions we finally ask UVU are few, precise, and cannot be answered from outside. Sweep date: September 3, 2026.
Integration surfaces
| Surface | What a self-contained service integrates with | Evidence |
|---|---|---|
| Microsoft Entra / Azure AD | Institutional SSO, user identity, group or role claims, MFA, lifecycle controls | Passwordless setup expressly requires Azure AD synchronization. UVU’s Microsoft job owns Entra, tenant management, Conditional Access, Defender, and DLP. Passwordless guide, job posting |
| MFA | UVU’s current page uses Microsoft MFA for myUVU/Microsoft apps and also documents Duo for some flows. | MFA page, Atlassian support flow. Exact endpoint-by-endpoint chain is UNKNOWN. |
| Canvas | LTI 1.3 placement, deep links, context/role claims, approved APIs, test courses, accessibility review | Canvas is UVU’s official LMS. UVU requires legal, stewardship, purchasing, accessibility, and systems review for integrations. Canvas, technology review |
| Banner | Authoritative enrollments and course roles; preferably consumed through the approved Canvas/Banner or API path | UVU says Banner automatically adds students and governs teacher/assistant roles in Canvas. Canvas roles |
| Banner version | Banner 9 Self-Service is the stated replacement for retired SSB8. Core Banner/Ethos/API versions are not public. | SSB8 retirement |
| Existing Canvas LTIs | Avoid duplicate tool placement and account for content/data flows | UEN’s April 2024–April 2025 report gives UVU 60,735 unique Canvas users and lists Proctorio, Kaltura, Watermark Insights, McGraw Hill Campus, and Cengage among its top LTIs. UEN Canvas report |
| Teams LTI | Evidence that UVU supports roster-aware Microsoft/Canvas integration patterns | Instructure reports Teams/Canvas schedule and roster synchronization at UVU. Case study, Microsoft 365/Canvas article |
| Pathify/myUVU | Portal card/deep link, notifications, service discovery, possibly SSO | Mobile moved to Pathify in April 2026. Desktop myUVU is scheduled for November 16, 2026. UVU mobile, FAQ |
| API layer | Private UVU API gateway, data lake and governed data products | The Cloud Data Engineer role says “Own API Management space at UVU.” UVU’s public tool list shows Banner AppNav, IDM Person Sync, and AD tools, many gated to campus/VPN. No public developer portal was found. Job, tools |
| IT service management | Incident, request, change, and service catalog integration | UVU Status confirms Atlassian/Jira Service Management. The SRE job confirms PagerDuty on-call. ServiceNow appears only as a possible skill. Status, SRE job |
| Monitoring/security | Splunk route, UVU monitoring, endpoint protection, security review | Internet2 lists UVU as a Splunk subscriber. CrowdStrike Falcon with OverWatch is the stated default on faculty/staff computers. Splunk, CrowdStrike |
| Data governance | Data classification, steward approval, minimum-access personas, architecture review, data catalog | Policy 445, effective 2024-03-28 |
| Privacy | Approved purpose, authorized storage, collection notice, third-party disclosure | Policy 446, effective 2023-05-09 |
| Security | Vendor review, MFA, least privilege, encryption, logging, inventory, tested changes, network controls | Policy 447, effective/reviewed 2026-08-10 |
| Records | Prompt, output, support, audit, roster-cache, and analytics schedules; GRAMA/legal-hold handling | UVU says chat may be logged and records may be subject to GRAMA. Retention follows record content, schedules, and holds—not a universal deletion period. Privacy statement, Policy 451 |
| AI policy | No approved Policy 441 appeared in the current manual at cutoff. A 2025 AI-policy proposal existed, but its PDF now returns 404. | Policy manual, pipeline, Library AI guide |
| Identity tenant (probed) | Microsoft Entra ID, cloud-managed ("Managed" namespace, no on-premises federation) — a standard app registration is all our sign-on needs | Public tenant-discovery endpoints, 2026-09-03 |
| AI Gateway hosting (probed) | Both Gateway hostnames are published through Microsoft Entra Application Proxy — a privately hosted app, pre-authenticated by the tenant, two separate app registrations. Our hook to it is an API endpoint its backend can call; we never sit inside it | Public DNS + one unauthenticated request per hostname, 2026-09-03 |
| Wilson hosting (probed) | Public Wilson answers from a Cloudflare-fronted host whose page title names a server node — self-hosted, node-named, consistent with the vendor-platform finding above | Public DNS + page title, 2026-09-03 |
| Campus address space (probed) | UVU owns its own public /16 block; core portals and VPN answer from it — an on-campus data center with public routing exists, so a rack for the fleet is a facilities conversation, not a hosting-model one | Public DNS + ARIN registry, 2026-09-03 |
| Advertised model endpoints (probed) | None — no public "llm", "ollama", "copilot", or similar hostnames exist. Nothing to collide with | Public DNS, 2026-09-03 |
Reading it together: a Microsoft-shaped house (the CIO's own words: "We're kind of a Microsoft shop. We have an A5 license"), Canvas with Banner-fed roles, a portal moving to Pathify this fall, Jira Service Management with PagerDuty on-call, CrowdStrike on managed devices, Splunk available, and Policies 445/446/447 as the written rulebook. Everything our design already assumed — now sourced.
The house's compute and network, as publicly known
| Area | Public fact | What it means for us | Grade |
|---|---|---|---|
| UVU internet uplink | UEN increased UVU internet capacity by a factor of five within ten days, completing the change March 18, 2025. Absolute pre/post capacity was not given. UEN | Size external model traffic and updates against a private bandwidth figure; do not interpret “5×” as a line rate. | fact |
| Research network | Internet2 lists UVU as a primary participant through UETN since 2018-01-01. Internet2 | UETN/Internet2 is an available institutional network path, but private routing and peering details are still needed. | fact |
| Cloud/service procurement | Internet2 currently lists UVU as a subscriber to AWS, Canvas, Pathify, and Splunk. | These are real institutional relationships and possible contract routes. They do not prove workload location, editions, spend, or Direct Connect. | fact |
| UVU-owned HPC | Three UVU-owned Notchpeak nodes at University of Utah CHPC: one 96-core/380 GB node and two 64-core/150 GB nodes. Total: 224 cores and 680 GB RAM. UVU accounts receive priority. No GPUs are listed. UVU HPC | Useful research CPU capacity exists, but it is not a proved production hosting surface. | fact |
| One-U H200 nodes | CHPC documents ten nodes, each with eight H200 GPUs, 96 CPU cores, and 2 TB RAM. It says all CHPC users can use guest access and AI researchers may request priority. CHPC | EST a UVU CHPC account may submit guest work. UVU entitlement, priority, protected-data approval, and service hosting rights remain unproved. | fact unknown |
| Redtail | Redtail is a statewide, University of Utah-managed AI supercomputer. No public UVU allocation, proposal, account, workload, or Gateway/Wilson connection was found. CHPC, funding account | Do not count Redtail as UVU capacity. | fact unknown |
| Smith Engineering Building | Opened January 22, 2026; nearly 200,000 square feet. Public descriptions name an AI lab, VR functions, and a machine-learning/autonomous-systems drone lab. No server or GPU specifications are published. UVU | Building and labs are real; production rack, power, cooling, network, and GPU capacity are private. | fact |
| SCaFL | Indexed/archived material describes a donated 10 Gb UTOPIA connection and local servers for a public/business lab. The current SCaFL pages and old impact report returned 404. | It is not evidence of UVU’s campus backbone or current service capacity. | HISTORICAL · C |
| Campus data center | UVU’s design guide follows data-center and server-room standards, including TIA-942 considerations, power, cooling, security, and growth space. It does not reveal an operating data-center location or capacity. Telecom guide | Rack location, available power/cooling, segmentation, and remote-hands service remain private. | fact unknown |
| Jamf | UVU’s page says Jamf handles Apple inventory, deployment, updates, encryption, and security. It says university-owned Apple devices were to be enrolled, but gives no count. Jamf | A managed Mac client may need Jamf packaging; this does not provide usable hardware inventory. | fact |
| Citrix/MyApps | UVU publicly lists Citrix-delivered MyApps and many student applications. No GPU-backed VDI profile, vGPU type, host count, or capacity was found. Software inventory | Browser-based delivery is safer than assuming a GPU-capable VDI client. | fact unknown |
Gateway and Wilson: what is actually known
Treat the employee Gateway, the public-website Wilson, and the in-course Wilson as three separate services until UVU shares an as-built diagram. The only safe shared assumption is Microsoft identity and governed UVU content — not a shared model layer. Notably, the public-website Wilson runs on a vendor platform inside UVU's own cloud tenant with a named UVU engineer operating it; whether the Canvas Wilson shares that backend is not public.
| Component | Public evidence | Grade |
|---|---|---|
| Gateway front door | UVU says every employee has a free, secure platform for approved AI tools. The public endpoint invokes Microsoft sign-in, but no login was attempted. UVU employee page | fact |
| Gateway launch and population | Reported employee launch July 1, 2026, after a 360-person pilot. Week-one registered users rose 24%; active users rose 31%. TechBuzz | fact |
| Gateway model families | Reported access to ChatGPT, Claude, Gemini, and locally hosted open-source models. Exact providers, versions, endpoints, and local models are UNKNOWN. TechBuzz | fact |
| Gateway functions | Multi-model comparison, projects/folders, personal knowledge bases, prompt chains, and agent creation/export are publicly reported. | fact |
| Gateway implementation | UVU’s Tyler Small said it used “FERPA-secure APIs and off-the-shelf templates.” TechBuzz describes the result as assembled in-house. | fact |
| Gateway allowance | About $5 of tokens per employee per week, resetting Monday; a paid tenant was reported for heavier use. Actual model rates and spend are private. | fact |
| Gateway student rollout | July reporting said students were next. No public student launch date, eligibility rule, or budget was found. | unknown |
| Gateway stack | Searches found no reliable tie to LibreChat, Open WebUI, LiteLLM, Portkey, Azure OpenAI, Bedrock, or any named router/database. “Locally hosted” does not identify where. | unknown |
| Gateway data terms | UVU’s employee page describes protective Microsoft terms, but those terms are not expressly the Gateway’s retention, logging, or provider-training contract. | unknown |
| Wilson product surface | UVU calls Wilson a UVU-developed assistant spanning its public website and selected Canvas courses. Course access excludes grades, tests, and private files; UVU says it does not store personal information. Wilson AI | fact |
| Wilson rollout | 14 classes in 2024; 26 courses and 1,800+ students in fall 2025; 44 courses and 2,000 students by February 11, 2026. No later count was found. PBA, board minutes, EdTech | fact |
| Wilson retrieval | Course Wilson is grounded in permitted course content. Baum specifically named Kaltura recordings and links to relevant points in a recording. The vector store, embeddings, chunking, reranking, and indexing cadence are private. | fact unknown |
| Public Wilson platform | Cloudforce names the public Wilson as a “nebulaONE-powered conversational agent” and says Daniel Razo “built it, supports it, manages it, and continually improves it.” Cloudforce | fact |
| Canvas Wilson relationship | No source proves whether public Wilson and the 44-course Canvas Wilson share a tenant, agent, retrieval store, codebase, or model route. | unknown |
| nebulaONE defaults | Current Cloudforce terms say nebulaONE ordinarily processes inside the client’s Azure tenant and does not train models on client data. Its privacy page describes anonymized telemetry retention. These are product defaults—not UVU’s executed terms or configuration. Terms, privacy | est |
| Wilson model/provider | nebulaONE documentation names Microsoft Foundry/Cognitive Services and gives Azure OpenAI and Claude as examples. No public source identifies UVU’s selected models. | unknown |
| Legacy Wilson | Ocelot powered the 2020 Wilson. UVU announced a replacement/transformation effective May 1, 2023. Ocelot is not evidence of the current stack. Ocelot, UVU transformation page | fact |
| Wilson roadmap | February 2026 plans included connecting course agents and bringing universitywide answers into course Wilson. Public evidence does not establish completion. | fact unknown |
The twelve-item data request, after the public sweep
None of the twelve could be fully closed from outside — each bundles public facts with private contracts or configuration — but most no longer need answering before we build. The service brings its own compute, front door, and operations; only the interface details block integration.
| # | Item | State | Closed by public evidence | What stays private — and whether it blocks us |
|---|---|---|---|---|
| 1 | Software and spend | unknown | Named public platforms now include M365/Entra, Canvas, Pathify, AWS, Splunk, Jira, CrowdStrike, Kaltura, and others. | Full SKU/seat/spend/renewal/reseller ledger. Does not block a self-contained service except where a contract governs an integration. |
| 2 | Microsoft | unknown | A5 environment, M365 tenant operations, Entra, Conditional Access, DLP, Defender, Power Platform, Graph direction. | Exact population mix, Copilot/ChatGPT SKUs, Azure OpenAI/Copilot Studio rights, $600,000 request outcome. Only identity/application registration blocks integration. |
| 3 | Gateway | unknown | Model families, employee rollout, pilot size, allowance, features, in-house/template build, FERPA-secure API claim, students-next direction. | Exact software, models, providers, router, hosting, APIs, connectors, spend, retention, logs, and training terms. Blocks coexistence/routing decisions, not independent hardware. |
| 4 | Wilson | unknown | Public Wilson is nebulaONE; Daniel Razo operates it. Canvas/Kaltura grounding and 44-course/2,000-student latest count are public. | Whether web/course products share a backend; model/provider, RAG, LTI/API, hosting, retention, cost, current count, assessment, oversight. Blocks overlap and Canvas design. |
| 5 | Canvas | unknown | Official LMS, Banner-fed roles, approval process, major LTIs, Teams sync pattern, Internet2 subscription. | Core/Plus/Next edition, IgniteAI flags, current Copyleaks state, contract/data terms, exact LTI/API configuration. Interface fields block integration; commercial tier mostly does not. |
| 6 | Compute | unknown | Three Notchpeak nodes: 224 CPU cores/680 GB; statewide One-U and Redtail resources described. | Current status/use, GPU entitlement, allocations, protected-data approval. Does not block because the planned service brings its own compute. |
| 7 | Campus hardware | unknown | Jamf function, MyApps/Citrix, Smith labs, SCaFL historical 10 Gb link, and room purposes. | Counts, GPU/server models, Citrix profiles, rack/power/cooling, operational state. Only the proposed service’s facility/network envelope blocks us. |
| 8 | NVIDIA | unknown | Three-year, no-funds collaboration; training, workshops, tools, cloud workstations, ambassador pathway. | Executed MOU, active deliverables, credits, equipment, usage. Does not block the self-contained service. |
| 9 | People and funding | unknown | Kahlert $5.2 million; $2 million one-time recommendation; AI/Dx reinvestment figures; public program counts from x29. | Full rosters, outcomes, grant ledger, placements, gift restrictions. Useful diligence, not a technical blocker. |
| 10 | Data and agreements | unknown | UVU Data Lake, Azure/Synapse/API Management, hybrid environment, governance/security/privacy policies. | As-built schemas, system-of-record map, current tools, API contracts, executed MOUs/DPAs/DUAs, retention and subprocessors. Data/interface parts block integration. |
| 11 | State and UNESCO | unknown | State task force/credential intent, Redtail facts, UNESCO chair and committee are public from x29. | Credential launch/platform/enrollment/RFP; chair funding, executed agreements, live projects. Not a core infrastructure blocker. |
| 12 | Payment reconciliation | STILL PRIVATE | No reliable FY2024–FY2026 provider/reseller payment export was completed. | Vendor payments, reseller routes, contracts and ledger reconciliation. Useful diligence; not needed to size our own service. |
The five questions that genuinely require UVU
Asked once, together, of the technical side of the house:
- Identity and application interface pack: What Entra registration, protocol, issuer/tenant, claims, groups, MFA/Conditional Access rules, test accounts, Canvas LTI 1.3 deployment details, Canvas API scopes, Banner-fed role source, Pathify placement, and UVU API Management route must the service use?
- Gateway/Wilson as-built boundary: Please provide one current data-flow diagram for Gateway, public Wilson, and course Wilson showing platform/version, hosting, models/providers, routing, retrieval sources, Canvas/Kaltura interfaces, and whether any components or stores are shared. Must our service hand off through any of them?
- Approved data and control envelope: Which Policy 445 data classes may enter the service; what are the required retention, deletion, legal-hold, logging, no-training, subprocessor, region, encryption, accessibility, privacy-notice, security-review, and final Policy 441 requirements?
- Network and facilities handoff: For owner-supplied hardware, what rack location, power/cooling, physical access, VLANs, firewall/egress rules, DNS/TLS, load-balancing, UETN/Internet2 path, private-cloud links, remote administration, monitoring, backup, and disaster-recovery rules apply?
- Operating contract and rollout: Who owns product approval, data stewardship, Canvas administration, security, networking, support, change control, incident response/JSM/PagerDuty, and service-level decisions; what pilot users/courses and current Wilson/Gateway rollout constraints must we plan around?
Everything else in the twelve-item ledger is commercial diligence that can be requested later as an export, or never. The four decisions that belong to Dr. Burns rather than to IT stay in Appendix H: whether a Moonshot or research-fund notice of intent was filed, whether a budget number exists or the dial stays the instrument, which college sponsors the pilot and how Dx co-sponsors, and whether a seat quote under $30 per user per year should flip the plan automatically.
The free path, honestly
Appendix L in five lines
- QuestionCan free online model access serve all 36,600 people at Utah Valley University?
- AnswerNo; three measurable offers cover only 1.38% of a normal day when run from university accounts.
- Deciding numbers78,690 turns per day neededest1,082 turns per day suppliedest1.38% coveredest
- What the plan doesUse these offers only for public or made-up data, tests, and a separate sandbox; keep the owned local service as the floor.
- Still unknownMany rate limits, retention terms, and end dates are UNKNOWN; no account was made and no request was sent.
A fair challenge was put to this plan: before buying anything, have we exhausted the free options? Free tiers exist from a dozen providers, and as of this week Meta's newest model is free through a coding gateway. This appendix takes that path seriously, on the providers' own published terms, and shows exactly where it holds and where it breaks for 36,000 people. Verified directly on September 3, 2026, then extended by a full provider survey the same day.
Meta Muse Spark 1.3 — the fact sheet
| Question | Answer | Grade |
|---|---|---|
| What is it? | Meta's newest frontier model, released September 2, 2026 — trained for long-horizon, agentic work; 1M-token context; text, image, and video input. | fact |
| Is it open-weight — can a campus run it on its own machines? | No. It is API-only today. Meta's own announcement lists an open-weights release on its roadmap with no date. Until weights ship, it cannot be part of an owned floor. | fact |
| How good is it? | Artificial Analysis Intelligence Index: 61 (xhigh) / 62 (max) — level with GPT-5.6 Sol (61) and Grok 4.6 (61), just under Claude Opus 5 (63), under Claude Fable 5.1 (66), and above the best open-weight model, GLM-5.3 (60). Meta's coding claims hold on two benchmarks: it ties Sol on Terminal-Bench 2.1 (88.8) and leads both Opus 5 and Sol on DeepSWE (75.4 vs 74.0 / 73.0); it trails Opus 5 on knowledge-work and computer-use benchmarks. | fact |
| What does it cost? | Standard tier: $1.25 per million input tokens, $4.25 output (cache hits $0.15). Roughly 40% cheaper per benchmark task than GPT-5.6 Sol at the same score. | fact |
| Where is it "free"? | Through the contributor tier: $0.10 / $0.20 per million tokens on Meta's platform, and $0 on OpenCode Zen "for a limited time" — in both cases in exchange for permission to use prompts and completions to train future Meta models ("used to improve our products"). | fact |
| What is the catch for a university? | The contributor tier is a data-for-discount trade. Every prompt becomes training data, which rules it out for anything touching student records, unpublished research, licensed course material, or personal data — the Sensitive and Restricted classes in UVU's own policy. The free window has no stated end date, and the offer is aimed at individual developers; there is no institutional, education, or admin-controlled program published. | fact terms · est fit |
| Where does it fit this plan? | As a candidate for the metered frontier pool under the standard (no-training) tier — a strong, cheaper coding and agent route beside Gemini Flash, Sol, and Fable — and as a legitimate free sandbox for CS students working on their own code with no UVU data. Not as a floor, and never on the contributor tier for real work. | est |
OpenRouter's free models — the limits, from the source
| Term | What OpenRouter states | Grade |
|---|---|---|
| Free-model rate limit | 20 requests per minute and 50 requests per day on an account that has never bought credits; 1,000 requests per day once at least $10 of credits has ever been purchased. Paid models carry no platform-level request cap. | fact |
| Can a campus multiply accounts? | No — by design and by contract. Docs: "Making additional accounts or API keys will not affect your rate limits, as we govern capacity globally." Terms (updated August 31, 2026) prohibit creating "multiple accounts as a single user, for purposes of bypassing or circumventing use limits" and reselling API access. | fact |
| What happens to the data? | "Some Models may store or train on your Inputs for improving their own large language models" per each model's terms; OpenRouter has opted out "where possible." On free models the upstream provider's terms govern — often training-permitted. | fact |
| Stability | Free models rotate in and out without notice and are throttled hardest at peak hours (community reports across 2025–26); the free-tier limits themselves were tightened in 2025. | est |
The arithmetic at university scale
| Scenario | Math | Verdict |
|---|---|---|
| Phase-1 daily demand | 36,629 people × ~2.15 prompts per person per day (whole-population campus telemetry, §06) ≈ 78,800 prompts per day; exam-week stress ≈ 9× the base concurrency. | The bar any "free" plan must clear. |
| One institutional OpenRouter account, funded | 1,000 requests per day ÷ 78,800 ≈ 1.3% of daily demand; 20 per minute ≈ 0.33 concurrent streams against a base need of ~53. | A sandbox, not a service. |
| One free account per person | 50 per day each would technically exceed 2.15 per person — but the terms forbid organized multiplication of accounts, capacity is "governed globally," data terms are per-user consumer terms, and there is no admin control, audit, budget, or SSO. | Not permitted, not governable, not private. |
| Muse Spark contributor tier at scale | Zero dollars — and every prompt trains Meta's models. At 78,800 prompts a day the campus would be donating the largest free labeled dataset in Utah higher education. | Only for code and questions with no UVU data in them. |
| Muse Spark standard tier in the frontier pool | $1.25 / $4.25 per million tokens; a typical 1,500-token research exchange ≈ $0.006. For the 200-researcher pool, adding Spark to the routed mix changes yearly cost by tens of dollars either way. | Belongs in the pool as a route; changes nothing about the floor. |
What free tiers are genuinely good for
- Sandboxes with no UVU data — a CS student trying a new model on their own code; a faculty member evaluating a release before the fleet qualifies it.
- Model evaluation — free tiers are the cheapest way to benchmark a candidate against the fleet's current model before pulling its weights.
- Teaching the routing lesson — the plan's data-class rules (§08) are exactly the discipline that makes free tiers safe: Public-class work may burst to labeled outside routes; nothing Sensitive ever does.
- Never the floor. No published free tier offers institutional terms, an admin plane, an SLA, or a promise to exist next semester. The owned floor is what makes the free options safe to use at all.
Sources verified directly: Meta's model page and September 2 announcement; Artificial Analysis, September 2; OpenCode Zen documentation; OpenRouter's rate-limit documentation and Terms of Service (updated August 31, 2026).
Every route to Muse Spark 1.3 — and its terms
Prices are per million tokens: input / cache read / output. The pattern to notice: the same model is sold at two prices, and the cheap one is paid for with your data.
| Route | Price and limits | Data terms | Eligibility, duration, institutional reading | Grade |
|---|---|---|---|---|
| OpenCode Zen Contributor Free; raw ID muse-spark-1.3-contributor-free, OpenCode config ID opencode/muse-spark-1.3-contributor-free | $0 / $0 / $0. Exact RPM, TPM, and request caps are UNKNOWN. | Prompts and completions may train future Meta models. OpenCode says prompts pass upstream without being stored by OpenCode, but its general unpaid-account term also permits service-improvement use. Meta retention is UNKNOWN. | “Limited time,” with no end date. General terms permit organizational accounts; no education program was found. Setup requests billing details. Post-free removal, conversion, and fallback are UNKNOWN. Zen, privacy, terms | fact unknown |
| OpenCode Zen workspace | Pay-as-you-go credits; default auto-reload adds $20 when balance drops below $5 unless disabled. Card fee: 4.4% + $0.30. | Admins can disable data-collecting models. | Team workspace administration is free during beta; future price is unknown. Admin/member roles and per-member/monthly limits exist. | fact |
| OpenCode Go; opencode-go/muse-spark-1.3-contributor | $10/month; accounting rate $0.10 / $0.002 / $0.20. Usage-value caps: $12 per five hours, $30/week, $60/month. One subscribed member per workspace. | Training: Yes. Zero-data-retention: No. Exact retention time unknown. | Region limited by Meta policy. After limits, users can use free models or optionally fall back to paid Zen balance. Not a campus-wide entitlement. Go documentation | fact |
| Meta Model API standard; muse-spark-1.3 | Meta-upstream public routes list $1.25 / $0.15 / $4.25. The unauthenticated Meta pricing page could not be read, so direct-account rate limits and max pricing remain UNKNOWN. | Public gateway disclosures describe the standard route as non-contributor/no-training. Direct Meta retention and ZDR wording could not be verified anonymously. | Public preview. Exact tier-based RPM/TPM and education programs are UNKNOWN. Meta release, OpenRouter, Artificial Analysis | fact unknown |
| Meta Model API Contributor; muse-spark-1.3-contributor | Public routes list $0.10 / $0.002 / $0.20. Conflicting secondary rate-limit figures were found; none is safe to publish as exact. | Meta may use prompts and completions to improve/train future products; not ZDR. Retention duration unknown. | Geography page was login-gated. No education exception was found. | fact unknown |
| OpenRouter standard and Contributor | Standard: $1.25 / $0.15 / $4.25. Contributor: $0.10 / $0.002 / $0.20. | Contributor route explicitly allows training. OpenRouter can enforce no-data-collection/ZDR routing only when a compliant provider is available; Contributor is not such a route. | Both currently relay Meta as the upstream, not separately hosted weights. Standard, Contributor, provider privacy controls | fact |
| Vercel AI Gateway | Standard public rate matches $1.25 input / $4.25 output; Contributor is also listed. | Inherits the Meta route’s data distinction. | Gateway, not an independent host. Vercel provider page | fact |
| Muse Code subscriptions | Current reporting lists $5 Everyday, $15 High, and $50 Power monthly plans. First-party subscription material was rate-limited during research. | Plan-specific institutional data terms were not verified. | Consumer/developer product, not a proven UVU entitlement. | est |
| AWS Bedrock, Vertex AI, Azure AI | No exact Spark 1.3 listing found. | — | Search miss, not proof of global absence. | unknown |
Two more cautions from the survey: independent Terminal-Bench runs score Spark at 85–86 rather than Meta's 88.8, and most of Meta's launch comparisons used a "max" reasoning mode that was not generally available. And an instructive earlier incident: a misconfigured third-party test of Spark 1.1 gave the model internet access to a real target, and it changed real data — not a jailbreak, but proof that isolation and target validation are the harness's job, not the model's.
Longevity — the dated record behind the "free" offer
| Date | Event | What it means for durability | Grade |
|---|---|---|---|
| 2023-02-24 | Original LLaMA used research-only, case-by-case access. Meta | Meta’s meaning of “open” has changed across generations. | fact |
| 2023-07-18 | Llama 2 added broad commercial use under a custom license, a 700M-user gate, and restrictions on training other LLMs. License | Open-weight did not mean OSI open source. | fact |
| 2024-04-18 to 2024-07-23 | Llama 3 retained restrictions; 3.1 loosened output use but added derivative naming duties. Llama 3, 3.1 | New generations can arrive with materially different terms. This is not evidence of retroactive relicensing. | fact |
| 2024-09-25 to 2025-04-05 | Llama 3.2 and Llama 4 added/retained EU multimodal restrictions. 3.2 policy, Llama 4 | Geography and use rights require version-specific review. | fact |
| 2025-04-29 | Meta launched hosted Llama API in limited free preview. Meta | Hosted Meta model access has precedent—but not permanence. | fact |
| 2026-07-06 | Reporting said Meta’s hosted Llama API closed after roughly 14 months. The official deprecation page existed but was unreadable anonymously; shutdown was not live-tested. Official page, contemporaneous report | Supports high endpoint-lifecycle risk, with incomplete primary proof. | fact unknown |
| 2026-03-01 | Older Llama repositories were consolidated; the Llama 3 repository was archived. Repository | Tooling/repository retirement did not withdraw downloaded weights. | fact |
| 2026-04-08 to 2026-09-02 | Four Spark releases in under five months: original, 1.1, 1.2, 1.3. Version intervals were 92, 27, and 28 days. | Fast improvement, but poor evidence for long pinned-version life. | fact |
| 2026-08-10 | Meta released separate Muse Glimmer weights under Apache-2.0. Meta | Meta can release Muse weights; Glimmer does not guarantee Spark weights or compatibility. | fact |
| 2026-09-02 | Spark 1.3 shipped while the Model API remained a public preview; future Spark weights remained only a roadmap item. | No support window, overlap period, EOL notice, or SLA was found. | fact |
| 2025-04 to 2025-06 | Current OpenCode code line had a 0.0.1 tag on April 22; a founder interview says public launch was June 19. Tag | Exact “first release” is unresolved; the project is about 17 months old. | fact unknown |
| 2025-09 | Zen was publicly visible during September; sources disagree on the exact launch day. | Hosted gateway is about one year old. | unknown |
| 2026-02-06 to 2026-08-05 | Zen documents 18 model retirements across GPT/Codex, Claude, Gemini, GLM, Kimi, Qwen, and MiniMax. Zen | Direct evidence of catalog churn. | fact |
| 2026-09-03 | OpenCode: about 203.4k stars, 26.5k forks, 15,650 commits, approximately 950–1,000 contributors, and a release on September 2. MIT client; company/core-team governance. Repository, license | Strong project momentum and an easy fork/exit path. | fact |
| 2026-09-03 | Operator is Anomaly Innovations, Inc.; YC lists an active 24-person W2021 company. Predecessor SST raised $1M in 2021. Terms, YC, funding report | Real backing, but audited current finances are unavailable. | fact |
| If the plan depended on… | Risk | Why |
|---|---|---|
| Pinned Muse Spark 1.3 endpoint | HIGH | Public preview, four versions in five months, no lifecycle policy, closed weights, and prior hosted-API sunset evidence. |
| Future Spark weights | HIGH/UNKNOWN | Only a roadmap statement; no checkpoint, date, license, size, or hardware requirement. |
| OpenCode open client | LOW–MEDIUM | MIT, unusually active, broad provider support, large contributor base, and readily forkable. Company-led governance remains a concentration risk. |
| OpenCode Zen/Go service | MEDIUM | Young, upstream-dependent, terms allow limits/discontinuation, and the catalog changes quickly. Paid products, backing, and client portability reduce exit risk. |
The harness landscape — what a university could actually stand on
A harness is the software people touch: the chat front door, the coding agent, the research tool. A powerful harness does not make a model smarter; it gives the model context, memory, tools, and authority — which is exactly why the choice matters for a campus. Snapshot September 3, 2026. Controls legend: I identity/SSO, B enforceable per-user budget, A audit/admin policy; Y verified, N absent, U not publicly proved. No product published a current accessibility conformance report — "a11y U" means unproven, not unusable.
Chat and assistant front doors
| Harness · license · maturity | Models and local fit | Controls | Data default · accessibility | University fit | Grade |
|---|---|---|---|---|---|
| LibreChat; MIT; ≥2023; ~42.8k★/~5.4k commits; active 0.8.8 RC line | OpenAI-compatible, Ollama/local, major clouds; works cleanly behind LiteLLM | I:Y OIDC/SAML/OAuth/LDAP; B:Y token-credit balances; A:Y, though admin UI is still marked Preview | App RUM off and GTM only when configured; usage/cost transactions stored. Bundled Meilisearch analytics needs explicit disabling. A11y U, but current keyboard, focus, contrast, and screen-reader work is documented. | Best broad front door. Keep the current plan; validate RC admin controls and telemetry settings. | fact unknown |
| Open WebUI; custom Open WebUI License, BSD-3 base plus branding requirement above the threshold; not ordinary OSI open source; created 2023-10-06; ~150.8k★/~18.4k commits | Ollama and any OpenAI-compatible endpoint | I:Y OIDC/LDAP/groups/RBAC; B:U; A:Y, audit export off by default | No outbound app telemetry by default; messages and usage remain in its DB. Community functions can execute unaudited Python. A11y U; high-contrast/text-scale controls exist. | Strong alternative pilot if branding terms are accepted and LiteLLM enforces spend. | fact unknown |
| AnythingLLM; MIT; ≥2023; ~65k★/~2.4k commits; Mintplex Labs, regular releases | Generic OpenAI, LiteLLM, llama.cpp, Ollama, LM Studio, LocalAI | Built-in multi-user accounts; I:U campus SSO; B:U; A:U | Anonymous PostHog usage is on unless DISABLE_TELEMETRY=true; vendor says documents/chats are excluded. No-auth installs and admin-enabled unrestricted agents are supported. A11y U. | Department/lab pilot, not first central choice. | fact unknown |
| Jan; Apache-2.0; ≥2023; ~44.3k★/~8.6k commits; active Jan/Menlo project | OpenAI/Anthropic-compatible, Ollama, LocalAI, TGI, llama.cpp, LiteLLM | Desktop I:N/B:N/A:N; separate Jan Server has Keycloak/OIDC | Local chats/settings/logs; no telemetry without consent. Cloud prompts go to the selected provider. A11y U. | Good personal faculty/research desktop, not a campus front door. | fact |
| Cherry Studio; AGPL-3.0 Community Edition; ≥2024; ~51.4k★; rapid 1.x→2.x cadence | 50+ providers, Ollama/LM Studio, MCP and RAG | CE desktop I:N/B:N/A:N; enterprise detail insufficient | Anonymous usage, errors, and crashes on by default; can be disabled, but upgrades reset the settings on. A11y U. | Individual or department use until enterprise controls are proved in writing. | fact unknown |
| Msty; proprietary per-user license; first changelog 2024-01-03; active 2.9.x | OpenAI-compatible, MLX/llama.cpp/Ollama, cloud providers | I:Y enterprise SSO; B:U; A:Y logs, allowlists, access controls | Msty says Studio has no analytics/telemetry. Enterprise retains team and audit metadata in a single-tenant service. A11y U. | Plausible managed pilot, subject to procurement and security evidence. | fact unknown |
| LobeHub, formerly LobeChat; LobeHub Community License, Apache-derived with commercial/distribution restrictions; ≥2023; ~82.2k★/~13.4k commits | Major clouds, OpenAI-compatible, Ollama; now an agent workspace | I:Y generic OIDC options; B:U; A:U | Analytics packages exist, but no complete official default/payload statement was found. A11y U. | Legal/security pilot only; license and control gaps outweigh feature breadth. | fact unknown |
| Onyx, formerly Danswer; MIT Community Edition, separate enterprise license; ≥2023; ~31.9k★/~8.1k commits; company-backed, active v4 | Major clouds, Ollama, LiteLLM, vLLM; deep connector catalog | I:Y, but CE/enterprise docs conflict; B:U; A:Y group/model/agent controls | Self-host sends opt-out aggregate telemetry; other data stays in UVU infrastructure except selected providers. A11y U. | Strongest governed knowledge-search candidate, after obtaining a versioned entitlement map. | fact unknown |
| Khoj; AGPL-3.0; ≥2023; ~37k★/~5.2k commits; 2.0 beta, cloud deprecation announced | OpenAI-compatible, Ollama/vLLM/llama.cpp, major providers | Google OAuth/magic link; generic OIDC/SAML U; B:U/A:U | Anonymous telemetry on by default; disable with KHOJ_TELEMETRY_DISABLE=True; vendor says prompts and PII are excluded. A11y U. | Personal/research-lab self-host only while service direction contracts. | fact unknown |
Coding and agent harnesses
| Harness · license · maturity | Provider posture | Controls and data posture | University fit | Grade |
|---|---|---|---|---|
| OpenCode; MIT; ≥2025; ~203k★/~15.7k commits; Anomaly; extremely active | Provider-neutral: OpenAI-compatible, llama.cpp, LM Studio, Ollama, broad clouds; direct LiteLLM Meta support | Enterprise I:Y and central configuration; generic client B:U/A:U. Direct/BYOK prompts reportedly not stored; /share uploads a full session. Hosted unpaid terms allow improvement use. A11y U. | Preferred CLI builder pilot. Disable sharing, use UVU keys/gateway, and ship a deny-first permission profile. | fact unknown |
| Cline; Apache-2.0; ≥2024; ~67.4k★/~7.2k commits; company-backed, rapid releases | Provider-neutral; custom OpenAI-compatible endpoints, Ollama/LM Studio and major clouds | I:Y SAML/OIDC/RBAC; B:U; A:Y tool/model policy, cost telemetry and selective audit. Anonymous telemetry opt-in; prompt archive is explicit. A11y U. | Top managed IDE pilot. Centrally disable automatic unrestricted execution and use sandboxes. | fact unknown |
| Roo Code; Apache-2.0; ≥2024; ~24.3k★ | Historically provider-neutral | All products and maintenance ended 2026-05-15. Sunset notice | Reject: discontinued. | fact |
| Continue; Apache-2.0; ≥2023; ~35.7k★ | Historically provider-neutral | Repository is read-only and not actively maintained; final 2.0 removed auth and telemetry. | Reject: retired/archive only. | fact |
| aider; Apache-2.0; ≥2023; ~48.7k★/~13.1k commits; slowing — last release Aug 2025, last code push May 2026 (GitHub API, Sept 3, 2026) | Almost any cloud or local model, including Ollama/OpenAI-compatible routes | I:N/B:N/A:N; enforce all controls at LiteLLM. Analytics asks consent; the selected prompt default is Yes, excluding code/chats/keys/PII. A11y U. | Good expert tool, not a managed campus platform by itself — and watch its maintenance cadence before standardizing on it. | fact unknown |
| Goose; Apache-2.0; ≥2025; ~53.9k★; 500+ contributors; Linux Foundation Agentic AI Foundation | Provider-neutral; Ollama, LiteLLM, 15+ providers, MCP/ACP | Central I/B/A:U. Telemetry off by default—but prompt-injection checks and desktop sandbox are also off by default. A11y U. | Strong sandboxed IT/CS/research pilot. | fact unknown |
| OpenHands; MIT; ≥2024; ~86k★/~8.1k commits; company-backed | Any LLM/BYOK; can supervise other coding agents; LiteLLM supported | Enterprise I:Y, SAML/SSO/RBAC; B:U; A:Y with event coverage U. Self-host can keep code/conversations at UVU; traces may contain tool and LLM I/O when enabled. A11y U. | Enterprise automation pilot only, with containers and narrow repository credentials. | fact unknown |
| Kilo Code; MIT core; ≥2025; ~27.2k★/~30k commits, some inherited fork history; acquired by Anaconda | Provider-neutral; 500+ hosted, BYOK and local models | I:Y SSO/SCIM/RBAC; B:Y user/sub-org limits; A:Y logs and model allowlists. Operational telemetry may contain linked diagnostic data. A11y U. | Strong managed candidate after residency and telemetry review. | fact |
| Codex CLI; Apache-2.0; public 2025-04; ~121k★/~10.2k commits; OpenAI, very active | OpenAI-first; --oss supports Ollama/LM Studio and custom Responses-compatible providers | Workspace I:Y/A:Y; client B:U. Business/Enterprise/Edu traffic is not trained on by default; Plus/Pro needs opt-out. A11y U. | Strong controlled engineering/CS pilot, but operationally vendor-centered. | fact unknown |
| Gemini CLI; Apache-2.0; 2025-06-25; ~106.8k★/~6.4k commits; Google, weekly stable cadence | Vendor-tied: Gemini/Google/Vertex models | Google identity; B:U in client; A:Y immutable policy, strict mode, MCP allowlist, OTEL. OTEL is off, but if enabled prompt logging defaults true. A11y U. | Good Google-centered pilot, not a neutral campus standard. | fact unknown |
| Claude Code; proprietary core; preview 2025-02-24; distribution/plugin repo ~144k★ | Vendor-tied: Claude through Anthropic, Bedrock, Vertex, Foundry, or a LiteLLM gateway | I:Y/A:Y; budget through provider/gateway. Retention is channel-specific; enterprise local-session transcripts can default to six years unless changed. ZDR may be available. A11y U. | Controlled vendor pilot only after retention and route are contractually fixed. | fact unknown |
Research, retrieval, and workflow harnesses
| Harness · license · maturity | Provider support | Controls, data, accessibility | University fit | Grade |
|---|---|---|---|---|
| Onyx/Danswer; MIT CE plus separately licensed enterprise; active v4 | Clouds, Ollama, LiteLLM/vLLM; 40+ enterprise connectors | Identity available but CE/enterprise docs conflict; hard user budget U; model/agent/document controls available; aggregate telemetry opt-out; a11y U | Best central library, policy, and department knowledge pilot. Require a written feature/entitlement map. | fact unknown |
| Kotaemon; Apache-2.0; ≥2024; ~25.7k★ but only ~327 commits | OpenAI/Azure/Cohere, Ollama and llama.cpp | Basic multi-user login; OIDC/SAML, budgets, audit and telemetry defaults U; a11y U | Research prototype, not a central service. | fact unknown |
| RAGFlow; Apache-2.0; ≥2024; ~90k★/~9k commits; InfiniFlow, active 0.27.x | OpenAI-compatible, Ollama, LocalAI, vLLM/Xinference and clouds | OIDC present; budgets/audit/telemetry U; a11y U. Heavy deployment; official docs require gVisor when code execution is enabled. | Advanced research/RAG pilot with dedicated platform engineering. | fact unknown |
| Dify; modified Apache-2.0, with multi-tenant permission and branding restrictions; ≥2023; ~154k★/~24k forks; active | Broad provider/plugin catalog; OpenAI-compatible and self-hosted models | Enterprise SSO and operation logs; hard user budget U. OTEL off by default; marketplace connections exist; complete analytics default U; a11y U. | Strong app/workflow builder, but likely needs commercial licensing for campus multi-tenancy. | fact unknown |
| n8n; Sustainable Use License, source-available rather than OSI open source; ≥2019; ~203k★/~60.5k forks/~23.8k commits | 1,500+ integrations, major model providers and Ollama | Enterprise SAML/OIDC/LDAP, RBAC and log streaming; per-user AI budget U. Anonymous diagnostics on by default. A 2026 issue reports an external banner request despite isolation flags; a11y U. | IT integration tier only, with least-privilege service accounts, approval gates, and outbound monitoring. | fact unknown |
Lifecycle matters more than popularity: Roo Code was discontinued in May 2026, Continue's repository is read-only, and Khoj's hosted direction is shrinking. Open WebUI, LobeHub, Dify, and n8n carry licenses that are not ordinary permissive open source and should not be presented as such.
What an agentic harness buys — and risks — in a classroom
| Beyond a chat box | Concrete value | Main risk | Likely users |
|---|---|---|---|
| Repository discovery and planning | Searches files, follows imports, reads tests and plans multi-file changes | Private code and secrets enter context | CS, IS, central IT, engineering |
| File creation and editing | Applies patches, refactors modules, writes tests and documentation | Wrong or destructive changes; licensing/IP contamination | Software and data teams |
| Shell and test execution | Installs dependencies, runs builds, linters, tests, notebooks and CLIs | Host compromise, malicious packages, resource exhaustion | CS labs, engineering, analytics |
| Browser/web tools | Reads documentation, investigates errors and interacts with web applications | Prompt injection from pages; accidental action on real systems | IT, cybersecurity, research |
| MCP and business tools | Calls databases, ticketing systems, cloud APIs and internal services | Credential misuse and cross-system data movement | IS, business analytics, integration teams |
| Iterative agent loop | Observes failures, revises its plan and continues without one prompt per step | Runaway token spend and compounding mistakes | Advanced builders and researchers |
| Git integration | Produces reviewable diffs, branches and commits | Unapproved pushes, releases or deployment | Managed development teams |
| RAG and permission-aware search | Answers across policy, library, research and department repositories | Stale ACLs, poisoned documents and cross-unit leakage | Library, research administration, service desks |
The controls that make it safe — each one is already a design commitment of this plan:
- Treat repository text, web pages, documents, issues, and MCP output as untrusted input. Prompt injection can become shell execution.
- Run every student or researcher agent in an ephemeral container/VM. Do not mount the home directory or credential stores.
- Deny outbound network access by default; allow only named package registries and approved endpoints.
- Use short-lived, task-scoped credentials. Default repository access to read-only; deny deploy, publish, push, billing, and production tools.
- Require a person to approve file writes outside the workspace, shell/network actions, package installation, and consequential tool calls.
- Put all model traffic through LiteLLM for identity, model allowlists, per-user budgets, cost alarms, and emergency shutoff.
- Keep FERPA, HR, finance, security, unpublished research, controlled data, credentials, and proprietary partner material out of contributor/free-training routes.
- Set course-level rules for disclosure, allowed tools, retained work logs, individual contribution, and assessment design.
- OpenCode’s current permission system supports allow/ask/deny rules, but most ordinary actions are permissive by default. UVU must ship its own deny-first profile. Permissions
- The UnderSpecBench study found that agents guess action boundaries when operational tasks are underspecified. Assignment and run instructions must explicitly define allowed targets and actions. Paper
Three tiers, one router
| Tier | Recommended harnesses | Local floor and frontier route | Free-tier policy and controls |
|---|---|---|---|
| Everyone | LibreChat as the supported front door; Onyx search exposed only for approved knowledge collections | LibreChat → LiteLLM → Qwen3.8-27B local default. Use GLM-5.3 only on infrastructure proven to fit it. Metered frontier routes selected by task/data class. | Do not expose Contributor Free in the general model menu. Use only institutionally accepted free entitlements with suitable no-training terms. Campus SSO, per-user budgets and logs are mandatory. |
| Builders | OpenCode default CLI; Cline managed IDE option; Goose or OpenHands for advanced sandboxed work; Kilo worth a procurement pilot | Harness → LiteLLM → local model first; route hard coding tasks to paid non-contributor Spark, Sol, Opus, or Gemini according to evaluation and budget | Contributor Free may appear only in a separate public-data sandbox with separate keys, a warning banner, no persistent connector credentials, disabled sharing, and no production access. |
| Researchers | Onyx for governed knowledge search; RAGFlow for advanced ingestion; Kotaemon for prototypes; Dify/n8n only as controlled workflow platforms | Local embeddings and local Qwen for protected corpora; paid no-training frontier route only when the data agreement allows it | No contributor/free-training route for source documents, unpublished work, grant material, participant data or licensed collections. Connector ACL tests, retention rules and source citations are required. |
The layers stay separate on purpose: a broad chat portal, a shell-capable coding agent, and a credentialed workflow engine do not belong inside one permission boundary.
Verdict on Spark
Muse Spark 1.3 changes the frontier-pool candidate list, not the architecture.
| Option | Result | Benefit | Main risk | Ease of undo |
|---|---|---|---|---|
| (a) Add paid, non-contributor Spark plus (b) a separate Contributor Free sandbox est | Pilot muse-spark-1.3 behind LiteLLM for coding/agent tasks. Offer free Contributor only in an isolated public-data environment. | Captures the strong model and the legitimate $0 experiment without trading UVU data for compute. | Meta endpoint churn; first-party retention/rate terms still need procurement confirmation; sandbox users may paste restricted data. | Easy if model names are abstracted behind LiteLLM and no workflow depends on Spark-only behavior. |
| (a) Paid non-contributor only est | Cleanest institutional route. | Simpler policy and lower data risk; still much cheaper than some frontier models. | Gives up free experimentation; direct Meta terms still need confirmation. | Easy. |
| (b) Contributor Free broadly est | Free campus access while promotion lasts. | Near-zero token spend. | Training use, non-ZDR status, unknown limits/retention/end date, behavioral leakage, and weak durability. | Technically easy, but data already disclosed cannot be recalled. Do not choose. |
| (c) Ignore Spark est | Keep the present model pool unchanged. | Least procurement and policy work. | Misses a capable, relatively inexpensive coding route and useful competitive pressure. | Easy. |
The paid pilot should not start until Meta's or the gateway's current terms confirm, in writing: no training, retention period, geography, incident handling, rate limits, and permitted educational use. Survey: working paper 34, September 3, 2026.
Every provider with free access — the inventory
Twenty-eight providers checked on their own pricing, limit, and terms pages on September 3, 2026. "Free" turns out to mean six different things: a renewable allowance, a one-time signup credit, a thirty-day trial, a limited-time promotion, a data-for-discount trade, or nothing at all. The survey's headline: OpenAI, Anthropic, xAI, Together, DeepInfra, and Meta's own API have no recurring free generation tier; GitHub Models — the most generous free gateway of 2025 — was fully retired on July 30, 2026, after about 23 months.
| Provider | Free access and capability anchor | Exact public limits and scope | Reliability / practice signal |
|---|---|---|---|
| OpenRouter | fact 18 literal free IDs above; random openrouter/free router. | 20 RPM; 50 RPD. After buying at least $10: 1,000 RPD, but no longer a zero-cash path. TPM/TPD/concurrency and exact org aggregation unknown. Limits, 2026-06-12 | Free model availability varies; current free endpoint showed 96.8% availability. Users report 429/500 errors and model switching. |
| Google AI Studio / Gemini API | fact Zero-price rows include Gemini 3.8/3.7/3.6/3.5 Flash; 3.5/3.1 Flash-Lite; Gemini 3 Flash Preview; Gemini 2.5 Pro, Flash and Flash-Lite; selected live/audio/TTS/transcription, embeddings, robotics, and Gemma models. Gemini 3.8 Flash: AA Intelligence 59 on 2026-09-02. | Limits apply per project, not key. Exact free RPM/RPD/TPM are now visible only in AI Studio and are not guaranteed: unknown. Representative Flash context: 1,048,576 input/65,536 output. Rate limits, updated 2026-09-02 | Google staff confirmed a 20-RPD reduction in Dec. 2025 and said free access was not designed as a long-term application base. |
| Groq | fact GPT‑OSS‑120B/20B, Qwen 3.6/3.8 27B, Compound/Mini, Prompt Guard, Whisper and Orpheus. GPT‑OSS endpoint: about 471 output tokens/s and 86% provider-test accuracy on 2026-09-03. | GPT‑OSS/Qwen: 30 RPM, 1,000 RPD, 8K TPM, 200K TPD. Compound: 30 RPM, 250 RPD, 70K TPM. Limits are per organization. Context generally 131K. Current table | Official status operational. Llama 3.1 8B and Llama 3.3 70B free/developer access ended 2026-08-16. |
| Cerebras | fact GPT‑OSS‑120B and Gemma 4 31B. GPT‑OSS endpoint: about 1,643 output tokens/s and 87% test accuracy. | Trial only: $5 credit, verified payment method, expires after 30 days. Each model: 5 RPM, 30K TPM, 1M tokens/hour and 1M TPD, per organization. Free/trial context 65K. Current limits | 100% trailing-90-day status on 2026-09-03. Former always-free service is no longer renewable. |
| SambaNova | conflicting sources Rate-limit docs list free DeepSeek V3.1, Llama 3.3 70B, GPT‑OSS‑120B, DeepSeek V3.2 and Gemma 4 31B. Its plans page says users must buy credits before the first request. | Docs say 20 RPM, 20 RPD, 200K TPD per user and model; most contexts 128K. Current practical zero-payment access was not login-tested. Limits · conflicting plan page | Status showed 99.93%–100% over 90 days. Treat free availability as unknown until a new account succeeds. |
| Together AI | fact No ordinary free trial. A separate $150-credit page exists, but eligibility and current availability are unknown. | Minimum $5 purchase; limits dynamic per organization/model. No-trial notice, updated 2026-06-01 | Serverless model availability ranged about 96.85%–100% over 30 days. |
| Fireworks AI | fact $1 one-time signup credit across eligible serverless models. Capability depends on chosen model. | Expiration, default RPM/TPM and concurrency are not public: unknown. Pricing | Serverless status operational. Self-service changed to prepaid on 2026-07-01. |
| Mistral Studio / La Plateforme | fact Free plan provides $10/month in API credits; model access is dynamic. labs-* models are free, experimental, and may silently change. Medium 3.5 AA Intelligence 30; Small 4, 20. | RPS/concurrency, TPM and monthly token caps are organization/model-specific and visible only after login: unknown. Current model contexts generally 128K–256K. Pricing · limits | API trailing-90-day uptime 99.395%; embeddings 94.337%. Labs are expressly not for production. |
| Hugging Face Inference Providers | fact $0.10 monthly credit for Free users; model and upstream-provider choice is broad. Capability and context depend on the selected model. | Dollar credit is exact; public RPM/TPM/concurrency are provider-specific and unknown. Pricing | Inference Endpoints API showed 100% over 30 days, but upstream availability differs. Free credit was reduced from a larger launch allowance. |
| Cloudflare Workers AI | fact Dynamic catalog within neuron allowance. Current free alternatives include GLM 4.7 Flash, Gemma 4 26B, Nemotron 3 120B and Llama 3.3 70B. | 10,000 neurons/day/account; general text 300 RPM; selected frontier models 20 RPM. Llama 3.3 70B example: about 125 standard turns/day. Pricing, updated 2026-08-28 | Workers AI operational; two incident-days in 30. Several frontier models became paid-only 2026-07-28. |
| GitHub Models | fact None. Fully retired 2026-07-30. | N/A. Retirement | The free gateway lasted about 23 months. Strongest direct warning against using previews as a campus floor. |
| NVIDIA build.nvidia.com / NIM | fact Developer Program members get dynamic hosted “Free Endpoint” models for prototyping. Recent examples include Kimi K3, DeepSeek V4 Pro and Nemotron 3.5 Lightning. | Public RPM/RPD/TPM/credit amount and context inventory: unknown. Community reports commonly see 1,000 credits and 40 RPM, but these are not contractual limits. Current program | No public hosted-NIM status page found. NVIDIA warns trial services and endpoints can change or end. |
| OpenCode Zen | fact Big Pickle, MiMo‑V2.5, Ling 3.0 Flash Fin, Nemotron 3 Ultra, Nemotron 3.5 Lightning, Muse Spark 1.3 Contributor and Muse 1.2 Contributor are listed free. Live API also exposed two other free IDs. Muse 1.3: AA Intelligence 62 on 2026-09-02. | RPM/RPD/TPM/TPD/concurrency/context: unknown. Every listed free model is limited-time. Zen, updated 2026-09-03 | Official-repository issue reported intermittent free endpoint unavailability on 2026-08-25. |
| Meta Model API / Muse | fact No verified direct recurring free tier. Muse 1.3 Standard is paid; Contributor is deeply discounted and trades data rights for price. Zen temporarily supplies Contributor free. | Direct current public limits and retention pages required login: unknown. OpenRouter shows 1,048,576 context. Muse 1.3 AA Intelligence 62. Meta release, 2026-09-02 | Product moved from private preview in April to public preview in July 2026. Durable paid availability is more likely than durable free availability. |
| Ollama Cloud | fact Recurring monthly Starter allowance and a smaller Starter model set, both undisclosed. | One concurrent request; additional requests queue. Monthly allowance, RPM/RPD/TPM and model roster: unknown. Pricing | Cloud started 2025-09-19; limits were redesigned 2026-08-31. Users reported large free models unexpectedly moving behind payment. |
| DeepInfra | fact No current free or trial allowance found. PAYG only. | Paid default is 200 concurrent requests/model; capacity can still return 429. Limits | API status 100% over 74 observed days; individual models about 98.78%–99.96%. |
| Novita AI | fact New-user voucher of unspecified value; some zero-price development models such as Llama 3.3 3B and Qwen 2.5 7B. Free-specific capability anchor unknown. | Amount, expiry, RPM/RPD/TPM and permanence: unknown. Quickstart | Model status about 98.4%–100%; zero-price model inventory changes. |
| Hyperbolic | fact $1 one-time credit after phone verification; broad paid model catalog. | 60 RPM on basic accounts; Pro 600 RPM after depositing $5. Token allowance depends on model price. Limits | No incident in prior seven days; no usable long-horizon uptime figure. Older $5/$10 signup pages are stale. |
| Chutes | fact No current free API. Its 200-RPD Early Access plan ended 2026-03-15. | N/A; current service is PAYG or paid subscription. Retirement notice, 2026-02-27 | API showed 100% over 30 days, but the free path is gone. |
| Kluster | unknown No current public self-service general inference offer or current free-price table found. | Limits, model roster and present signup credits: unknown. Historical $25 credits expired in Feb. 2025. | No public current status evidence found. Do not use cached 2025 offers. |
| Cohere trial/evaluation | fact Trial keys provide access to evaluation models; Command-family availability is account-dependent. | 1,000 API calls/month; Chat 20 RPM; Rerank 10 RPM. Non-production only. Limits | Official status operational. Trial allowance fell from 5,000 calls/month in 2023 to 1,000. |
| xAI API | fact No current free tier. The $25/month launch promotion ended in 2024. | Prepaid; $25 minimum top-up. Tier 0 means zero prior spend, not free inference. Billing | Official status includes API and billing incidents. |
| Anthropic API | fact No renewable tier. New users receive a small, unspecified one-time test credit. | Evaluation limits and credit value are unknown; paid limits are organization/model based. Pricing | Official 90-day API uptime was 99.52%; selective research grants exist. |
| OpenAI API | fact No general-generation free tier. One test request and free moderation only. | omni-moderation-latest: 250 RPM, 5,000 RPD, 10K TPM. GPT generation Free tier is not supported. Model page | Durable paid API; free generation cannot be planned as capacity. |
| Vercel AI Gateway | fact Every free team receives $5 every 30 days, usable across the catalog. GPT‑OSS‑120B is 131K context and starts at $0.10/M input, $0.50/M output. | No public free RPM/concurrency cap found. At a 700-input/300-output turn: about 22,727 turns/30 days. Pricing · model page | Gateway adds retries/fallbacks, but upstream status and terms remain relevant. |
| Z.AI | fact GLM‑4.7‑Flash and GLM‑4.5‑Flash are free; both advertise 200K context. | Numeric rate limits appear only in the account console: unknown. Model overview | Users report tight dynamic throttling. Free availability is a conversion path to paid coding plans. |
| Alibaba Model Studio | fact New users receive model-specific quotas, typically 1M tokens/model. OAuth separately advertises 2,000 calls/day. Qwen Max/Plus/Flash models are included depending on region. | Account and RAM users share quotas. Signup quota is time-limited; a 90-day policy takes effect 2026-09-08. Re-registering does not issue new quota. Free quota, updated 2026-09-01 | Trial, not durable floor. Regional restrictions and rapid model sunset notices apply. |
| IBM watsonx.ai | fact Lite provides up to 300,000 foundation-model tokens/month and open-model playground/API access. Current catalog includes Granite and selected third-party models. | 300K combined tokens/month; one Lite instance and one authorized user. RPM/concurrency/context depend on model: unknown. Runtime plan, updated 2026-07-30 | Evaluation plan; idle Lite instances may be deleted after 30 days. The token allowance increased in July 2025. |
Data terms and account rules — in the providers' own words
| Provider | Free-tier data treatment | Account / university-use rule |
|---|---|---|
| OpenRouter | Router logging is off by default; upstream providers still differ. ZDR routing is available. | Organizational authorized users are supported, but terms prohibit creating multiple accounts to bypass limits: “create multiple accounts…for purposes of bypassing or circumventing use limits.” Terms, updated 2026-08-31 |
| Unpaid prompts and responses may improve Google products and may be human-reviewed. Retention for ordinary unpaid calls is unknown. | 18+; no under-18-directed API client. “Do not submit sensitive, confidential, or personal information to the Unpaid Services.” Additional Terms, effective 2026-03-23 | |
| Groq | No training without permission; inference is not retained by default. Reliability/abuse logs can last 30 days; ZDR is available. | Organization-wide quota. “beyond published…limitations, including by registering multiple accounts or orchestrating usage between multiple organizations.” AUP, effective 2025-10-15 |
| Cerebras | Does not train on or retain ordinary inference I/O; operational retention duration is not stated. | Per organization; verified payment required. “Buy, sell or transfer API keys without our prior written consent.” Terms, effective 2024-08-27 |
| SambaNova | EULA limits customer-content processing to service delivery or law; staff says prompts are not stored or trained on. Output/log retention unknown. | Free docs say per user. Access may be given only to authorized users; resale or third-party availability is prohibited. |
| Together | Inputs/outputs are not stored by default; temporary caching can occur; training is opt-in. | Project keys may be shared only with trusted collaborators and have full project rights. Multi-account-evasion clause unknown. |
| Fireworks | Open-model I/O is not logged or stored by default unless the customer opts in. | “Buying, selling, or transferring API keys is prohibited without prior written consent.” Terms |
| Mistral | Standard API I/O is not used for training and normally has 30-day abuse-monitoring retention. Labs/Preview can train and offer no opt-out. | Limits are per organization. Buying, selling, or transferring keys/accounts is prohibited. |
| Hugging Face | HF does not store bodies/responses; content-free debugging logs remain up to 30 days. Upstream terms also apply. | Organization tokens exist. Explicit multi-account evasion rule unknown. |
| Cloudflare | No model training or service improvement without explicit consent; persistence occurs only if the customer invokes storage products. | Terms prohibit quota circumvention and “automated agents or scripts to create multiple accounts.” Terms, updated 2025-09-12 |
| GitHub Models | Retired. Historical service said model providers did not receive training rights. | One free account per person/entity; users could not share tokens to exceed limits. |
| NVIDIA | Trial terms can permit deidentified content to improve services; fine-tune input can persist 30 days and generated content 90 days. | Keys are tied to one user. Trial access may not be transferred or made available to third parties. |
| OpenCode Zen | OpenCode says it does not store prompt content, but free upstreams have exceptions. Muse Contributor explicitly permits Meta training. | Terms prohibit bulk or multiple accounts that evade usage or promotions. |
| Meta Muse | Contributor permits training on prompts/completions. Standard is described as no-training, but current primary retention terms were login-blocked. | Direct account-sharing, limit scope, and education terms: unknown. |
| Ollama Cloud | Prompts/responses are transient, not retained after fulfillment, and not used for training. | “Ollama is one account per person.” Teams require a separate shared account. Pricing |
| DeepInfra | I/O is not stored or trained on without consent; support/security exceptions may last 30 days. | Scoped keys and spending caps exist. Explicit multi-account rule unknown. |
| Novita | Terms state default ZDR, transient processing, and no training; account metadata has longer retention. | Multiple accounts for extra quotas are prohibited. |
| Hyperbolic | Inference input is transient, discarded, and not used for training. | Account/password sharing and multiple-account limit evasion are prohibited. |
| Chutes | No current free tier. Public-model I/O is not stored, logged, or trained on; third-party chutes may differ. | Multiple subscriptions used to evade caps are prohibited. |
| Kluster | Current public self-service data terms: unknown. Enterprise page claims zero prompt logging. | Current consumer/inference account rules: unknown. |
| Cohere | Trial I/O may support R&D, performance, and safety work. Trial users are told not to submit personal data. | One account; user IDs cannot be shared or used to bypass limits. |
| xAI | API data is not used for training without permission; default retention is 30 days; eligible teams can use ZDR. | Team-scoped limits and confidential credentials; explicit multi-account clause unknown. |
| Anthropic | Commercial/API I/O is not used for training by default. Public pages conflict between default non-retention and deletion within 30 days. | Each user/service should have its own key; automated account creation and ban evasion are prohibited. |
| OpenAI | API data is not used for training by default; abuse-monitoring logs can contain content for up to 30 days. Approved organizations can request ZDR. | Keys must remain secret; institutional projects/service accounts are supported. |
| Vercel | Vercel says Gateway I/O is not stored or trained on; upstream policies still apply. ZDR/no-training routing controls exist. | AI terms allow authorized users but prohibit sensitive personal information in input. Mass free-team farming is not expressly authorized. |
| Z.AI | A portion of request content may be cached; cache retention is undisclosed. Free-model training treatment: unknown. | Paid coding plans are individual and cannot be shared. General free-API account rules are unknown. |
| Alibaba | Free-tier training and detailed retention terms were not found: unknown. | One Alibaba account and its RAM users share quota; new accounts cannot be re-registered for another grant. |
| IBM watsonx.ai | Lite-specific prompt-training and retention terms were not clearly stated on the pricing/plan pages: unknown. | Lite is limited to one instance and one authorized user; it is an evaluation plan, not a campus entitlement. |
The capacity math, done properly
One "turn" below is one prompt-and-response of 700 input and 300 output tokens; a four-turn chat is four turns. Phase-1 demand: 36,600 people × 2.15 turns per day ≈ 78,690 turns per day; exam-week stress ≈ 708,000. The three most generous recurring, publicly measurable free offers, combined:
| Provider / model | Binding free limit | One institution-controlled account | 36,600 genuine personal accounts (raw maximum) |
|---|---|---|---|
| Vercel / GPT‑OSS‑120B | $5 per 30 days; $0.00022/standard turn | 22,727/30 days = 758/day average | 27,742,800/day |
| Groq / GPT‑OSS‑120B | 200K TPD; 1K tokens/turn | 200/day | 7,320,000/day |
| Cloudflare / Llama 3.3 70B | 10K neurons/day; 80.109 neurons/turn | 124/day | 4,538,400/day |
| Combined | — | 1,082/day | 39,601,200/day |
| Scenario | Base demand covered | Exam-stress demand covered |
|---|---|---|
| One UVU-controlled account/org at all three | 1.38% | 0.153% |
| 36,600-account raw multiplication | 50,326% raw capacity | 5,592% raw capacity |
| Genuine, self-managed personal accounts | 100%, subject to eligibility and each person’s terms | 100%, because 19.35 turns/person/day is below each allowance |
| Centrally created/poolable account farm | 100% only by violating or evading terms | Mathematically sufficient, contractually unusable |
The answer in one sentence: free tiers cover about 1.4% of the campus's base day as a centrally operated service; the only way to 100% is to centralize thousands of personal accounts, which the providers' terms forbid — and the daily numbers are generous, because the same providers cap requests per minute at 20–30 with no promise of burst capacity against a base of ~53 simultaneous conversations and an exam-week peak near 474. Separately, a real person's own free account can cover that person's own demand — a genuine personal benefit, but not campus infrastructure UVU can promise, audit, or revoke.
What breaks at university scale
- Account farming is not a lawful architecture. Groq expressly prohibits orchestrating usage across multiple organizations. Cloudflare prohibits scripted multi-account creation and quota circumvention. OpenRouter and OpenCode prohibit multiple accounts used to bypass limits.
- A real person’s account is not UVU capacity. The person controls consent, settings, deletion, model selection, and account continuity. UVU cannot promise access, audit activity, revoke all access centrally, or move unused personal quota into a campus pool.
- 36,600 credentials become a security system of their own. Providers say keys must remain confidential and must not be embedded client-side. Central collection would create a high-value credential store; sharing keys often violates the applicable terms.
- Sensitive data leaves campus under inconsistent terms. Google explicitly warns against confidential, sensitive, or personal information in its unpaid API. Muse Contributor permits training. Cohere trial data may support R&D and safety. NVIDIA trial terms allow some service-improvement use.
- FERPA needs control, not merely a privacy promise. The FERPA school-official exception requires the institution to exercise direct control over the vendor’s use and maintenance of education records. Personal consumer accounts normally do not create that relationship. U.S. Department of Education FERPA guidance
- UVU’s own rules follow the data wherever it is stored. UVU classifies grades, schedules, disciplinary, financial, payroll, research and similar data as Sensitive/Confidential, including when a contractor or cloud stores it. UVU data classification · Policy 445, effective 2024-03-28 · Policy 447, effective 2026-08-10
- Free services lack institutional controls. Some have team dashboards, but personal free accounts generally lack UVU SSO, SCIM, domain claims, legal hold, central retention, eDiscovery, audit export, data-region choice, and contract remedies.
- Models rotate silently or with short notice. OpenRouter’s router chooses among available models; Mistral Labs permits silent updates; NVIDIA trial endpoints can change; OpenCode says free models are limited-time. GitHub Models demonstrates that an entire service can disappear.
- No SLA means no floor. Status histories show good averages for several vendors, but free/trial terms do not promise capacity during registration, class-change bursts, finals, or provider-wide shortages.
- Age and consent complicate rollout. Google’s developer API requires users to be 18+. Groq’s current service agreement is also adult/business focused. UVU still has some under-18 students.
Governance conclusion, not legal advice: Public-class prompts may use personal or free tiers after clear notice; anything Sensitive or Confidential needs an approved institutional service with a data-processing agreement and a reviewed processing chain — which is what the owned floor is.
Free, metered, or owned — the honest three-way comparison
| Choice | Privacy and terms | Predictability and control | Three-year cost at modeled base demand |
|---|---|---|---|
| Free cloud tiers | Mixed: from no-training/ZDR to explicit training and human review. Mostly self-service terms. | Lowest. Quotas, models and access can change; little SLA or central administration. | $0 inference, but the three measurable providers cover only 1.38% of base demand centrally. Integration, support and governance still cost staff time. |
| Metered hosted open APIs | Paid terms usually offer no-training defaults, better ZDR choices, DPAs and institution-owned keys. Requests still leave campus and may traverse a gateway plus upstream provider. | Stronger: central gateway, budgets, rate controls, audit, fallbacks and negotiated capacity. Vendor/model churn remains. | est 86.17M standard turns over three years. GPT‑OSS‑120B raw inference is about $6.6K at a low OpenRouter route, $19.0K through Vercel’s displayed route, or $24.6K at Groq’s rate. With a 30% reserve: roughly $8.6K–$31.9K, excluding gateway and labor. |
| Current owned Apple fleet | Prompts remain on campus if all routing, storage and telemetry are local. Best control over retention and logs. | Strongest governance and fixed local capacity, but UVU owns uptime, upgrades, security, model qualification and staffing. | est 26 M5 Pro/48GB minis plus kit: $63.9K upfront / $79.6K three-year machine TCO, excluding labor. Rated around 104 streams: enough for the 53-stream base, not the 474-stream whole-campus exam case. |
| Owned fleet scaled to whole-campus exam peak | Same local benefits. | Fixed 9× peak capacity, much of it idle on ordinary days. | est ceil(474÷4)+1 spare = 120 minis; roughly $367K three-year machine TCO by linear scaling, before larger network/rack/operations costs. This is a planning bound, not a quote. |
The finding worth sitting with: at the observed prompt rate, a cheap hosted open model can cost less than the machines — roughly $9k–$32k over three years for inference alone against about $80k of owned hardware. The owned floor earns its place through privacy, continuity, and control, not because local inference is always cheaper. The guide has said this since its first version (§12); this survey puts numbers under it from the free-provider side.
Longevity — twenty dated events
| Date | Event | What it proves |
|---|---|---|
| 2023-09-27 | Cloudflare launched Workers AI for all plans and warned access/limits could change. | A recurring allowance can last, but model eligibility is movable. |
| 2024-07-29 | NVIDIA launched free hosted NIM credits for Developer Program members. | Durable prototyping funnel; no production promise. |
| 2024-08-01 | GitHub Models public preview launched. | Start of a free gateway later retired. |
| 2024-08-27 | Cerebras launched an always-free inference tier. | Historical free promise did not survive unchanged. |
| 2024-09-10 | SambaNova Cloud launched with free Llama 3.1 405B access. | Model later disappeared; catalog churn is real. |
| 2024-09-17 | Mistral launched its free experimentation tier. | Now framed as monthly credits plus experimental Labs. |
| 2024-11-04 | xAI offered $25/month only through the end of 2024. | Promotion had an explicit end and is not current. |
| Feb. 2025 | Hugging Face launched Inference Providers with a larger small allowance. | Current Free allowance is only $0.10/month. |
| Apr. 2025 | Users reported OpenRouter free limits falling from 200 to 50 RPD. | Anecdotal exact cut; current 50-RPD primary limit is confirmed. |
| 2025-07-10 | OpenRouter said two providers withdrew free capacity and that it was subsidizing replacements. | Free availability depends on upstream marketing budgets. |
| July 2025 | Together ended ordinary free trials. | Mature inference shifted to prepaid use. |
| 2025-09-19 | Ollama Cloud preview began. | Free cloud history is only about one year. |
| 2025-12-07 | Google staff confirmed a 20-RPD free reduction. | Google free API quotas can shrink without supporting production use. |
| 2026-03-15 | Chutes retired its 200-RPD Early Access plan. | Loss-making free inference was removed. |
| 2026-07-30 | GitHub Models fully retired. | Whole free gateway can disappear, not just one model. |
| 2026-07-28 | Cloudflare moved Kimi/GLM/DeepSeek frontier models to paid access. | Free allowance remained, but strongest models moved out. |
| 2026-08-16 | Groq ended free/developer access to Llama 3.1 8B and Llama 3.3 70B. | Model availability is not durable even when the provider survives. |
| 2026-08-31 | Ollama changed usage accounting to monthly token pools. | Capacity mechanics can change quickly. |
| 2026-09-03 | OpenCode lists Muse Spark 1.3 Contributor Free as limited-time. | Today’s attractive offer carries no longevity promise. |
| 2026-09-30 | OpenRouter’s Dots 3 Note free model is scheduled to expire. | Current catalog contains explicit near-term expiry. |
Reading the business models: small renewable credits that lead naturally to pay-as-you-go (Vercel, Mistral, Cloudflare, IBM) are the more durable kind; subsidized GPU giveaways to win developers (OpenRouter free routes, NVIDIA's hosted trial, OpenCode promotions) are the less durable kind; one-time signup credits are not capacity at all. GitHub, Chutes, and xAI ended their free offers rather than maintaining them.
Verdict, provider by provider
| Provider | Best legitimate UVU role | Hard limit | Data-term risk | Longevity risk | Recommendation |
|---|---|---|---|---|---|
| OpenRouter | Evaluation, Public-data fallback, model discovery | 50 RPD/20 RPM | MED: upstream varies | HIGH: subsidized rotation and 2025 withdrawals | Use experimentally; pay for controlled routing if adopted. |
| Google Gemini API Free | CS labs and voluntary experiments | Exact quota private/per project | HIGH: unpaid data may train; human review | HIGH: staff-confirmed quota cuts | Never send Sensitive data; do not size infrastructure from it. |
| Groq Free | Fast open-model labs and public prototypes | 200K TPD/org for strong text models | LOW: no training, ZDR available | MED: free models rotate | Best direct free API, but not a campus floor. |
| Cerebras | Thirty-day performance evaluation | $5/30 days | LOW | HIGH: renewable tier already ended | Trial only. |
| SambaNova | Conditional model-speed lab | 20 RPD/model in docs, but signup conflict | MED | HIGH: current access contradicted | Verify manually before any course promise. |
| Together | Metered open-model provider | No ordinary free tier | LOW–MED | HIGH for free | Consider paid, not free. |
| Fireworks | One-time benchmark/evaluation | $1 credit | LOW–MED | HIGH: finite credit | Trial only; paid candidate. |
| Mistral | EU-hosted model evaluation and course labs | $10/month; private rate caps | MED; Labs HIGH | HIGH: hidden caps/silent Labs updates | Useful sandbox; contract/pay for production. |
| Hugging Face | Small interoperability tests | $0.10/month | MED: upstream terms | HIGH: allowance was cut | Catalog/research hub, not capacity. |
| Cloudflare | Public-data edge inference and burst valve | 10K neurons/day/account | LOW | MED: recurring since 2023, but frontier gating | Strong free complement; paid plan for reliability. |
| GitHub Models | None | Retired | N/A | HIGH: already gone | Remove from all plans. |
| NVIDIA hosted NIM | Prototype models before self-hosting | Numeric allowance unknown | HIGH for trial | HIGH: pre-release/changeable | Evaluation only. |
| OpenCode Zen | Developer evaluation, especially Muse | Limits unknown; limited-time | HIGH: contributor training | HIGH | Public/synthetic prompts only. |
| Meta Model API | Paid Muse evaluation | No direct verified free tier | Contributor HIGH | HIGH for free | Treat Contributor as paid data-for-discount, not free. |
| Ollama Cloud | Personal experimentation | 1 concurrent; allowance undisclosed | LOW | HIGH: recent limit redesign | Good personal benefit; cannot capacity-plan. |
| DeepInfra | Cheap metered open API | No free tier | LOW–MED | HIGH for free | Serious paid candidate. |
| Novita | Small-model development | Unknown voucher/free limits | LOW | HIGH: dynamic zero-price models | Sandbox only. |
| Hyperbolic | Short evaluation | $1, 60 RPM | LOW | HIGH: one-time | Trial only. |
| Chutes | Paid specialist inference | Free plan retired | LOW–MED | HIGH: already ended | Do not count as free. |
| Kluster | None pending evidence | Current service unknown | UNKNOWN | HIGH | Exclude until vendor supplies current terms. |
| Cohere Trial | Course demonstration/evaluation | 1,000 calls/month; non-production | HIGH: trial may support R&D | HIGH: allowance cut | Synthetic/Public data only. |
| xAI | Paid Grok testing | No free tier | LOW–MED | HIGH for free | Paid-only consideration. |
| Anthropic | Research grants or paid Claude | No recurring free API | LOW commercial | HIGH for free | Pursue grants; price Claude for Education separately. |
| OpenAI | Research grants, Codex student benefit, paid API | No free generation tier | LOW commercial | HIGH for free | Use education/research programs, not imagined API quota. |
| Vercel AI Gateway | $5/month model evaluation and fallback testing | Dollar cap; throughput unknown | MED: upstream chain | MED: renewable conversion funnel | Strong missed free option; institutional paid use needs review. |
| Z.AI | Public/synthetic GLM experiments | Quotas private | UNKNOWN–HIGH | HIGH: free conversion path | Experimental only. |
| Alibaba Model Studio | 90-day Qwen course/research trial | Usually 1M tokens/model | UNKNOWN | HIGH: new-user grant | Time-boxed course or benchmark. |
| IBM watsonx.ai Lite | Granite/open-model teaching labs | 300K tokens/month; one user | UNKNOWN | MED–HIGH: evaluation plan | Worth a controlled faculty pilot, not campus capacity. |
Survey: working paper 33, September 3, 2026 — every limit read from the provider's own page, every quoted clause from its terms; "unknown" marks where a number lives only behind a login. No accounts were created and no requests were sent; the numbers are what the providers publish, not what they deliver under load.
The harness market, extensively
Appendix M in five lines
- QuestionWhich chat, coding, research, and work tools are used, rising, or dying, and which fit the university?
- AnswerThe review supports the current core tools while new tools are tested around them.
- Deciding numbers171 toolsest86 code projects checkedest4 archived projects and 9 with no code update in over 90 daysest
- What the plan doesKeep the current core, test new tools for 90 days, and reject shut-down or hype-only tools.
- Still unknownFast risers still lack proven users, owners, safety, and campus controls; the social scan found no campus-wide university use.
Appendix L compared twenty-five harnesses in depth. This appendix widens the lens to the whole market as of September 3, 2026: 171 products and projects, each with its form, license and first release, owner or backer, public scale, growth and release cadence, disclosed capital, stability signals, any rating or benchmark placement, provider and local-model support, identity and budget controls, data and telemetry defaults, and a university verdict. Three feeds were merged: a dedicated survey lane over GitHub trending, awesome-lists, marketplaces, package registries, funding announcements, and press; a live pull from the GitHub API for 86 repositories (below); and an X trend scan whose findings appear as their own section below. Every cell carries its grade: fact a source states it, est derived, unknown not independently verified. Stars are an interest signal, never proof of adoption, quality, or safe operation.
The short version
- Reach is concentrated in the big consumer and enterprise assistants — ChatGPT, Gemini, Meta AI, GitHub Copilot, Microsoft 365 Copilot, Replit, Gemini Notebook, Claude. Open-source attention concentrates in Ollama, AutoGPT, Dify, Langflow, Open WebUI, LangChain, llama.cpp, OpenCode, and n8n.
- Funding is concentrating around foundation-model vendors and coding agents: OpenAI, Anthropic, Cognition, Replit, Lovable, Cursor, CodeRabbit, Ollama, Braintrust, Qodo.
- The sharpest new GitHub spikes — DeepSeek Harness, Ponytail, Graphify, Caveman, ECC, Paperclip, OpenMAIC, openclaude — need provenance and real-usage validation before any procurement conversation; several show star patterns no organic project produces.
- Clear retirement or migration risk: Flowise, Roo Code, Continue, OpenAI Swarm, the old AutoGen and Semantic Kernel agent paths, Sourcegraph Cody's retired tiers, Amazon Q CLI, and the consumer Gemini CLI path.
- The plan's base still maps well: LiteLLM as the routing boundary, LibreChat as the front door, llama.cpp as the floor, OpenCode/Cline-class coding, Onyx-class research. The market evidence argues for testing more harnesses around that boundary, not replacing it.
Four rankings — how the market is moving
All four are market-motion scores on a 0–100 scale est, not quality, security, or procurement scores. Method:
- Adoption: 55% active/paid-user evidence, 20% installs or package use, 15% repository reach, 10% organization/revenue evidence. Missing dimensions were renormalized. Vendor claims received a confidence discount.
- Growth: 70% percentile rank of the best recent user, install, download, or star-growth measure; 30% recency and evidence quality. Shorter intervals were annualized only for ranking and are shown explicitly.
- Funding momentum: 65% log-scaled disclosed amount, 20% recency, 15% strategic or revenue evidence. Acquisitions without a disclosed price received no invented amount.
- New and rising: limited to products first released or materially relaunched within nine months; 60% recent velocity, 20% age, 20% release activity.
- These are market-motion scores, not security, quality, accessibility, or procurement scores.
| # | Absolute adoption | 90-day growth (or closest public interval) | Funding momentum | New and rising, ≤9 months |
|---|---|---|---|---|
| 1 | est 100 ChatGPT — >1B WAU | est 100 DeepSeek Harness — +210,708 stars/22d; anomaly | est 100 OpenAI — $122B, 2026-03-31 | est 100 DeepSeek Harness — 210,708 stars/22d; validation required |
| 2 | est 98 Gemini — >1B MAU | est 96 Ponytail — 122,907 stars/84d | est 98 Anthropic — $65B, 2026-05-28 | est 96 Ponytail — 122,907 stars since 2026-06-12 |
| 3 | est 92 Meta AI — 700M MAU, stale 2025 evidence | est 91 Dify — ≈16.8K stars/month over 97d | est 88 Cognition — >$1B, 2026-05-27 | est 93 Graphify — 114,218 stars since 2026-04-03 |
| 4 | est 90 GitHub Copilot — 50M users | est 88 OpenCode — ≈12K stars/month over 118d | est 86 Lovable — $400M, 2026-08-12 | est 91 Caveman — 102,927 stars since 2026-04-04 |
| 5 | est 88 Replit — >50M platform users | est 86 Codex — +30,401 stars/86d | est 84 Replit — $400M, 2026-03-11 | est 89 ECC — 246,773 stars since 2026-01-18 |
| 6 | est 86 Microsoft 365 Copilot — >30M paid seats | est 84 OpenMAIC — +9,426 stars/7d | est 82 Cursor/SpaceX — acquisition 2026-08-14; prior $900M round | est 87 Paperclip — 79,934 stars since 2026-03-02 |
| 7 | est 85 Gemini Notebook — >30M users | est 82 qm — +6,712 stars/7d | est 80 CodeRabbit — $143M, 2026-08-12 | est 84 OpenMAIC — +9,426 stars/week |
| 8 | est 84 Claude/Claude Code — 39% survey share; use doubled | est 80 Cloudflare OS — +9,549/month | est 78 Ollama — $88M, 2026-07-09 | est 82 openclaude — +1,035/week; 32,228 stars |
| 9 | est 82 LangChain — 229.9M PyPI downloads/month | est 79 Cloudflare Computer — +8,868/month | est 76 Braintrust — $80M, 2026-02-17 | est 80 nanobot — 47,685 stars since 2026-02-01 |
| 10 | est 81 Lovable — 60M projects; 900M visits/month | est 76 Paperclip — ≈4,962 stars/month | est 74 Qodo — $70M, 2026-03-30 | est 78 codebase-memory-mcp — 42,025 stars since 2026-02-24 |
| 11 | est 80 Codex — 1.6M weekly users | est 75 Browser Use — ≈4,860 stars/month | est 72 Arize — $70M, 2025-02 | est 76 QwenPaw — 34,846 stars since 2026-02-24 |
| 12 | est 78 Cursor — 12% survey; hundreds of millions weekly requests | est 74 llama.cpp — ≈4,260 stars/month | est 70 Gumloop — $50M, 2026-03-12 | est 74 Prime Agent — 19,733 stars since 2026-05-08 |
| 13 | est 76 Vercel AI SDK — 92.1M npm downloads in Aug | est 72 Browser Harness — ≈3,730 stars/month | est 68 Consensus — $30M, 2026-05-11 | est 72 Cloudflare OS — 9,539 stars since 2026-04-15 |
| 14 | est 74 OpenCode — 7% survey; 1.85M npm/week | est 70 OpenHands — ≈3,330 stars/month | est 67 Dify — $30M, 2026-03-10 | est 71 Cloudflare Computer — 8,954 stars since 2026-06-05 |
| 15 | est 73 Ollama — 8.9M developers claimed | est 69 Open WebUI — ≈3,100 stars/month | est 66 Mastra — $22M, 2026-04-09 | est 70 qm — 14,523 stars since 2026-07-29 |
| 16 | est 71 CrewAI — 26.5M PyPI/month | est 64 RAGFlow — ≈2,520 stars/month | est 65 LangChain — $125M, 2025-10-20 | est 66 OpenHarness — 15,633 stars but already stale |
| 17 | est 70 Cline — >5M installs | est 63 Ollama — ≈2,430 stars/month | est 63 OpenHands — $18.8M, 2025-11-18 | est 64 MiMo Code — ≈4,580 stars/month |
| 18 | est 68 Strands — 14M downloads claimed | est 59 CrewAI — ≈1,470 stars/month | est 61 Cline — $32M, 2025-07-31 | est 62 Kimi Code — 7,238 stars since 2026-05-22 |
| 19 | est 67 Dify — 1.4M machines claimed; 154K stars | est 58 Cline — ≈1,450 stars/month | est 60 Glean — $150M prior round plus $300M ARR | est 60 open-connector — 5,524 stars since 2026-06-29 |
| 20 | est 66 Open WebUI — 150.8K stars and strong container use | est 57 Langflow — ≈1,360 stars/month | est 58 Browserbase — $40M, 2025-06 | est 58 fx — new 2026-08-11 |
| 21 | est 65 n8n — 203K stars; strong commercial growth | est 54 Crush — ≈952 stars/month | est 56 Relevance AI — $24M, 2025-05 | est 57 Browser Harness — 17.3K stars since 2026-04-17 |
| 22 | est 64 LangGraph — 13.15M npm downloads in Aug | est 53 Aider — ≈872 stars/month | est 54 LlamaIndex — $19M, 2025-03 | est 56 Antigravity CLI — 6% survey within four months |
| 23 | est 63 Glean — $300M ARR and 45% wDAU/wMAU | est 51 Qwen Code — ≈804 stars/month | est 52 Onyx — $10M seed, 2025 | est 54 IBM Bob — GA 2026-04; 80K internal IBM users |
| 24 | est 62 Elicit — >400K monthly researchers | est 49 Mastra — ≈663 stars/month sampled | est 50 Elicit — $22M, 2025-02 | est 52 Bedrock AgentCore — GA 2026-06-18 |
| 25 | est 60 AnythingLLM — >5M Docker pulls | est 47 LobeHub — ≈663 stars/month sampled | est 48 CrewAI — >$18M, 2024-10 | est 50 Meta Muse Code — beta 2026-08-05 |
The growth and new-entrant columns deserve special caution: the DeepSeek Harness, Ponytail, ECC, Graphify, Caveman, OpenClaw, and qm star patterns are large enough that stars should not be accepted as adoption without package use, identifiable users, repository provenance, and contribution-distribution checks.
Declining or dead — with dates
| Product / path | State | Effective date | Evidence | UVU action |
|---|---|---|---|---|
| Flowise | fact frozen, archived and ended | 2026-08-31 | Official EOL discussion | Do not start; export existing flows |
| Roo Code | fact repository archived/shutdown | 2026-05-15 | Archived repository | Do not start; migrate maintained forks only after review |
| Continue | fact no longer actively maintained | current 2026-09-03 | Repository notice | Do not make a campus standard |
| OpenAI Agent Builder | fact scheduled unavailable | 2026-11-30 | AgentKit notice | Do not build new durable workflows |
| OpenAI Swarm | fact replaced by Agents SDK | 2025 | Swarm repository | Use Agents SDK or another framework |
| AutoGen | fact maintenance mode | current 2026-09-03 | AutoGen repository | New Microsoft work goes to Agent Framework |
| Semantic Kernel agent path | fact new agent work moving to Agent Framework | 2026 | Microsoft Agent Framework | Keep only supported existing workloads |
| Sourcegraph Cody Free/Pro/Enterprise Starter | fact unavailable | 2025-07-23 | Cody FAQ | Do not plan around retired tiers |
| Amazon Q CLI | fact superseded by Kiro CLI | current 2026 | Migration guide | Evaluate Kiro separately; do not assume compatibility |
| Gemini CLI consumer service path | fact ended/transitioned to Antigravity CLI | 2026-06-18 | Google transition notice | Freeze before adopting either path campus-wide |
| SWE-agent full harness | fact directs new users to mini-SWE-agent | current 2026 | Repository | Use for research replication, not a campus default |
| HuggingChat original service | fact shut down, later replaced by Omni | 2025-07-01 / 2025-10-16 | Hugging Face Chat | Treat Omni as a new service review |
| Quivr | fact latest release 2025-02-04 | 2025-02-04 | Repository | Avoid new standardization |
| MetaGPT | fact last release 2025-03; last push 2026-01-21 | 2026-01-21 | Repository | Research reference only |
| GPT4All | fact latest release 2025-02-25 | 2025-02-25 | Repository | Prefer Ollama, llama.cpp, Jan or LM Studio |
| Ragas | fact no release in last 90 days; last push 2026-02-24 | 2026-02-24 | Repository | Do not rely on it as the sole evaluation layer |
| AgentOps | fact latest GitHub release 2025-08 | 2025-08 | Repository | Prefer Langfuse, Phoenix, Promptfoo or Opik |
| OpenHarness | fact last push 2026-06-04, despite April launch | 2026-06-04 | GitHub snapshot | Treat as stalled until maintenance resumes |
Open-source status also needs care:
- Open WebUI’s custom license restricts rebranding for deployments with 50 or more users. License explanation
- Dify’s modified Apache license restricts certain multi-tenant and branding uses. Dify license
- n8n moved to its Sustainable Use License. License announcement
- Crush uses FSL-1.1-MIT, which is source-available rather than immediately permissive.
- Phoenix’s server uses the Elastic License 2.0.
- These are not necessarily blockers for internal university use, but they are procurement and redistribution constraints.
Stability and longevity — the top 40, scored on six axes
L, M, H mean low, medium, or high risk, not product quality. Axes: VD vendor/platform dependence · LIC license or future-use risk · RUN funding/runway · BUS maintainer concentration · GOV governance concentration · BRK breaking-change churn. A dated desk assessment, not a security audit.
| # | Harness | Overall | VD | LIC | RUN | BUS | GOV | BRK | Dated basis |
|---|---|---|---|---|---|---|---|---|---|
| 1 | ChatGPT | M | H | H | L | L | H | M | fact >1B WAU and $122B committed by 2026-03-31; closed service |
| 2 | Claude | M | H | H | L | L | H | M | fact $65B Series H, 2026-05-28; closed Anthropic service |
| 3 | Gemini | M | H | H | L | L | H | M | fact >1B MAU, 2026-08-11; Google-controlled |
| 4 | Microsoft 365 Copilot | M | H | H | L | L | H | M | fact >30M paid seats, 2026-07-29; Microsoft-controlled |
| 5 | GitHub Copilot | M | H | H | L | L | H | M | fact 50M users, 2026-07-29; service and policy controlled by Microsoft/GitHub |
| 6 | OpenAI Codex | M | M | L | L | L | H | H | fact OSS CLI plus closed service; version 0.153 by 2026-09-03 |
| 7 | Claude Code | M | H | H | L | L | H | H | fact fast releases through 2.1.259; Anthropic routes only |
| 8 | Cursor | H | H | H | L | M | H | H | fact joined SpaceX 2026-08-14; survey share declined from 18% to 12% |
| 9 | Replit Agent | H | H | H | L | M | H | H | fact $400M round 2026-03-11; vertically integrated platform |
| 10 | Lovable | H | H | H | L | M | H | H | fact $400M Series C, 2026-08-12; fast-moving hosted platform |
| 11 | Open WebUI | H | L | H | M | H | H | H | fact custom license; multiple advisories patched June–Aug 2026 |
| 12 | LibreChat | M | L | L | H | M | M | M | fact MIT and active; unknown institutional funding; R90 only 4 |
| 13 | Ollama | M | M | L | L | M | H | H | fact MIT, $88M raised 2026-07-09, R90 29 |
| 14 | llama.cpp | M | L | L | L | M | M | H | fact MIT/community; R90 ≥100; HF organization since 2026-02-20 |
| 15 | Cline | M | L | L | M | M | M | H | fact Apache-2.0; $32M; npm incident and later advisories patched in 2026 |
| 16 | OpenCode | M | L | L | H | M | H | H | fact MIT, R90 55; unknown funding and institutional governance |
| 17 | OpenHands | M | L | M | M | M | M | H | fact MIT core, $18.8M Series A, R90 33 |
| 18 | Goose | M | L | L | M | L | L | M | fact Apache-2.0 and foundation home; safety switches off by default |
| 19 | Dify | H | L | H | M | M | H | M | fact modified license; $30M round 2026-03-10 |
| 20 | Langflow | M | L | L | L | M | M | H | fact MIT/active; August 2026 advisories patched |
| 21 | n8n | H | L | H | L | M | H | H | fact source-available license, $240M raised, R90 ≥100 |
| 22 | LangChain | M | L | L | L | M | H | H | fact MIT, $125M Series B, R90 86 |
| 23 | LangGraph | M | L | L | L | M | H | M | fact MIT; parent well-funded; active evolution |
| 24 | LlamaIndex | M | L | L | M | M | H | M | fact MIT; $27.5M total as of 2025-03 |
| 25 | CrewAI | M | L | L | M | M | H | H | fact MIT; >$18M; R90 37; telemetry on |
| 26 | Microsoft Agent Framework | M | L | L | L | L | H | M | fact MIT, Microsoft-backed, 1.0 on 2026-04-03 |
| 27 | PydanticAI | M | L | L | M | M | M | H | fact MIT; parent funded; R90 56 |
| 28 | Google ADK | M | M | L | L | L | H | M | fact Apache-2.0, Google-backed, R90 20 |
| 29 | Vercel AI SDK | M | L | U | L | M | H | H | fact 23.6M npm/week; R90 ≥100; license not frozen in evidence set |
| 30 | Mastra | M | L | M | M | M | H | H | fact core/EE split; $35M total; R90 22 |
| 31 | Onyx | M | L | M | M | M | M | M | fact MIT core/EE split; $10M seed; active university reference |
| 32 | RAGFlow | M | L | L | H | M | H | M | fact Apache-2.0; 1,581 open issue/PR count; funding unknown |
| 33 | Elicit | M | H | H | M | M | H | M | fact $22M Series A; closed research service |
| 34 | Glean | M | H | H | L | L | H | M | fact $300M ARR on 2026-05-28; closed enterprise service |
| 35 | Gemini Notebook | M | H | H | L | L | H | M | fact >30M users; Google-controlled/rebranded 2026-07-16 |
| 36 | Langfuse | M | L | M | L | M | H | H | fact acquired by ClickHouse 2026-01; R90 ≥100 |
| 37 | Phoenix | H | L | H | L | M | H | H | fact ELv2 server, auth off and analytics on by default; R90 97 |
| 38 | Braintrust | M | M | H | L | M | H | M | fact $80M Series B 2026-02-17; managed control plane |
| 39 | Promptfoo | M | L | L | L | M | H | M | fact MIT; OpenAI acquisition agreement 2026-03-09 |
| 40 | Paperclip | H | L | L | H | H | H | H | fact created 2026-03-02; 5,370 open issue/PR count; funding/governance unknown |
Category maps
| Category | Current leader | Fastest riser | Safest university choice | Wildcard |
|---|---|---|---|---|
| Chat/front door | est ChatGPT by weekly reach; Gemini close by monthly reach | est Gemini/Claude by current user and business growth | est LibreChat behind LiteLLM for provider control; ChatGPT Edu for managed SaaS | est Duck.ai for low-risk, privacy-oriented public use |
| Coding | est GitHub Copilot by users and installs | est Codex and Claude Code by current use growth; OpenCode in OSS | est Cline Enterprise or governed OpenCode through LiteLLM | est Goose because it is provider-neutral and foundation-backed |
| Browser/computer use | est ChatGPT agent by parent-product reach | est Browser Harness and Cloudflare Computer by new-project velocity | est Playwright MCP inside a locked-down managed runtime | est Cloudflare Computer |
| Frameworks | est LangChain by downloads and ecosystem | est Microsoft Agent Framework among governed entrants; Paperclip by raw stars | est Haystack or Microsoft Agent Framework, depending language/cloud | est Paperclip, after provenance and maintenance checks |
| RAG/research | est Glean for enterprise knowledge; Gemini Notebook for personal research reach | est Onyx/Dify among deployable platforms; Elicit among research tools | est Onyx for campus search; Gemini Notebook for governed individual study | est Elicit Research Agent |
| Workflow | est n8n by OSS reach and commercial growth | est Gumloop by recent funding; Activepieces by package growth | est n8n self-hosted with license review, or Activepieces | est Gumloop |
| Local runtime/desktop | est Ollama by developer reach | est llama.cpp by repository velocity | est llama.cpp as the technical floor; Ollama as the managed developer layer | est Jan for a telemetry-off desktop |
| Evaluation/observability | est Langfuse by OSS/package reach | est Opik by releases; Promptfoo by use/acquisition | est Promptfoo offline plus Phoenix behind authentication | est MLflow where UVU already operates ML infrastructure |
Fit to this plan — pilot, watch, avoid
The market evidence does not justify replacing the architecture. The smaller, safer move is to keep the router as the boundary and test additional harnesses around it.
| Category | Pilot (2–3) | Watch | Avoid now | Why |
|---|---|---|---|---|
| Coding | Cline Enterprise; OpenCode through LiteLLM; Goose or OpenHands | Antigravity; IBM Bob; Paperclip | Roo Code; Continue; Grok Build | The pilots cover IDE, terminal and autonomous modes while retaining provider choice. The avoided set is retired or too new to govern. |
| Browser/computer use | Playwright MCP in a sandbox; ChatGPT agent Enterprise with narrow permissions; Browserbase with contract limits | Cloudflare Computer; Browser Harness | Logged-in Browser Use deployments without telemetry and egress controls | Browser actions create a larger data and authority boundary than ordinary chat. |
| Frameworks | Microsoft Agent Framework; LangGraph; Haystack | Paperclip; Antigravity SDK; Bedrock AgentCore | New AutoGen, Semantic Kernel agent work, OpenAI Swarm | The pilots cover .NET/Python, graphs and mature RAG while avoiding announced migration paths. |
| RAG/research | Onyx; Gemini Notebook under Workspace; Elicit | RAGFlow; Consensus | Flowise; Quivr | Onyx matches UVU’s current research tier. Gemini Notebook and Elicit cover individual source-grounded research. |
| Workflow | n8n self-hosted; Activepieces; Dify self-hosted | Gumloop; Relevance AI | OpenAI Agent Builder; unreviewed acquired platforms | These pilots offer useful visual workflows while preserving deployment choice. License terms must be checked before shared-service use. |
| Local | llama.cpp; Ollama; Jan or LM Studio | LocalAI | GPT4All and Text generation web UI as campus standards | llama.cpp remains the strongest Apple-silicon floor. Ollama improves developer experience. Jan offers a telemetry-off desktop path. |
| Evaluation | Promptfoo offline; Phoenix with authentication enabled; Langfuse with telemetry disabled or contractually approved | Braintrust; Opik | Ragas, AgentOps and Portkey as the sole standard | This gives UVU regression testing, traces, red-team checks and a self-host route without tying all proof to one vendor. |
Pilot gates that apply to every harness:
- Route provider calls through LiteLLM unless the pilot requires a documented exception.
- Use university identity, role separation, audit logs, budget caps and short retention.
- Test FERPA exposure, accessibility, source attribution, prompt-injection handling, export and deletion.
- Deny browser/file/shell/network permissions by default; enable only the minimum needed.
- Record model, harness, version, policy and evaluation set together. A model benchmark is not a harness benchmark.
- Require an exit path: configuration export, data export, provider substitution and rollback.
- Run pilots for 90 days before changing the campus standard.
Live GitHub snapshot — 86 repositories, pulled directly from the API
Our own pull on September 3, 2026, independent of the survey: stars, forks, creation and last-push dates, latest release, archive state. Four projects are archived and nine had no code pushed in over 90 days — including two the earlier survey had described as active. This is the check that keeps "popular" and "maintained" from being confused.
| Repository | Category | Stars | Forks | Created | Last push | Latest release | Flag | License |
|---|---|---|---|---|---|---|---|---|
| anomalyco/opencode | Coding / agent | 203,444 | 26,541 | 2025-04-30 | 2026-09-03 | 2026-09-02 | MIT | |
| n8n-io/n8n | Workflow builder | 203,223 | 60,535 | 2019-06-22 | 2026-09-03 | 2026-09-03 | NOASSERTION | |
| ollama/ollama | Local runner / serving | 180,043 | 17,668 | 2023-06-26 | 2026-09-03 | 2026-08-27 | MIT | |
| langgenius/dify | Workflow builder | 154,326 | 24,397 | 2023-04-12 | 2026-09-03 | 2026-08-25 | NOASSERTION | |
| langflow-ai/langflow | Workflow builder | 154,184 | 10,012 | 2023-02-08 | 2026-09-03 | 2026-09-01 | MIT | |
| open-webui/open-webui | Chat front door | 150,808 | 22,036 | 2023-10-06 | 2026-09-02 | 2026-08-31 | NOASSERTION | |
| anthropics/claude-code | Coding / agent | 143,897 | 23,003 | 2025-02-22 | 2026-09-02 | 2026-09-02 | — | |
| ggml-org/llama.cpp | Local runner / serving | 126,897 | 22,703 | 2023-03-10 | 2026-09-03 | 2026-08-25 | MIT | |
| openai/codex | Coding / agent | 121,174 | 18,567 | 2025-04-13 | 2026-09-03 | 2026-09-03 | Apache-2.0 | |
| browser-use/browser-use | Browser agent | 112,156 | 12,334 | 2024-10-31 | 2026-09-03 | 2026-08-16 | MIT | |
| google-gemini/gemini-cli | Coding / agent | 106,800 | 14,525 | 2025-04-17 | 2026-09-03 | 2026-09-01 | Apache-2.0 | |
| vllm-project/vllm | Local runner / serving | 90,881 | 21,653 | 2023-02-09 | 2026-09-03 | 2026-08-26 | Apache-2.0 | |
| modelcontextprotocol/servers | Other | 90,049 | 11,548 | 2024-11-19 | 2026-09-03 | 2026-08-31 | NOASSERTION | |
| infiniflow/ragflow | Research / RAG | 89,982 | 10,610 | 2023-12-12 | 2026-09-03 | 2026-08-28 | Apache-2.0 | |
| zed-industries/zed | Coding / agent | 89,700 | 10,420 | 2021-02-20 | 2026-09-03 | 2026-09-02 | NOASSERTION | |
| OpenHands/OpenHands | Coding / agent | 86,069 | 11,286 | 2024-03-13 | 2026-09-03 | 2026-08-27 | MIT | |
| lobehub/lobehub | Chat front door | 82,198 | 15,857 | 2023-05-21 | 2026-09-03 | 2026-08-28 | NOASSERTION | |
| daytonaio/daytona | Agent framework / infra | 71,829 | 5,650 | 2024-02-06 | 2026-07-24 | 2026-06-23 | — | |
| FoundationAgents/MetaGPT | Agent framework / infra | 70,197 | 8,921 | 2023-06-30 | 2026-01-21 | 2024-04-22 | stale >90d | MIT |
| openinterpreter/openinterpreter | Other | 68,230 | 5,873 | 2023-07-14 | 2026-08-20 | 2026-08-20 | Apache-2.0 | |
| cline/cline | Coding / agent | 67,402 | 7,280 | 2024-07-06 | 2026-09-03 | 2026-09-02 | Apache-2.0 | |
| Mintplex-Labs/anything-llm | Chat front door | 65,565 | 7,247 | 2023-06-04 | 2026-09-03 | 2026-08-27 | MIT | |
| microsoft/autogen | Agent framework / infra | 60,791 | 9,178 | 2023-08-18 | 2026-04-15 | 2025-09-30 | stale >90d | CC-BY-4.0 |
| crewAIInc/crewAI | Agent framework / infra | 58,045 | 8,326 | 2023-10-27 | 2026-09-03 | 2026-08-27 | MIT | |
| BerriAI/litellm | Observability / gateway | 57,934 | 11,131 | 2023-07-27 | 2026-09-03 | 2026-09-02 | NOASSERTION | |
| zylon-ai/private-gpt | Research / RAG | 57,489 | 7,614 | 2023-05-02 | 2026-09-02 | 2026-06-18 | Apache-2.0 | |
| FlowiseAI/Flowise | Workflow builder | 55,404 | 24,971 | 2023-03-31 | 2026-08-13 | 2026-07-29 | ARCHIVED | NOASSERTION |
| AntonOsika/gpt-engineer | Coding / agent | 55,111 | 7,293 | 2023-04-29 | 2025-05-14 | 2024-06-06 | ARCHIVED | MIT |
| aaif-goose/goose | Coding / agent | 53,879 | 6,161 | 2024-08-23 | 2026-09-03 | 2026-08-27 | Apache-2.0 | |
| CherryHQ/cherry-studio | Chat front door | 51,401 | 4,909 | 2024-05-24 | 2026-09-03 | 2026-09-03 | AGPL-3.0 | |
| mudler/LocalAI | Local runner / serving | 48,845 | 4,412 | 2023-03-18 | 2026-09-03 | 2026-08-20 | MIT | |
| Aider-AI/aider | Coding / agent | 48,698 | 4,920 | 2023-05-09 | 2026-05-22 | 2025-08-09 | stale >90d | Apache-2.0 |
| oobabooga/textgen | Local runner / serving | 47,614 | 5,983 | 2022-12-21 | 2026-08-17 | 2026-05-20 | AGPL-3.0 | |
| exo-explore/exo | Local runner / serving | 47,226 | 3,492 | 2024-06-24 | 2026-08-25 | 2026-04-23 | Apache-2.0 | |
| janhq/jan | Chat front door | 44,313 | 3,007 | 2023-08-17 | 2026-09-03 | 2026-07-23 | NOASSERTION | |
| danny-avila/LibreChat | Chat front door | 42,770 | 8,849 | 2023-02-12 | 2026-09-03 | — | MIT | |
| agno-agi/agno | Agent framework / infra | 42,031 | 5,865 | 2022-05-04 | 2026-09-03 | 2026-09-01 | Apache-2.0 | |
| chatboxai/chatbox | Chat front door | 41,633 | 4,224 | 2023-03-06 | 2026-09-02 | 2026-09-02 | GPL-3.0 | |
| langchain-ai/langgraph | Agent framework / infra | 40,992 | 6,916 | 2023-08-09 | 2026-09-03 | 2026-08-27 | MIT | |
| The-Vibe-Company/quivr | Research / RAG | 39,490 | 3,732 | 2023-05-12 | 2026-08-31 | 2025-02-04 | NOASSERTION | |
| CopilotKit/CopilotKit | Coding / agent | 37,181 | 4,601 | 2023-06-19 | 2026-09-03 | 2026-09-01 | MIT | |
| khoj-ai/khoj | Research / RAG | 37,036 | 2,452 | 2021-08-16 | 2026-08-02 | 2026-03-26 | AGPL-3.0 | |
| continuedev/continue | Coding / agent | 35,739 | 5,322 | 2023-05-24 | 2026-09-03 | 2026-06-19 | Apache-2.0 | |
| langfuse/langfuse | Observability / gateway | 34,155 | 3,688 | 2023-05-18 | 2026-09-03 | 2026-09-03 | NOASSERTION | |
| TabbyML/tabby | Coding / agent | 33,860 | 1,788 | 2023-03-16 | 2026-06-30 | 2026-01-25 | NOASSERTION | |
| sgl-project/sglang | Local runner / serving | 33,779 | 8,513 | 2024-01-08 | 2026-09-03 | 2026-08-22 | Apache-2.0 | |
| Pythagora-io/gpt-pilot | Coding / agent | 33,680 | 3,473 | 2023-08-16 | 2026-06-18 | — | NOASSERTION | |
| mckaywrigley/chatbot-ui | Chat front door | 33,342 | 9,421 | 2023-03-11 | 2024-08-03 | — | stale >90d | MIT |
| onyx-dot-app/onyx | Research / RAG | 31,910 | 4,399 | 2023-04-27 | 2026-09-03 | 2026-09-02 | NOASSERTION | |
| openai/openai-agents-python | Agent framework / infra | 29,169 | 4,658 | 2025-03-11 | 2026-09-02 | 2026-08-19 | MIT | |
| huggingface/smolagents | Agent framework / infra | 29,140 | 2,918 | 2024-12-05 | 2026-08-25 | 2026-05-29 | Apache-2.0 | |
| voideditor/void | Coding / agent | 28,814 | 2,644 | 2024-09-11 | 2026-06-02 | — | ARCHIVED | Apache-2.0 |
| charmbracelet/crush | Coding / agent | 27,876 | 2,218 | 2025-05-21 | 2026-09-03 | 2026-08-31 | NOASSERTION | |
| mastra-ai/mastra | Agent framework / infra | 27,668 | 2,725 | 2024-08-06 | 2026-09-03 | 2026-08-28 | NOASSERTION | |
| Kilo-Org/kilocode | Coding / agent | 27,156 | 3,114 | 2025-03-10 | 2026-09-03 | 2026-09-02 | MIT | |
| vercel/ai | Agent framework / infra | 26,564 | 5,069 | 2023-05-23 | 2026-09-03 | 2026-09-02 | NOASSERTION | |
| huggingface/open-r1 | Local runner / serving | 26,449 | 2,448 | 2025-01-24 | 2026-04-02 | — | stale >90d | Apache-2.0 |
| Cinnamon/kotaemon | Research / RAG | 25,729 | 2,155 | 2024-03-25 | 2026-07-14 | 2026-05-31 | Apache-2.0 | |
| letta-ai/letta | Agent framework / infra | 24,601 | 2,612 | 2023-10-11 | 2026-08-23 | 2026-05-14 | Apache-2.0 | |
| browserbase/stagehand | Browser agent | 24,135 | 1,665 | 2024-03-24 | 2026-09-03 | 2026-08-28 | MIT | |
| Skyvern-AI/skyvern | Browser agent | 22,925 | 2,153 | 2024-02-28 | 2026-09-03 | 2026-08-31 | AGPL-3.0 | |
| PromtEngineer/localGPT | Research / RAG | 22,206 | 2,462 | 2023-05-24 | 2026-08-26 | — | MIT | |
| stackblitz-labs/bolt.diy | Coding / agent | 19,839 | 10,561 | 2024-10-13 | 2026-02-07 | 2025-05-12 | stale >90d | MIT |
| pydantic/pydantic-ai | Agent framework / infra | 19,699 | 2,634 | 2024-06-21 | 2026-09-03 | 2026-09-03 | MIT | |
| agent0ai/agent-zero | Coding / agent | 19,077 | 3,779 | 2024-06-10 | 2026-09-02 | 2026-08-27 | NOASSERTION | |
| avante-corp/avante.nvim | Coding / agent | 18,147 | 848 | 2024-08-14 | 2026-08-28 | 2026-08-27 | Apache-2.0 | |
| camel-ai/camel | Agent framework / infra | 17,672 | 2,066 | 2023-03-17 | 2026-09-03 | 2026-03-22 | Apache-2.0 | |
| plandex-ai/plandex | Coding / agent | 15,619 | 1,177 | 2023-10-24 | 2025-10-03 | 2025-07-16 | stale >90d | MIT |
| nanobrowser/nanobrowser | Browser agent | 13,720 | 1,451 | 2024-12-31 | 2026-08-18 | 2025-11-22 | Apache-2.0 | |
| e2b-dev/E2B | Agent framework / infra | 13,664 | 1,018 | 2023-03-04 | 2026-09-02 | 2026-09-02 | Apache-2.0 | |
| Portkey-AI/gateway | Observability / gateway | 12,890 | 1,282 | 2023-08-23 | 2026-05-25 | 2026-01-12 | stale >90d | MIT |
| bentoml/OpenLLM | Local runner / serving | 12,525 | 837 | 2023-04-19 | 2026-08-31 | 2025-04-21 | Apache-2.0 | |
| simonw/llm | Chat front door | 12,459 | 972 | 2023-04-01 | 2026-09-02 | 2026-09-02 | Apache-2.0 | |
| LostRuins/koboldcpp | Local runner / serving | 11,603 | 758 | 2023-03-16 | 2026-09-03 | 2026-08-29 | AGPL-3.0 | |
| Arize-ai/phoenix | Observability / gateway | 11,307 | 1,094 | 2022-11-09 | 2026-09-03 | 2026-09-03 | NOASSERTION | |
| github/copilot-cli | Coding / agent | 11,136 | 1,910 | 2023-01-06 | 2026-09-02 | 2026-08-29 | NOASSERTION | |
| sigoden/aichat | Chat front door | 10,425 | 743 | 2023-03-03 | 2026-02-23 | 2025-07-06 | stale >90d | Apache-2.0 |
| microsoft/magentic-ui | Coding / agent | 10,083 | 1,016 | 2025-05-05 | 2026-09-03 | 2026-05-21 | MIT | |
| microsoft/vscode-copilot-chat | Coding / agent | 9,971 | 2,011 | 2025-06-10 | 2026-05-20 | 2026-04-07 | ARCHIVED | MIT |
| ml-explore/mlx-lm | Local runner / serving | 6,882 | 1,018 | 2025-03-11 | 2026-09-03 | 2026-04-22 | MIT | |
| olimorris/codecompanion.nvim | Coding / agent | 6,833 | 449 | 2023-12-27 | 2026-08-31 | 2026-08-24 | Apache-2.0 | |
| AgentOps-AI/agentops | Observability / gateway | 5,810 | 618 | 2023-08-15 | 2026-06-25 | 2025-08-29 | MIT | |
| lmstudio-ai/lms | Local runner / serving | 5,258 | 450 | 2024-04-15 | 2026-09-01 | — | MIT | |
| ag2ai/ag2 | Agent framework / infra | 4,900 | 714 | 2024-11-11 | 2026-09-02 | 2026-08-28 | Apache-2.0 | |
| Fannovel16/comfyui_controlnet_aux | Other | 4,171 | 373 | 2023-08-17 | 2026-08-27 | — | Apache-2.0 | |
| OpenHands/OpenHands-Cloud | Coding / agent | 76 | 44 | 2025-02-17 | 2026-09-03 | 2026-09-02 | NOASSERTION |
The master table — all 171
Chat and assistant front doors — 16
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact ChatGPT | fact consumer/enterprise assistant | fact closed; 2022-11-30 | fact OpenAI | fact >1B weekly users, 2026-08-31 | fact active continuous service | fact $122B committed, 2026-03-31; unknown team | est high vendor dependence; strong runway | unknown no comparable marketplace score | fact OpenAI-hosted; no local | fact Enterprise/Edu identity, admin, retention and usage controls | fact plan-dependent; enterprise data excluded from training by default | est everyone, governed tier |
| fact Claude | fact consumer/enterprise assistant | fact closed; 2023-03-14 | fact Anthropic | fact Claude Code weekly use doubled since 2026-01; business subscriptions quadrupled | fact active | fact $30B Series G, 2026-02-12; $65B Series H, 2026-05-28 | est low runway risk; high vendor dependence | unknown no comparable rating | fact Anthropic/Bedrock/Vertex; no local | fact Enterprise SSO, SCIM, audit and spend controls | fact commercial/API data policy differs from consumer plan | est everyone, governed tier |
| fact Gemini | fact consumer/Workspace assistant | fact closed; Bard 2023-03-21; Gemini rename 2024-02-08 | fact Google | fact >1B monthly users, 2026-08-11 | fact active | fact Alphabet-backed; unknown standalone funding/team | est strong longevity; high ecosystem dependence | unknown no comparable rating | fact Google models; no local | fact Workspace identity, audit, DLP and admin controls | fact Workspace protections differ from consumer use | est everyone, governed tier |
| fact Microsoft 365 Copilot | fact enterprise assistant | fact closed; 2023-11-01 GA | fact Microsoft | fact >30M paid seats, 2026-07-29 | fact active | fact Microsoft-backed | est strong longevity; Microsoft lock-in | unknown no comparable rating | fact Microsoft-hosted; multiple partner models in some features | fact Entra, Purview, audit, DLP and budget administration | fact commercial data protections | est everyone where M365 is standard |
| fact Perplexity | fact search/research assistant | fact closed; 2022 | fact Perplexity AI | unknown current active-user count | fact active | unknown complete current funding/team verification | est medium vendor risk | unknown no comparable rating | fact hosted multi-model; no local | fact Enterprise SSO/SCIM, roles, credit and audit controls | fact enterprise no-training commitments; consumer terms differ | est researchers |
| fact Poe | fact multi-bot chat front door | fact closed; 2022-12 | fact Quora | unknown current active-user count | fact active | fact Quora-backed; unknown Poe-specific funding | est medium/high third-party-bot risk | unknown no comparable rating | fact many hosted providers; no local | unknown complete campus-grade control set | fact third-party bot operators can have separate data practices | est not suitable for regulated use |
| fact Le Chat | fact chat/enterprise assistant | fact closed; 2024-02 | fact Mistral AI | unknown active-user count | fact active; Vibe coding brand changed 2026-05-28 | unknown current round total in this sweep | est medium vendor risk | unknown no comparable rating | fact Mistral hosted; enterprise/on-prem options | fact enterprise identity and deployment options | unknown consumer telemetry default not fully verified | est everyone, controlled pilot |
| fact Meta AI | fact consumer assistant | fact closed; 2023-09-27 | fact Meta | fact 700M MAU reported 2025-03; stale | unknown current standalone growth | fact Meta-backed | est strong runway; high account/data coupling | unknown no comparable rating | fact Meta/Llama hosted; no local front door | unknown university control plane | fact consumer service tied to Meta privacy terms | est not suitable as campus standard |
| fact Grok | fact consumer/enterprise assistant | fact closed; 2023-11 | fact xAI | unknown standalone active users | fact active | fact $20B Series E announced 2026-01-06; unknown team | est strong runway; high vendor/platform risk | unknown no comparable rating | fact xAI-hosted; no local | unknown full university control set | unknown default telemetry and training treatment by plan | est not suitable pending controls review |
| fact DeepSeek Chat | fact consumer assistant | fact closed service; 2025-01 | fact DeepSeek/High-Flyer | unknown active-user count | fact active | unknown outside funding; privately backed | est jurisdiction and data-governance risk | unknown no comparable rating | fact DeepSeek-hosted; models can be run separately locally | unknown campus-grade identity/audit | fact policy covers prompts, files and history collection | est not suitable for protected data |
| fact Qwen Chat | fact consumer assistant | fact closed service; 2023 | fact Alibaba Cloud | unknown active-user count | fact active | fact Alibaba-backed | est high jurisdiction/vendor dependence | unknown no comparable rating | fact Qwen hosted; separate open-weight local models | unknown campus-grade controls | unknown default consumer telemetry not fully verified | est not suitable for protected data |
| fact Kimi | fact consumer assistant | fact closed; 2023-10 | fact Moonshot AI | unknown active-user count | fact active | unknown current cumulative capital in this sweep | est high jurisdiction/vendor dependence | unknown no comparable rating | fact hosted; no supported campus-local front door | unknown campus-grade controls | unknown default data use not fully verified | est not suitable for protected data |
| fact Z.ai | fact consumer/developer assistant | fact closed service; unknown first date | fact Zhipu AI | unknown active-user count | fact active | unknown current funding/team | est high jurisdiction/vendor dependence | unknown no comparable rating | fact hosted; some GLM weights available separately | unknown campus controls | unknown data default | est not suitable pending review |
| fact You.com | fact search/assistant front door | fact closed; 2021 | fact You.com | unknown current active users | fact active | unknown current capital/team | est medium vendor risk | unknown no comparable rating | fact hosted multi-model | fact enterprise controls advertised; depth not fully tested | unknown plan-specific data defaults | est researchers |
| fact HuggingChat / Omni | fact open-model chat front door | fact service; original 2023-04; Omni 2025-10-16 | fact Hugging Face | unknown active users | fact original service ended 2025-07-01 and returned as Omni | fact Hugging Face-backed | est medium product-continuity risk | unknown no comparable rating | fact multiple open models; separate self-host paths | unknown complete campus controls | unknown current Omni telemetry default | est researchers |
| fact Duck.ai | fact privacy-oriented assistant | fact closed front door; unknown first date | fact DuckDuckGo | unknown active-user count | fact active | fact DuckDuckGo-backed | est medium vendor risk | unknown no comparable rating | fact hosted third-party models; no local | unknown SSO/budget/audit | fact chats are not stored or used for training by default | est everyone for low-risk work |
Coding agents, editors and review systems — 50
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact Claude Code | fact CLI/cloud coding agent | fact proprietary client; 2025 | fact Anthropic | fact 143,875/23,001/≈14,668; npm 20,032,224/week; 39% JetBrains survey use | fact latest 2.1.259, 2026-09-02; fact use up from 18% in 2026-01 | fact Anthropic $65B Series H | est rapid change; strong runway | fact Terminal-Bench scores are agent+model, not harness-only | fact Anthropic/Bedrock/Vertex; no local | fact enterprise identity, policy and audit options | fact route/plan dependent | est builders |
| fact OpenAI Codex | fact CLI/cloud coding agent | fact Apache-2.0 CLI; 2025 | fact OpenAI | fact 121,101/18,554/≈15,011; npm 22,794,236/week; 1.6M weekly users | fact latest 0.153, 2026-09-03; >3× users since 2026-01 | fact OpenAI $122B committed | est rapid breaking-change risk | fact Terminal-Bench entry is agent+model | fact OpenAI-first; compatible/custom and local endpoints in CLI | fact enterprise admin; CLI policy surfaces | fact OpenTelemetry off and prompt logging false by default | est builders |
| fact GitHub Copilot | fact IDE/CLI/cloud agent | fact closed; 2021 preview | fact Microsoft/GitHub | fact 50M users; VS Code 74,501,253 installs; 4.6/5 from 1,054 ratings | fact users up from 26M in 2025-10 | fact Microsoft-backed | est low runway; high vendor dependence | fact marketplace rating; benchmarks depend on model/mode | fact hosted multi-model; no campus-local inference | fact enterprise policy, audit, content exclusions and spend | fact business data protections; telemetry remains service dependent | est builders |
| fact Cursor | fact AI editor/cloud agent | fact closed; 2023 | fact Cursor; joined SpaceX 2026-08-14 | fact 12% JetBrains survey; hundreds of millions weekly requests reported | fact survey down from 18% in 2026-01 | fact $900M Series C, 2025-06-06; >$500M ARR then | est ownership transition and vendor concentration risk | unknown comparable marketplace rating | fact hosted multi-model; no supported local control plane | fact enterprise SSO, privacy and admin features | fact privacy mode available; default depends on plan | est builders |
| fact Windsurf | fact editor/coding agent | fact closed; 2024 | fact Cognition since 2025-07-14 | fact vendor reported hundreds of thousands of DAU and ≈$82M ARR at acquisition | unknown current 90-day use growth | fact acquired by Cognition; price not disclosed | est integration/ownership risk | unknown comparable rating | fact hosted multi-model | fact enterprise identity/admin | unknown verified defaults after acquisition | est builders |
| fact Google Antigravity | fact IDE/agent environment | fact closed preview; 2026-05-19 | fact Google | fact 6% JetBrains survey | fact new Q2 2026 | fact Google-backed | est preview/breaking-change risk | unknown comparable rating | fact Gemini/Google hosted | unknown full production control set | fact usage telemetry enabled in preview | est builders, pilot only |
| fact Gemini CLI | fact coding CLI | fact Apache-2.0; 2025 | fact Google | fact 106,800/14,525/850 | fact R90 96 including nightly; latest 2026-09-02 | fact Google-backed | fact consumer service path ended 2026-06-18; Antigravity transition | unknown harness-only rating | fact Gemini/Vertex; no first-class local | fact Google Cloud/Workspace controls by route | unknown CLI telemetry default in current transition | est builders |
| fact Jules | fact cloud coding agent | fact closed; 2025 | fact Google | unknown active users | fact active | fact Google-backed | est medium vendor/preview risk | unknown public rating | fact Gemini-hosted | fact Google-account controls; unknown full campus audit | unknown data default | est builders |
| fact Devin | fact cloud software agent | fact closed; 2023-12 | fact Cognition | unknown public active users | fact active | fact Cognition raised >$1B at $26B valuation, 2026-05-27 | est strong runway; proprietary lock-in | unknown independent current harness score | fact hosted; no local | fact enterprise controls advertised | unknown default task-data retention | est builders |
| fact Replit Agent | fact browser IDE/application agent | fact closed; 2024 | fact Replit | fact >50M platform users; 85% of Fortune 500 represented | unknown agent-active count | fact $400M round, 2026-03-11, $9B valuation | est strong runway; platform lock-in | unknown independent harness rating | fact Replit-hosted models/tools | fact teams, roles and enterprise controls | unknown precise default telemetry | est builders |
| fact Amazon Q Developer | fact IDE/cloud coding assistant | fact closed; 2023 | fact AWS | unknown active users | fact Q CLI superseded by Kiro CLI | fact Amazon-backed | est product-line migration risk | unknown current independent score | fact AWS-hosted | fact IAM, organization and logging integration | fact service telemetry under AWS terms | est builders |
| fact Junie | fact JetBrains coding agent | fact closed; 2025 | fact JetBrains | fact 9% JetBrains survey | fact active | fact JetBrains-backed | est medium IDE lock-in | unknown separate marketplace score | fact hosted providers; no local standard | fact JetBrains enterprise controls vary by IDE | unknown data default | est builders |
| fact Kiro | fact IDE/CLI coding agent | fact closed; 2025 | fact AWS | unknown active users | fact current successor path for Amazon Q CLI | fact Amazon-backed | est new/product-transition risk | unknown rating | fact AWS-hosted | fact AWS identity/policy integration | unknown detailed default | est builders |
| fact Amp | fact coding agent | fact closed; 2025 | fact Sourcegraph | unknown active users | fact active | fact Sourcegraph-backed | est medium vendor risk | unknown rating | fact hosted models | unknown complete control matrix | unknown default | est builders |
| fact Factory Droids | fact cloud coding agents | fact closed; unknown first date | fact Factory | unknown active users | fact active | fact $50M Series B, 2025-09 | est medium vendor risk | unknown independent rating | fact hosted | fact enterprise controls advertised | unknown default | est builders |
| fact Augment Code | fact IDE coding agent | fact closed; 2024 | fact Augment | fact VS Code 776,022 installs; 3.5/5 from 340 ratings | fact active | fact $227M Series B, 2024-04; $252M total then | est strong funding; proprietary lock-in | fact marketplace rating | fact hosted | fact enterprise security/admin | unknown default telemetry | est builders |
| fact Warp | fact agentic terminal | fact closed client/service; 2020 | fact Warp | unknown active users | fact active | unknown current capital in this sweep | est medium vendor risk | unknown comparable rating | fact hosted agent; shell remains local | fact team/admin options | fact terminal telemetry configurable; exact current default not rechecked | est builders |
| fact TRAE | fact AI editor | fact closed; 2025 | fact ByteDance | unknown active users | fact active | fact ByteDance-backed | est jurisdiction/vendor risk | unknown rating | fact hosted models | unknown university controls | unknown default data use | est not suitable pending review |
| fact IBM Bob | fact enterprise coding agent | fact closed; GA 2026-04-28 | fact IBM | fact used by about 80,000 IBM employees | fact new Q2 2026 | fact IBM-backed | est low runway; new-product risk | unknown independent rating | fact hosted/enterprise IBM stack | fact enterprise identity, governance and audit | unknown exact telemetry default | est builders |
| fact Meta Muse Code | fact coding agent | fact closed beta; 2026-08-05 | fact Meta | unknown beta users | fact new Q3 2026 | fact Meta-backed | est high preview risk | unknown rating | fact Meta-hosted | unknown production controls | unknown telemetry | est builders, watch |
| fact Grok Build | fact coding agent/client | fact source published 2026-07-15; license unknown here | fact xAI | fact ≈26.4K stars; issues disabled | fact periodic upstream sync; new Q3 2026 | fact xAI-backed | est high governance and issue-transparency risk | unknown rating | fact Grok-hosted | unknown campus controls | fact egress reporting described; unknown full default | est not suitable yet |
| fact OpenCode | fact CLI/TUI coding agent | fact MIT; unknown first date | fact Anomaly | fact 203,446/26,541/5,643; npm 1,848,183/week; 7% survey | fact R90 55; est ≈12K stars/month over 118 days | unknown funding/team | est rapid-change and maintainer-concentration risk | unknown independent harness-only rating | fact provider-agnostic; local models | unknown native enterprise control plane | unknown current default telemetry | est builders |
| fact Aider | fact terminal coding agent | fact Apache-2.0; 2023 | fact independent | fact 48,698/4,920/1,849 | fact no GitHub release since 2025-08-09; development/tags continue | unknown funding/team | est medium maintainer/bus-factor risk | fact public model benchmark suite; mostly model+prompt results | fact provider-agnostic; local models | unknown SSO/audit/budget | fact consent telemetry | est builders |
| fact Cline | fact IDE coding agent | fact Apache-2.0; 2024 | fact Cline | fact 67,402/7,280/1,194; VS Code 5,198,510; 4.0/5, 312 ratings | fact R90 ≥100; est ≈1.45K stars/month | fact $32M seed+Series A, 2025-07-31 | fact npm compromise 2026-02-17; patched; later advisories patched | fact marketplace rating; vendor benchmark not independent | fact provider-agnostic; local models | fact Enterprise SSO, RBAC, audit and cost controls | fact code remains local in extension; optional telemetry, default unknown | est builders |
| fact Roo Code | fact IDE coding agent | fact Apache-2.0; 2024 | fact Roo Code Inc. | fact 24,313/3,415/1,034; 1,978,032 installs; 4.5/5, 348 | fact archived 2026-05-15 | unknown funding/team | fact shutdown/archived | fact marketplace rating | fact provider-agnostic/local historically | unknown continuing controls | unknown post-shutdown data state | est not suitable |
| fact Kilo Code | fact IDE coding agent | fact MIT; unknown first date | fact Kilo; acquired by Anaconda 2026-07-15 | fact 27,156/3,114/552; 1,494,211 installs; 4.3/5, 205 | fact active | fact acquisition; price unknown | est medium integration risk | fact marketplace rating | fact multi-provider/local | unknown complete campus control matrix | unknown default | est builders |
| fact Continue | fact IDE coding assistant | fact Apache-2.0; 2023 | fact Continue | fact 35,739/5,322/941; 4,061,691 installs; 3.5/5, 181 | fact repository says no longer actively maintained | unknown current funding/team | fact maintenance retreat/read-only direction | fact marketplace rating | fact provider-agnostic/local historically | unknown continuing enterprise support | unknown default | est not suitable |
| fact OpenHands | fact autonomous coding platform | fact MIT core; 2024 | fact All Hands AI | fact 86,069/11,286/640; PyPI 623,690/month | fact R90 33; latest 1.16, 2026-08-27 | fact $18.8M Series A, 2025-11-18 | est active but fast-moving | fact benchmark results are agent+model | fact multi-provider/local | fact self-host and enterprise controls | unknown telemetry default | est builders |
| fact Goose | fact local coding/general agent | fact Apache-2.0; 2024 | fact Block; LF AI & Data/AAIF | fact 53,879/6,161/267 | fact active | fact corporate/foundation backed | est favorable governance; active-change risk | unknown independent rating | fact provider-agnostic/local | unknown native SSO/budget; self-host policy possible | fact telemetry off; sandbox and prompt-injection defenses off by default | est builders |
| fact Qwen Code | fact coding CLI | fact Apache-2.0; 2025 | fact Alibaba/Qwen | fact 27,611/2,983/1,264 | est ≈804 stars/month over sampled interval | fact Alibaba-backed | est medium governance/jurisdiction risk | unknown rating | fact Qwen/OpenAI-compatible endpoints; local possible | unknown enterprise controls | unknown telemetry | est builders |
| fact Mistral Vibe CLI | fact coding CLI | fact Apache-2.0; unknown first date | fact Mistral AI | fact 4,913/684/285 | fact active | fact Mistral-backed | est medium provider dependence | unknown rating | fact Mistral-first | unknown native controls | unknown telemetry | est builders |
| fact Crush | fact terminal coding agent | fact FSL-1.1-MIT; 2025 | fact Charmbracelet | fact 27,877/2,218/687 | est ≈952 stars/month over 97 days | unknown capital/team | est source-available license risk | unknown rating | fact multi-provider/local | unknown enterprise controls | fact telemetry on by default | est builders |
| fact SWE-agent | fact research coding agent | fact MIT; 2024 | fact Princeton/NLP research community | unknown current count not frozen in final set | fact directs new users toward mini-SWE-agent | fact academic/community backed | est migration risk | fact SWE-bench results are agent+model/configuration | fact multi-provider/local possible | unknown enterprise controls | unknown telemetry | est researchers |
| fact mini-SWE-agent | fact compact research coding agent | fact MIT; 2025 | fact SWE-agent team | unknown exact snapshot count | fact active successor path | fact academic/community backed | est medium research-project risk | fact benchmark results remain agent+model | fact multi-provider/local possible | unknown enterprise controls | unknown telemetry | est researchers |
| fact CodeRabbit | fact code-review agent | fact closed service; unknown first date | fact CodeRabbit | fact >2M reviews/week; 17K customers | fact active | fact $143M Series C, 2026-08-12; unknown lead in this sweep | est strong funding; hosted lock-in | unknown independent review-quality rating | fact hosted | fact enterprise SSO/policy/audit advertised | unknown code-retention default by plan | est builders |
| fact Qodo | fact coding/review agent | fact closed plus OSS PR-Agent; unknown first date | fact Qodo | unknown active users | fact active | fact $70M Series B, 2026-03-30; $120M total | est strong funding; vendor risk | unknown independent rating | fact hosted multi-model | fact enterprise controls | unknown default | est builders |
| fact PR-Agent | fact pull-request agent | fact AGPL-3.0; 2023 | fact Qodo | unknown exact final snapshot | fact active | fact parent raised $120M total | est copyleft and parent-product coupling risk | unknown rating | fact multi-provider | unknown native complete controls | unknown telemetry | est builders |
| fact ECC | fact coding-agent toolkit | fact MIT; 2026-01-18 | unknown owner provenance not fully verified | fact 246,773/37,187/136 | fact R90 3; all growth occurred in <8 months | unknown capital/team | est very high provenance/star-quality risk | unknown rating | unknown provider/local matrix | unknown controls | unknown telemetry | est not suitable yet |
| fact DeepSeek Harness | fact coding harness | fact MIT; 2026-08-13 | unknown verified organizational relationship | fact 210,708/24,655/0 | fact R90 10; est 291,543 stars/month annualized from 22 days; anomaly | unknown capital/team | est extreme provenance/manipulation risk until validated | unknown rating | unknown | unknown | unknown | est not suitable |
| fact Ponytail | fact coding/agent harness | fact MIT; 2026-06-12 | unknown owner/backer | fact 122,907/6,643/202 | est 44,539 stars/month from 84-day lifetime | unknown | est extreme new-project risk | unknown | unknown | unknown | unknown | est not suitable yet |
| fact Caveman | fact coding harness | unknown license; 2026-04-04 | unknown owner/backer | fact 102,927/5,986/121 | fact pushed 2026-09-02 | unknown | est high license/provenance risk | unknown | unknown | unknown | unknown | est not suitable yet |
| fact Graphify | fact coding/context harness | fact Apache-2.0; 2026-04-03 | unknown owner/backer | fact 114,218/11,101/1,243 | fact pushed 2026-08-30 | unknown | est high new-project/provenance risk | unknown | unknown | unknown | unknown | est not suitable yet |
| fact ruflo | fact agent/coding orchestration | fact MIT; unknown first date | fact community project | fact 70,316/8,380/900 | fact R90 ≥100 | unknown | est rapid-change and maintainer risk | unknown | fact multi-provider/local options | unknown | unknown | est builders, watch |
| fact oh-my-openagent | fact coding-agent configuration | unknown license; 2025-12 | unknown owner/backer | fact 68,648/5,639/910 | fact R90 68, beta | unknown | est high license/beta risk | unknown | unknown | unknown | unknown | est not suitable yet |
| fact Prime Agent | fact coding agent | fact MIT; 2026-05-08 | unknown owner/backer | fact 19,733/2,154/81 | fact active new project | unknown | est high new-project risk | unknown | unknown | unknown | unknown | est builders, watch |
| fact OpenHarness | fact coding harness | fact MIT; 2026-04-01 | unknown owner/backer | fact 15,633/2,545/86 | fact one release 2026-05-07; last push 2026-06-04 | unknown | est high staleness risk despite age | unknown | unknown | unknown | unknown | est not suitable |
| fact Kimi Code | fact coding agent | fact MIT; 2026-05-22 | fact Moonshot/Kimi | fact 7,238/1,161/1,298 | fact active new project | fact parent-backed; exact project capital unknown | est high new-project/jurisdiction risk | unknown | fact Kimi-first | unknown | unknown | est builders, watch |
| fact MiMo Code | fact coding agent | fact MIT; 2026-06-10 | unknown owner/backer | fact 12,939/1,334/975 | est ≈4,580 stars/month over lifetime | unknown | est high new-project risk | unknown | unknown | unknown | unknown | est builders, watch |
| fact fx | fact coding harness | fact Apache-2.0; 2026-08-11 | unknown owner/backer | fact 2,710/312/184 | fact new Q3 2026 | unknown | est high early-stage risk | unknown | unknown | unknown | unknown | est builders, watch |
| fact qm | fact coding harness | fact MIT; 2026-07-29 | unknown owner/backer | fact 14,523/1,764/382 | est +6,712 stars in sampled seven-day interval | unknown | est very high spike/provenance risk | unknown | unknown | unknown | unknown | est not suitable yet |
General, browser and computer-use agents — 18
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact ChatGPT agent | fact browser/computer/research agent | fact closed; 2025-07-17 | fact OpenAI | unknown agent-specific active users | fact absorbed Operator/deep-research paths | fact OpenAI-backed | est strong runway; high action/vendor risk | unknown independent operational score | fact OpenAI-hosted | fact Enterprise/Edu policy and admin controls | fact plan-dependent retention/training | est everyone, limited pilot |
| fact Manus | fact cloud general agent | fact closed; 2025-03 | fact acquired by Meta 2025-12-29 | fact vendor reports millions of users, 147T tokens and 80M virtual machines | fact active | fact Meta-backed; acquisition price unknown | est ownership and hosted-action risk | unknown independent rating | fact hosted | unknown complete university controls | unknown task-data default | est not suitable pending review |
| fact Perplexity Comet | fact agentic browser | fact closed; 2025 | fact Perplexity | unknown active users | fact active | unknown capital in this sweep | est browser-data/vendor risk | unknown rating | fact hosted | fact enterprise MDM and policy controls advertised | unknown detailed browsing-data default | est researchers, limited pilot |
| fact Gemini Agent / Project Mariner | fact browser/computer agent | fact closed research/preview; 2024 | fact Google DeepMind | unknown active users | fact capabilities moving into Gemini Agent | fact Google-backed | est preview/product-transition risk | unknown independent rating | fact Google-hosted | fact managed access is off by default in relevant enterprise paths | unknown detailed data default | est everyone, watch |
| fact Browser Use | fact browser-agent framework/front end | fact MIT; 2024 | fact Browser Use | fact 112,156/12,334/402 | fact R90 9; latest 0.13.8, 2026-08-16; est ≈4.86K stars/month | fact $17M seed, 2025-03-22 | est active; telemetry and browser-action risk | unknown independent score | fact multi-provider; local possible | unknown native campus control plane | fact telemetry true by default; reported collection can include task/action context | est not suitable without isolation |
| fact Browser Harness | fact browser-agent harness | unknown license here; 2026-04-17 | unknown owner/backer | fact ≈17.3K/1.7K | fact latest 0.1.10, 2026-08-26; est ≈3.73K stars/month | unknown | est high early-stage risk | unknown | unknown | unknown | fact telemetry opt-out | est builders, watch |
| fact Stagehand | fact browser automation SDK | fact MIT; 2024 | fact Browserbase | fact ≈24.1K/1.7K; npm 1,403,303/week | fact V4, 2026-08-10 | fact Browserbase $40M Series B, 2025-06 | est medium platform dependence | unknown independent rating | fact multi-model; Browserbase or local browser | fact Browserbase enterprise controls by route | unknown SDK/platform telemetry default | est builders |
| fact Browserbase Agents | fact managed browser-agent platform | fact closed; GA 2026-06-30 | fact Browserbase | fact platform reports 35M browser sessions/month; not agent users | fact new Q2 2026 | fact $40M Series B, 2025-06 | est medium hosted-browser dependence | unknown rating | fact hosted browser, multi-model | fact enterprise session/admin controls | unknown exact default retention | est builders |
| fact Skyvern | fact browser automation agent | fact AGPL-3.0; 2023 | fact Skyvern | fact ≈22.9K stars | fact active | unknown current capital/team | est copyleft and browser-action risk | unknown | fact multi-model/self-host | unknown full native campus controls | fact telemetry on by default | est builders |
| fact Playwright MCP | fact browser MCP server | fact Apache-2.0; 2025 | fact Microsoft | fact ≈36.8K stars | fact active | fact Microsoft-backed | est low project-runway risk; browser-action risk remains | unknown | fact model-agnostic; local browser | unknown native SSO/budget; host supplies policy | est no model telemetry intrinsic; browser data reaches chosen host/model | est builders |
| fact Open Interpreter | fact local computer/code agent | fact AGPL-3.0; 2023 | fact Open Interpreter | fact ≈68.2K stars | fact product direction rewritten around Rust/Codex; old Python community persists | unknown | est high architecture/migration risk | unknown | fact multi-provider/local historically | unknown | unknown | est builders, watch |
| fact UI-TARS / Agent TARS | fact computer-use agent | fact Apache-2.0; 2025 | fact ByteDance | fact ≈38.8K stars | fact active | fact ByteDance-backed | est jurisdiction and action-safety risk | unknown benchmark results depend on model/environment | fact TARS models/local options | unknown campus controls | unknown telemetry | est researchers |
| fact Agent S | fact computer-use research agent | fact Apache-2.0; 2024 | fact Simular/academic collaborators | fact ≈12.2K stars | fact active | unknown capital/team | est research-system risk | fact OSWorld results are model+agent+environment | fact multi-model | unknown enterprise controls | unknown | est researchers |
| fact OpenClaw | fact general autonomous agent | unknown license; 2025-11-24 | unknown owner/backer | fact 388,724/81,627/6,089 | fact active | unknown | est extreme provenance, governance and open-issue risk | unknown | unknown | unknown | unknown | est not suitable |
| fact AutoGPT | fact general-agent platform | fact MIT; 2023 | fact Significant Gravitas | fact 187,099/46,040/544 | fact R90 13 | unknown | est medium product-direction risk | unknown current independent rating | fact multi-provider/local possible | unknown complete controls | unknown telemetry | est researchers |
| fact nanobot | fact compact agent framework | fact MIT; 2026-02-01 | fact HKU Data Science group | fact 47,685/8,416/758 | fact active new project | fact academic backing | est high early-stage/bus-factor risk | unknown | fact multi-provider/local possible | unknown | unknown | est researchers |
| fact Cloudflare Computer | fact computer-use runtime | fact MIT; 2026-06-05 | fact Cloudflare | fact 8,954/501/21 | fact +8,868 stars in monthly trending window | fact Cloudflare-backed | est preview and platform-dependence risk | unknown | fact Cloudflare runtime; model choice varies | fact Cloudflare account controls | unknown telemetry/retention | est builders, watch |
| fact Cloudflare OS | fact agent operating environment | fact Apache-2.0; 2026-04-15 | fact Cloudflare | fact 9,539/1,123/117 | fact +9,549 monthly-trending signal | fact Cloudflare-backed | est high preview/breaking-change risk | unknown | fact Cloudflare platform | fact account controls | unknown telemetry | est builders, watch |
Agent frameworks, SDKs and orchestration — 23
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact LangChain | fact agent/LLM framework | fact MIT; 2022 | fact LangChain Inc. | fact 145,576/24,297/443; PyPI 229,927,751/month | fact R90 86; latest alpha 2026-09-02 | fact $125M Series B, 2025-10-20, $1.25B valuation | est low runway; high change/vendor-direction risk | fact benchmark claims configuration-dependent | fact provider-agnostic/local | fact LangSmith adds SSO/audit/spend; core does not | fact core tracing off; CLI analytics on | est builders |
| fact LangGraph | fact stateful agent framework | fact MIT; 2023 | fact LangChain Inc. | fact 40,992/6,916/743; npm 13,152,611 in Aug | fact active | fact parent $125M Series B | est medium vendor/change risk | fact deep-agent results are agent+model/configuration | fact provider-agnostic/local | fact managed platform adds enterprise controls | unknown core telemetry default | est builders |
| fact AutoGen | fact multi-agent framework | fact MIT; 2023 | fact Microsoft | fact 60,791/9,178/1,031 | fact maintenance mode | fact Microsoft-backed | fact successor is Microsoft Agent Framework | unknown current rating | fact provider-agnostic/local | unknown core native controls | unknown telemetry | est not suitable for new builds |
| fact Semantic Kernel | fact agent/application SDK | fact MIT; 2023 | fact Microsoft | fact 28,527/4,752/271 | fact new work directed toward Microsoft Agent Framework | fact Microsoft-backed | est migration risk | unknown | fact multi-provider/local connectors | fact Azure supplies enterprise controls | unknown SDK telemetry | est not suitable for new agent builds |
| fact Microsoft Agent Framework | fact successor agent SDK | fact MIT; 2026 | fact Microsoft | fact 13,306/2,252/649 | fact 1.0 on 2026-04-03; R90 24; Python 1.17 on 2026-09-03 | fact Microsoft-backed | est medium early-version risk | unknown independent score | fact multi-provider; Azure/local connectors | fact Azure identity, audit and policy integration | unknown core default telemetry | est builders |
| fact CrewAI | fact multi-agent framework/platform | fact MIT; 2023 | fact CrewAI | fact 58,045/8,326/710; PyPI 26,457,068/month; vendor says 10M agents/month | fact R90 37 GitHub releases | fact >$18M raised by 2024-10 | est active; medium vendor/change risk | unknown independent rating | fact provider-agnostic/local | fact enterprise platform controls; core limited | fact telemetry on by default | est builders |
| fact LlamaIndex | fact data/agent framework | fact MIT; 2022 | fact LlamaIndex | fact 51,998/8,076/696; PyPI 6,405,226/month | fact active | fact $19M Series A, 2025-03; $27.5M total | est medium vendor/change risk | unknown comparable rating | fact provider-agnostic/local | fact managed platform controls; core limited | unknown core telemetry default | est researchers |
| fact Haystack | fact RAG/agent framework | fact Apache-2.0; 2019 | fact deepset | fact 26,403/3,071/103; PyPI 760,882/month | fact active | unknown current parent capital/team | est comparatively mature | fact ADK Arena preprint reports strong build-task performance; not production proof | fact provider-agnostic/local | fact deepset platform provides enterprise controls | unknown core telemetry | est researchers |
| fact PydanticAI | fact typed agent framework | fact MIT; 2024 | fact Pydantic | fact 19,699/2,634/780; PyPI 11,040,846/month | fact R90 56 | fact parent $12.5M Series A, 2024-10 | est active; rapid-change risk | unknown | fact provider-agnostic/local | unknown core enterprise controls | unknown telemetry | est builders |
| fact OpenAI Agents SDK | fact agent SDK | fact MIT; 2025 | fact OpenAI | fact 29,169/4,658/78; PyPI 33,424,617/month | fact R90 17 | fact OpenAI-backed | est low runway; vendor-direction risk | unknown harness-only score | fact OpenAI-first; custom model adapters | fact OpenAI org controls by route | fact tracing on by default and can include prompts/results | est builders |
| fact Google ADK | fact agent SDK | fact Apache-2.0; 2025 | fact Google | fact 21,390/3,937/517; PyPI 19,429,084/month | fact R90 20 | fact Google-backed | est medium vendor/change risk | fact ADK Arena results are framework+configuration, preprint | fact multi-model; Google-optimized; local possible | fact Google Cloud controls by deployment | unknown SDK telemetry | est builders |
| fact Strands Agents | fact agent SDK | fact Apache-2.0; 2025 | fact AWS; foundation participation | fact 7,141/1,089/699; vendor says 14M downloads | fact active | fact AWS-backed | est medium ecosystem dependence | unknown independent score | fact model-agnostic; AWS-optimized | fact IAM/CloudTrail by deployment | unknown SDK telemetry | est builders |
| fact Mastra | fact TypeScript agent framework | fact Apache-2.0 core plus commercial EE; 2024 | fact Mastra | fact 27,668/2,725/541; npm ≈1.58M/week | fact R90 22 | fact $22M Series A, 2026-04-09; $35M total; >35 staff | est strong growth; license/product-boundary risk | unknown | fact provider-agnostic/local | fact enterprise tier controls | unknown telemetry | est builders |
| fact Agno | fact agent framework/runtime | fact Apache-2.0; 2023 | fact Agno | fact 42,031/5,865/1,312; est ≈2.18M monthly downloads from secondary counter | fact 31 stable/47 total releases in 90 days | unknown capital/team | est high release/breaking-change risk | unknown | fact provider-agnostic/local | unknown complete enterprise controls | fact telemetry on | est builders |
| fact smolagents | fact compact agent framework | fact Apache-2.0; 2024 | fact Hugging Face | fact 29,140/2,918/756; PyPI 580,640/month | fact no GitHub releases in R90; source active | fact Hugging Face-backed | est medium cadence risk | unknown | fact provider-agnostic/local | unknown native controls | unknown telemetry | est researchers |
| fact DSPy | fact program/agent optimization framework | fact MIT; 2023 | fact Stanford research/community | unknown exact final snapshot omitted | fact active | fact academic/community-backed | est medium research-governance risk | fact evaluations are task/model/program dependent | fact provider-agnostic/local | unknown native controls | unknown telemetry | est researchers |
| fact BeeAI | fact agent framework | fact Apache-2.0; 2024 | fact LF AI & Data; originated at IBM | fact ≈3,390 stars | fact IBM no longer sole maintainer; foundation continues | fact foundation/corporate backing | est medium governance-transition risk | unknown | fact provider-agnostic/local | unknown | unknown | est researchers |
| fact MetaGPT | fact multi-agent framework | fact MIT; 2023 | fact FoundationAgents/community | fact 70,197 stars | fact last push 2026-01-21; latest release 2025-03 | unknown capital/team | est high staleness risk | unknown | fact multi-provider | unknown | unknown | est not suitable for new builds |
| fact Letta | fact stateful/memory agent platform | fact Apache-2.0 core; 2023 | fact Letta | fact 24,601 stars | fact latest release 2026-05-14 | unknown current capital/team | est medium cadence/vendor risk | unknown | fact multi-provider/local | fact hosted controls vary by tier | unknown telemetry | est researchers |
| fact Vercel AI SDK | fact application/agent SDK | unknown license assertion not frozen; 2023 | fact Vercel | fact 26,564/5,069/1,545; npm 23,560,904/week | fact R90 ≥100 | fact Vercel-backed | est high change rate; ecosystem dependence | unknown harness rating | fact provider-agnostic; local endpoints possible | fact Vercel platform controls; core limited | unknown core telemetry | est builders |
| fact Antigravity SDK | fact agent SDK | unknown preview license; 2026-04-29 | fact Google | fact 3,275/1,295/32 | fact no formal releases; preview | fact Google-backed | est high preview/API-change risk | unknown | fact Google-optimized | unknown production controls | unknown telemetry | est builders, watch |
| fact Paperclip | fact agent orchestration/workspace | fact MIT; 2026-03-02 | unknown owner/backer verified only at repository level | fact 79,934/14,675/5,370 | fact R90 11; est ≈4,962 stars/month from sampled baseline | unknown capital/team | est very high newness, issue-load and bus-factor risk | unknown | fact multi-provider claims; local unknown | unknown | unknown | est builders, watch |
| fact Amazon Bedrock AgentCore | fact managed agent runtime/harness | fact closed service; preview 2026-04; GA 2026-06-18 | fact AWS | unknown active customers | fact new Q2 2026 | fact Amazon-backed | est low runway; high AWS dependence | unknown | fact model-flexible inside AWS | fact IAM, CloudTrail, budgets and policy | fact CloudWatch tracing on in configured paths; memory default 30 days | est builders |
RAG, research and knowledge systems — 16
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact Gemini Notebook | fact source-grounded research notebook | fact closed; NotebookLM 2023; renamed 2026-07-16 | fact Google | fact >30M people and 600K organizations | fact active/rebranded | fact Google-backed | est strong longevity; Google lock-in | unknown comparable research rating | fact Gemini-hosted; no local | fact Workspace identity/admin by edition | fact enterprise protections differ from consumer | est researchers |
| fact Elicit | fact literature-research assistant | fact closed; 2017 | fact Ought/Elicit | fact >400K monthly researchers | fact Research Agent launched 2026-08-04 | fact $22M Series A, 2025-02 | est medium vendor risk | unknown independent comprehensive rating | fact hosted models | fact organization controls; complete campus matrix unknown | unknown exact default by tier | est researchers |
| fact Consensus | fact scholarly search/answer system | fact closed; 2022 | fact Consensus | fact 2.5M MAU | fact active | fact $30M funding announced 2026-05-11 | est medium vendor risk | unknown independent rating | fact hosted | unknown complete SSO/audit set | unknown default | est researchers |
| fact Glean | fact enterprise knowledge/search agent | fact closed; 2019 | fact Glean | fact $300M ARR, 2026-05-28; nearly 2× Fortune 500 reach; 45% wDAU/wMAU | fact ARR tripled from $100M in 15 months | fact $150M Series F, 2025-06, $7.2B valuation | est strong runway; high vendor dependence | unknown comparable rating | fact hosted multi-model | fact enterprise identity, ACL preservation, audit and governance | unknown exact telemetry/retention by contract | est everyone where budget permits |
| fact Hebbia | fact enterprise document research | fact closed; unknown first date | fact Hebbia | unknown active seats | fact active | fact $130M Series B, 2024-07 | est strong capital; proprietary lock-in | unknown independent rating | fact hosted | fact enterprise controls advertised | unknown default | est researchers |
| fact SciSpace | fact scholarly reading/search assistant | fact closed; unknown first date | fact SciSpace | fact vendor reports 9.6M researchers; independent confirmation unknown | unknown 90-day growth | unknown current capital/team | est medium vendor/evidence risk | unknown independent rating | fact hosted | unknown complete campus controls | unknown data default | est researchers |
| fact Onyx | fact enterprise search/RAG assistant | fact MIT core plus proprietary EE; 2023 | fact Onyx | fact 31,910/4,399/427; vendor references 14K Netflix and 37K UCSD users | fact active | fact $10M seed, 2025; unknown team | est medium commercial-boundary risk | unknown comparable rating | fact provider-agnostic; local/air-gapped | fact SSO, RBAC, audit and enterprise controls | unknown telemetry default; self-host permits stronger containment | est everyone |
| fact Khoj | fact personal/team knowledge assistant | fact AGPL-3.0; 2023 | fact Khoj | fact 37,033/2,451 | fact active | unknown capital/team | est copyleft and small-team risk | unknown | fact multi-provider/local | unknown complete campus controls | unknown telemetry | est researchers |
| fact RAGFlow | fact document RAG platform | fact Apache-2.0; 2023 | fact InfiniFlow | fact 89,983/10,610/1,581 | est ≈2.52K stars/month in sampled 11-day interval | unknown capital/team | est medium issue-load/vendor risk | unknown | fact multi-provider/local | fact self-host; enterprise controls vary | unknown telemetry | est researchers |
| fact Dify | fact RAG/agent application builder | fact modified Apache; 2023-04-12 | fact Dify | fact 154,326/24,397/1,021; vendor claims 1.4M machines | fact R90 5; latest 1.17.0, 2026-08-25; est ≈16.8K stars/month over sampled 97 days | fact $30M pre-Series A, 2026-03-10 | fact license restricts some multi-tenant/rebranding uses | unknown independent rating | fact provider-agnostic/local | fact enterprise SSO/audit/budget features | fact telemetry present/on in default feature configuration | est builders |
| fact Flowise | fact visual RAG/agent builder | fact Apache-2.0; 2023 | fact Flowise | fact 55,404/24,971/1,046 | fact R90 2; frozen 2026-07-29; archived 2026-08-13; EOL 2026-08-31 | unknown | fact ended | unknown | fact historically multi-provider/local | unknown continuing controls | unknown | est not suitable |
| fact Langflow | fact visual RAG/agent builder | fact MIT; 2023 | fact Langflow/Astra ecosystem | fact 154,184/10,012/1,021 | fact R90 12; latest 1.12, 2026-09-01 | unknown current capital/team | fact security advisories in 2026-08; patches released | unknown | fact provider-agnostic/local | fact enterprise deployment options | unknown telemetry default | est builders |
| fact Quivr | fact personal/team RAG | fact Apache-2.0; 2023 | fact Quivr | fact 39,379 stars | fact latest release 2025-02-04 | unknown | est high staleness risk | unknown | fact multi-provider/local historically | unknown | unknown | est not suitable for new standard |
| fact DocsGPT | fact document chat/RAG | fact MIT; 2023 | fact Arc53/community | fact 18,236 stars | fact active | unknown | est medium small-project risk | unknown | fact multi-provider/local | unknown enterprise control depth | unknown telemetry | est researchers |
| fact Kotaemon | fact document RAG UI | fact Apache-2.0; 2024 | fact Cinnamon/community | fact 25,707 stars | fact active | unknown | est medium maintainer risk | unknown | fact multi-provider/local | unknown | unknown | est researchers |
| fact PrivateGPT | fact local document RAG | fact Apache-2.0; 2023 | fact Zylon/community | fact 57,487 stars | unknown current release cadence not frozen | unknown | est medium cadence/maintainer risk | unknown | fact local-first; multi-provider | unknown enterprise controls | est local operation can avoid external telemetry | est researchers |
Workflow and automation builders — 15
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact n8n | fact workflow/agent builder | fact Sustainable Use License; 2019 | fact n8n | fact 203,223/60,535/1,137; npm 94,686/week | fact R90 ≥100 | fact $180M Series C, 2025-10-09; $240M total; 6× users/10× revenue reported | fact 2026 security advisories patched; license limits commercial hosting | unknown comparable rating | fact many providers/local connectors | fact SSO, RBAC, audit and external secrets in enterprise | fact telemetry on unless disabled | est builders |
| fact Activepieces | fact workflow/agent builder | fact MIT core plus EE; 2022 | fact Activepieces | fact 24,210/4,136/519; npm 2,305,328 in Aug | fact R90 29 | unknown current capital/team | est medium commercial-boundary risk | unknown | fact multi-provider/self-host | fact enterprise SSO/RBAC/audit | unknown telemetry default | est builders |
| fact Pipedream | fact developer automation platform | fact closed platform; integration repo OSS; 2019 | fact Workday acquisition completed by 2026-01-31 | fact integrations repo 11,668 stars | fact active | fact acquired; price unknown | est integration/ownership risk | unknown | fact hosted multi-provider | fact enterprise identity/audit | unknown default telemetry | est builders |
| fact Gumloop | fact no-code agent workflow builder | fact closed; unknown first date | fact Gumloop | unknown active users | fact active | fact $50M Series B, 2026-03-12; at least $70M total | est strong funding; proprietary lock-in | unknown | fact hosted multi-model | fact enterprise controls advertised | unknown default | est builders |
| fact Relevance AI | fact agent/workforce builder | fact closed; unknown first date | fact Relevance AI | unknown active users | fact active | fact $24M Series B, 2025-05 | est medium vendor risk | unknown | fact hosted multi-model | fact enterprise controls advertised | unknown | est builders |
| fact Lindy | fact no-code assistant/agent builder | fact closed; unknown first date | fact Lindy | unknown active users | fact active | fact ≈23 staff publicly indicated; funding unknown here | est small-team/vendor risk | unknown | fact hosted | fact SOC 2 and Safe Mode advertised; full audit matrix unknown | unknown default | est builders |
| fact StackAI | fact enterprise workflow/agent builder | fact closed; unknown first date | fact acquired by Asana 2026-05-28 | unknown active users | fact active/acquisition integration | fact ≈$75M upfront reported; ≈62 staff | est ownership/integration risk | unknown | fact hosted multi-model | fact enterprise controls advertised | unknown | est builders |
| fact OpenAI Agent Builder | fact hosted visual agent builder | fact closed; 2025 | fact OpenAI | unknown active builders | fact service unavailable after 2026-11-30 | fact OpenAI-backed | fact sunset announced | unknown | fact OpenAI-hosted | fact OpenAI organization controls | fact tracing/data follows platform settings | est not suitable |
| fact Zapier Agents | fact SaaS agent builder | fact closed; 2024 | fact Zapier | fact 50K teams reported before rename; >9K app integrations | fact active | fact Zapier-backed | est medium vendor dependence | unknown | fact hosted multi-model | fact enterprise identity, app policy and audit | unknown detailed default | est builders |
| fact Make AI Agents | fact SaaS workflow/agent builder | fact closed; unknown first date | fact Make/Celonis | unknown active users | fact active | fact Celonis-backed | est medium platform dependence | unknown | fact hosted | fact enterprise controls by plan | unknown | est builders |
| fact Bardeen | fact browser/workflow automation | fact closed; 2021 | fact Bardeen | unknown active users | fact active | unknown current capital/team | est browser-data/vendor risk | unknown | fact hosted | fact team controls advertised | unknown | est builders |
| fact Relay.app | fact human-in-loop workflow builder | fact closed; unknown first date | fact Relay | unknown active users | fact active | unknown capital/team | est small-vendor risk | unknown | fact hosted multi-model | unknown complete campus controls | unknown | est builders |
| fact Dust | fact enterprise assistant/agent builder | fact closed plus OSS components; 2023 | fact Dust | unknown active users | fact active | unknown current capital/team | est medium vendor risk | unknown | fact multi-model | fact SSO, permissions and enterprise administration | unknown | est everyone, controlled pilot |
| fact Coze Studio | fact agent workflow builder | fact Apache-2.0; 2025 | fact ByteDance/Coze | unknown exact final snapshot | fact active | fact ByteDance-backed | est jurisdiction and platform risk | unknown | fact multi-model/self-host claims | unknown complete controls | unknown telemetry | est builders |
| fact Workato Agentic | fact enterprise automation/agent platform | fact closed | fact Workato | unknown agent-specific active users | fact active | unknown current capital in this sweep | est strong enterprise position; proprietary lock-in | unknown | fact hosted | fact mature identity, governance and audit controls | unknown | est builders |
Local runners and desktop front doors — 18
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact Ollama | fact local model runner/API | fact MIT; 2023 | fact Ollama | fact 180,043/17,668/3,889; vendor says 8.9M developers | fact R90 29; latest RC 2026-09-02 | fact $88M raised 2026-07-09 | est strong funding; company dependence | unknown harness rating | fact local-first, including Apple silicon | unknown native SSO/audit; external gateway needed | fact telemetry disabled by default; cloud path offers stated zero-data-retention | est everyone |
| fact LM Studio | fact local desktop runner/chat | fact closed/free; 2023 | fact LM Studio | fact vendor reports millions of downloads | fact 0.4.17 on 2026-06-27; Bionic update 2026-07-16 | unknown capital/team | est medium closed-client risk | unknown public comparable rating | fact local-first, Apple silicon | unknown centralized campus controls | fact local privacy claims; exact analytics default unknown | est everyone |
| fact GPT4All | fact local desktop/chat runner | fact MIT; 2023 | fact Nomic AI | fact ≈77.4K/8.3K | fact latest release 2025-02-25 | unknown current capital/team | est high staleness risk | unknown | fact local-first | unknown enterprise controls | unknown telemetry | est not suitable as new standard |
| fact Jan | fact local desktop assistant | fact Apache-2.0; 2023 | fact Menlo Research | fact 44,313/3,007/509; vendor says 6M downloads | fact R90 2 | unknown capital/team | est medium cadence risk | unknown | fact local-first and hosted providers | unknown centralized SSO/audit | fact analytics opt-in/off by default | est everyone |
| fact LocalAI | fact local OpenAI-compatible runtime | fact MIT; 2023 | fact community | fact 48,845/4,412 | fact active | unknown capital/team | est medium maintainer/bus-factor risk | unknown | fact 60+ local backends | fact RBAC/quota capabilities; full campus audit unknown | unknown telemetry | est builders |
| fact llama.cpp | fact local inference runtime | fact MIT; 2023 | fact ggml community; Hugging Face organization since 2026-02-20 | fact 126,898/22,705/2,389 | fact R90 ≥100; est ≈4.26K stars/month in sampled 29 days | fact community/corporate ecosystem; no single funding round | est low license/runway risk; high change rate | unknown harness rating | fact local-first; strong Apple silicon support | unknown native SSO/budget/audit | fact no hosted telemetry intrinsic | est everyone |
| fact llamafile | fact portable local model runner | fact Apache-2.0; 2023 | fact Mozilla Ocho/community | fact ≈25.9K stars | fact 0.10.5 on 2026-08-03 | fact Mozilla-backed project | est medium project-scope risk | unknown | fact local-first | unknown native controls | fact no hosted telemetry intrinsic | est builders |
| fact Open WebUI | fact self-hosted chat/RAG front door | fact custom license; 2023-10-06 | fact Open WebUI | fact 150,807/22,036/243 | fact R90 7; latest 0.11.3, 2026-08-31 | unknown funding/team | fact 50+ user branding restriction; several 2026 advisories patched | unknown | fact provider-agnostic/local | fact SSO/roles/admin analytics; deeper controls vary | fact external tracing off; admin analytics on | est everyone, governed self-host |
| fact LibreChat | fact self-hosted multi-model chat | fact MIT; 2023 | fact community/project company | fact 42,771/8,849/732 | fact R90 4; latest release candidate | unknown capital/team | est medium maintainer/release-cadence risk | unknown | fact provider-agnostic/local | fact OAuth, LDAP, SAML, roles, token-spend and audit functions | unknown optional providers determine data path | est everyone |
| fact LobeHub | fact chat/agent front end | fact custom license; 2023 | fact LobeHub | fact 82,183 stars; vendor says 6M users | est ≈663 stars/month in sampled 11-day interval | unknown capital/team | est license and vendor-claim risk | unknown | fact multi-provider/local | fact enterprise options; exact campus matrix unknown | unknown default telemetry | est everyone, review license |
| fact AnythingLLM | fact desktop/self-host RAG chat | fact MIT; 2023 | fact Mintplex Labs | fact 65,565/7,247/326; >5M Docker pulls | fact R90 6 | unknown capital/team | est medium company dependence | unknown | fact multi-provider/local | fact enterprise SSO, RBAC and logs | fact desktop states no data phones home | est everyone |
| fact Msty | fact local/multi-model desktop | fact closed; unknown first date | fact Msty | unknown active users | fact active | fact ≈8 staff publicly indicated | est small-team/closed-client risk | unknown | fact local and hosted | unknown complete campus controls; certifications incomplete | fact zero product telemetry stated | est everyone, limited pilot |
| fact Cherry Studio | fact desktop multi-model client | fact AGPL-3.0; 2024 | fact Cherry Studio/community | fact 51,397 stars | fact active | unknown | est medium governance/copyright review risk | unknown | fact multi-provider/local | unknown centralized controls | unknown telemetry | est everyone |
| fact PeerLLM | fact local desktop assistant | fact proprietary; local mode 2026-07-19 | unknown owner/backer | unknown public scale | fact new Q3 2026 | unknown | est very high early-stage risk | unknown | fact local mode claimed | unknown | unknown | est not suitable yet |
| fact Mantle Chat | fact local desktop assistant | fact closed alpha; 2026-08-06 | unknown owner/backer | unknown public scale | fact new Q3 2026 | fact ≈2-person team | est very high bus-factor/alpha risk | unknown | fact local claims | unknown | unknown | est not suitable yet |
| fact Chatbox | fact desktop multi-model chat | fact GPL; 2023 | fact community/company | fact 41,633/4,224/1,261 | fact R90 7 | unknown | est medium issue-load risk | unknown | fact multi-provider/local | unknown campus controls | unknown telemetry | est everyone |
| fact Text generation web UI | fact local model web UI | fact AGPL-3.0; 2022 | fact community | fact ≈47.2K stars | fact latest release 2026-05-20 | unknown | est medium cadence/maintainer risk | unknown | fact local-first | unknown enterprise controls | unknown telemetry | est researchers |
| fact KoboldCpp | fact local model runner/UI | fact AGPL-3.0; 2023 | fact community | fact ≈11.4K stars | fact release 2026-08-16 | unknown | est medium bus-factor risk | unknown | fact local-first | unknown enterprise controls | unknown telemetry | est researchers |
Evaluation, observability and harness-adjacent systems — 15
| Harness | Form | License · first release | Owner / backer | Public scale | Growth / release | Capital / team | Stability | Rating / benchmark | Provider / local | SSO · budget · audit | Data / telemetry default | UVU verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| fact Langfuse | fact tracing/evaluation platform | fact MIT core plus EE; 2022 | fact ClickHouse since 2026-01 | fact 34,155/3,688/881; npm 7,933,285 in Aug; PyPI 25,881,197/month | fact R90 ≥100 | fact acquired by ClickHouse; ≈16+ staff before/around acquisition | est strong parent; integration and high-change risk | unknown independent rating | fact provider-agnostic/self-host | fact SSO/RBAC/audit in enterprise | fact product telemetry on; some enterprise self-host telemetry cannot be disabled | est builders |
| fact Phoenix | fact tracing/evaluation UI | fact ELv2 server; some clients Apache-2.0; 2023 | fact Arize AI | fact 11,307/1,094/969 | fact R90 97 | fact Arize $70M Series C, 2025-02 | est license and rapid-change risk | unknown independent rating | fact provider-agnostic/self-host | fact authentication available but off by default | fact analytics on by default | est builders |
| fact Helicone | fact LLM gateway/observability | fact Apache-2.0; 2023 | fact Mintlify since 2026-03-03 | fact 6,131/662/156 | fact source active; latest GitHub release 2025-08 | fact acquired; price unknown | est acquisition/release-process risk | unknown | fact provider-agnostic | fact hosted enterprise controls | unknown telemetry beyond observed requests | est builders |
| fact Braintrust | fact evaluation/observability platform | fact closed platform; SDK components OSS | fact Braintrust | unknown public active-user count | fact active | fact $80M Series B, 2026-02-17; at least $121M total | est strong capital; proprietary control-plane risk | unknown independent rating | fact provider-agnostic | fact enterprise SSO/audit/access controls | fact managed system/billing telemetry remains on | est builders |
| fact W&B Weave | fact evaluation/tracing platform | fact Apache-2.0 SDK; hosted platform closed | fact Weights & Biases | fact ≈1,125 stars | fact active | fact W&B-backed | est medium platform dependence | unknown | fact provider-agnostic | fact enterprise identity/audit | unknown default | est builders |
| fact Opik | fact evaluation/observability platform | fact Apache-2.0; 2024 | fact Comet | fact 21,763/1,744/238; npm 156,598 in Aug | fact R90 90 | fact Comet-backed | est high release-change risk | unknown | fact provider-agnostic/self-host | fact enterprise controls in commercial tier | unknown telemetry | est builders |
| fact Promptfoo | fact red-team/evaluation CLI and platform | fact MIT; 2023 | fact Promptfoo; OpenAI acquisition agreement 2026-03-09 | fact 24,785/2,258/571; npm 2,581,681 in Aug; vendor says 350K developers/130K MAU | fact R90 11 | fact acquisition agreement; price unknown | est medium integration/vendor risk | fact evaluation results are configuration-specific | fact provider-agnostic/local/offline | fact enterprise controls; CLI can run without hosted service | fact telemetry on by default | est builders |
| fact DeepEval | fact evaluation framework | fact Apache-2.0; 2023 | fact Confident AI | fact 18,081 stars | fact active | unknown current capital/team | est medium vendor/change risk | fact outputs depend on judge model and rubric | fact provider-agnostic/local | unknown core enterprise controls | unknown telemetry | est builders |
| fact Ragas | fact RAG evaluation framework | fact Apache-2.0; 2023 | fact Exploding Gradients | fact 15,606 stars | fact last push 2026-02-24; R90 0 | unknown | est high staleness risk | fact judge/model dependent | fact provider-agnostic/local | unknown | unknown | est not suitable as sole standard |
| fact OpenLLMetry | fact OpenTelemetry instrumentation | fact Apache-2.0; 2023 | fact Traceloop | fact 7,414 stars; npm 653,998 in sampled month | fact active | unknown current capital/team | est medium vendor/maintainer risk | unknown | fact provider-agnostic | unknown control plane; backend supplies controls | est sends configured traces to selected backend | est builders |
| fact AgentOps | fact agent observability SDK | fact MIT; 2023 | fact AgentOps | fact 5,810 stars | fact latest GitHub release 2025-08 | unknown | est high cadence risk | unknown | fact provider-agnostic | fact hosted controls vary | unknown telemetry beyond task traces | est not suitable as standard |
| fact LangWatch | fact evaluation/observability | fact Apache-2.0 core plus EE; 2023 | fact LangWatch | fact 3,522 stars | fact R90 70 | unknown capital/team | est high change/small-team risk | unknown | fact provider-agnostic/self-host | fact enterprise tier controls | unknown telemetry | est builders |
| fact MLflow | fact model/app evaluation and tracking | fact Apache-2.0; 2018 | fact Linux Foundation/Databricks ecosystem | fact 27,795 stars | fact current 3.13 line | fact foundation/corporate-backed | est favorable longevity; broad-scope complexity | unknown harness rating | fact provider-agnostic/self-host | fact RBAC and SSO plugin paths | fact usage tracking on by default | est builders |
| fact Portkey | fact gateway/observability | fact MIT; 2023 | fact Portkey | fact 12,890 stars | fact last push 2026-05-25; no R90 release | unknown current capital/team | est medium/high cadence risk | unknown | fact provider-agnostic/self-host gateway | fact enterprise controls | unknown telemetry | est builders, watch |
| fact Galileo | fact evaluation/observability platform | fact closed | fact Galileo | unknown active customers | fact active | fact $45M Series B, 2024-10; $68M total | est medium proprietary-vendor risk | unknown independent rating | fact provider-agnostic | fact enterprise SSO/audit/governance | unknown default | est builders |
What X says — the practitioner signal, last 90 days
The third feed: a Grok Build research thread searched public X posts from June 5 to September 3, 2026, weighting the last thirty days, for what practitioners are adopting, abandoning, praising, and complaining about. Distinct-account mention counts are estimates from sampled queries, not a census; funding and incidents are recorded only when a post or linked announcement states them. It ran the morning of a simultaneous Claude Code, Codex, Cursor, and Grok Build outage — which is itself the argument for owning a floor.
- FACT (2026-09-03): Claude Code, Codex, Cursor, and Grok Build failed together. Anthropic staff: partial outage across Claude Code, API, claude.ai (@cjav_dev).
- FACT (JetBrains, cited on X): workplace use Claude Code 39% (from 18% Jan), Copilot 21%, Codex 16% (from 3%), Cursor 12% (from 18%), OpenCode 7%.
- EST: Default stack is Claude Code + Codex. OpenCode Zen is the outage/free-model hatch. IDEs (Cursor, Copilot) remain, but the argument is harness, not autocomplete.
- FACT: Muse Spark 1.3 Contributor Free on OpenCode Zen is today’s meme; a minority *do* notice Meta may train (@_yama_yu_, @CDGalpha). Go vs Zen: users pick free if training is already assumed.
- FACT: Antigravity ToS names OpenClaw-style third-party use as Google-account-suspendable. Staff say stale; written ToS still wins the thread (@GergelyOrosz).
- FACT as posted: Cursor = SpaceX/xAI ($60B cited via Bloomberg). Cognition/Devin ~$1B at $47B, ARR >$900M. Kilo = Anaconda (price unknown). Cline enterprise = Spec Driven + LG CNS, quiet.
- FACT: FrontierHarness (360 runs): same model, pass 50–67%, $1.05 (Exo) to $18.34 (Claude Code). Codex if you don’t want to think (@guanlan).
- EST star-spikes: DeepSeek Harness suspect; Graphify mixed; Ponytail/ECC thin; Caveman is a skill; qm is real YC with launch-star inflation; OpenClaw has real users and abandonment.
- FACT: Langflow is the CVE story (CISA RCE). Goose is AAIF/LF plumbing, not the daily CLI. OpenHands is the OSS autonomous list + one faculty local stack. Aider is named more than used.
- FACT: Open WebUI license still draws “not real OSS” drops. LibreChat vs Open WebUI relicensing war is thin this window.
- FACT: Higher-ed on X is Stanford CS146S (agents course, 2026-09-22) and Claude Campus Ambassadors — not campus-wide harness deployments. UVU unnamed.
- EST for UVU: keep a local/OpenCode floor; treat Spark contributor-free as public-data only; do not buy star-spike harnesses; Antigravity is a ToS risk until Google rewrites it.
Rankings as seen on X
- Method (5 lines). (1) JetBrains workplace % when an X post cites it. (2) First-person “I use / I switched / it is down so I cannot work” over listicles. (3) Official maintainer posts count as product-alive, not as adoption. (4) GitHub star spikes on X count for *rising interest*, not *adopted now*, unless diaries exist. (5) Last-30d weighted over last-90d.
- Most adopted now (est): Claude Code; GitHub Copilot (installed, share falling); Codex CLI; Cursor; OpenCode; n8n (automation lane); ChatGPT/Claude.ai as the chat door; Browser Use in the browser lane; Hermes in the OSS-always-on lane.
- Fastest rising last 90d (est): Codex (3%→16% cited); Claude Code (18%→39%); OpenCode Zen + Spark 1.3 free; Grok Build (SpaceX stack); Mastra (TS shipping); DeepSeek Harness/Graphify/qm *stars*; Cognition capital.
- Newest ≤6 months with real traction: Antigravity (conversation, negative ToS); Muse Spark 1.3 (model, not harness); Stagehand v4; Mastra computer-use; qm (YC, launch spike); AQ meta-harness; Factory public-sector. “Real” here means named workflows or eval inclusion, not stars alone.
- Declining or dead (est): Windsurf as a brand (Cognition/Google split residue); Copilot *share*; Aider vs new CLIs; Roo Code (thin/likely archived); Continue (listicle only); personal OpenClaw at Every; Langflow *trust*; AutoGen for new builds.
Method: Native X keyword + semantic search, public posts only. Windows: last 30 (since:2026-08-04) and last 90 (since:2026-06-05). Latest plus min_faves filters. No DMs, no posting. Distinct-account mention counts are est: unique handles in returned samples (max 10 per query), scaled by how many distinct queries hit the name and whether official accounts dominate. Not a census. Funding, acquisitions, license changes, and incidents are fact only when a post or linked announcement states them. No invented numbers. Rankings prefer practitioner “I use / I switched” over listicle spam. Star counts on X are interest, not adoption. Misses are logged. A miss is not proof of absence.
Star-spike provenance — are the new "100k-star" harnesses real?
| Project | X verdict | Why |
|---|---|---|
| DeepSeek Harness | Suspect-until-proven | 100k stars in <48h claimed; FrontierHarness includes “DSH” as a real harness under test (good). Counter: Julian Goldie SOP-for-DM videos, crypto “TaiYi Super Agent” spam, almost no named production teams. Pin versions if you touch it. |
| Ponytail | Thin users / mixed | Trending (+1.3k/day cited). Metric claims (54/20/27%) are promo. A few “old oil” skill posts. Not a daily-driver. |
| ECC | Unverified | Trending explainers only. No first-person ops posts in this sample. |
| Graphify | Mixed — demo real, scale unverified | Viral 100k-star video (208k views). One engineer: it cuts reread-token burn. Used as a clone-demo target. Not seen as a workplace standard. |
| Caveman | Not the spike people think | In this window it is a token-saving *skill* (“talks like caveman”), not a 100k-star platform. Full spike check still open. |
| OpenClaw | Real users, messy ops | ToS-named, Steinberger “we all use it,” Every killed personal Claws for a shared Slack agent. Stars are inflated relative to maintainability, but this is not a ghost repo. |
| qm | Real org, marketing-heavy spike | YC open-sourced Jul 2026 as the firm’s Slack+web multiplayer harness (MIT, yc-software/qm). Launch posts claim 1.6k–13k stars in days and internal accounting/legal/eng use. Last-30d first-person ops almost absent; Sep conversation moved on. Treat as real product with inflated launch stars. |
| Paperclip | Real niche, not a spike scam | Official release cadence, adapters, orchestrator comparisons. Name-collides with paperclip-AI-risk essays and a patent search tool. |
Funding and ownership moves stated on X (dated)
| Date | Item | Amount | Source | Grade |
|---|---|---|---|---|
| 2026-09-03 | NVIDIA intends to acquire Hugging Face | $12,930,300,000 | @ClementDelangue | fact (announcement post). Not a coding harness; cited because X grouped it with Cursor/Kilo. |
| 2026-09-03 | Anaconda acquired Kilo Code | price not stated | @GargEtisha; @kilocode bio | fact acquisition; unknown price |
| last 30d (cited 2026-09-03) | SpaceX acquired Cursor | price not stated in sampled posts | multiple, e.g. @KevinhoMorales | fact acquisition; unknown price |
| 2026-09-02 | SpaceX Cursor deal | $60B | @wallstengine citing Bloomberg | fact as that post states; primary Bloomberg not opened |
| 2026-09-02 | Cognition (Devin) new round | ~$1B at $47B val; ARR >$900M (from $492M late May); ~$10B demand | @wallstengine citing Bloomberg | fact as posted; primary unknown |
| last 30d (cited 2026-09-03) | Stripe acquired OpenRouter | price not stated | @GargEtisha | fact as stated by that post; not independently confirmed in this batch |
| 2026-08-29 (Grok summary) | Google DeepMind hired Windsurf founders + $2.4B tech license; Cognition took remainder | $2.4B | @grok | est/secondary until primary 2025 posts are fetched |
Incidents and stability signals on X (dated)
| Date | Item | Source | Grade |
|---|---|---|---|
| 2026-09-03 | Partial outage: Claude Code + API + claude.ai | @cjav_dev | fact |
| 2026-09-03 | Codex/ChatGPT down same morning | @ProductHunt, many | fact (widespread user + PH) |
| 2026-09-03 | Cursor + Grok Build also failing in user screenshots | @sora_biz | fact as user report |
| 2026-09-03 | OpenCode Zen Spark 1.3 free live; some 500s on chat-completions | @hppzyy | fact user |
| 2026-09-03 | Antigravity ToS / Google-account ban debate | @theo, @GergelyOrosz | fact posts; enforcement unknown |
| 2026-09-02 | n8n agent silently removed 92% of AI nodes (HN via X) | @echo_vic | fact as cited; original HN not opened this batch |
| 2026-09-02 | GitSpawn: goose 1.44.0 patched; Hermes/Qwen Code/Grok Build still open on retest | @mkeremturhan | fact researcher claim |
| 2026-09-01 | “One flaw affects nearly every major AI coding agent” — 8 findings, 4 unpatched | @Ax_Sharma | fact as journalist thread; details not fully fetched |
| 2026-06-08 | Cline Spec Driven enterprise with LG CNS | @cline | fact |
| 2026-08-05 (cited 08-27) | Langflow unauth RCE, CISA active exploitation | @SnowCrashLabs | fact as cited |
| 2026-08-10 | Stagehand v4 | @Stagehanddev | fact |
| 2026-09-02 | Mastra sandbox computer-use | @calcsam | fact |
| 2026-09-02 | FrontierHarness Eval published | @guanlan | fact |
| 2026-09-03 | Browser Use × Link agent cards | @browser_use | fact |
| 2026-09-03 | Muse Spark 1.3 contributor-free on OpenCode Zen | many; training caveat @CDGalpha | fact |
Higher education sightings
| Date | Sighting | Source | Grade |
|---|---|---|---|
| 2026-09-02 | Stanford CS146S *The Modern Software Developer* (Fall 2026): 85% of 2025 material thrown out; agent skills, MCP, AGENTS.md, software factories; OSS PR requirement with OpenHands, Browserbase, CrewAI, Warp, Vercel, Pi, etc.; guests from Cursor, Claude Code, Factory, Cognition, Replit. Campus start 2026-09-22; materials public. | @mihail_eric | fact (instructor announcement) |
| 2026-09-02 | Claude Campus Ambassadors 2026–27: undergrad/grad/PhD tracks; $3,600 scholarship cited; apply by 2026-09-12 | @claudeai | fact |
| 2026-09-01 | Iowa State: SpaceXAI Grok Bot build night 2026-09-03; free Cursor Pro for attendees | @jackxlau | fact event post |
| 2026-09-01 | Hult International Business School: 2026 stack includes Cursor Boston, Lovable, ChatGPT, Claude, Gemini, Copilot | @HULT_JP | fact school comms |
| 2026-08-31 | University security faculty: DGX Spark ×2 + vLLM + qwen3.8-flash-next + OpenHands local stack | @valdzone | fact post; campus-wide unknown |
| 2026-09-03 | Student moved housing to afford extra Claude Code Max | @enjojoyy | fact anecdote, not a deployment |
| — | UVU-named public deployment | — | unknown |
No campus-wide harness deployment by any university surfaced on X in the window; the sightings are courses, ambassador programs, build nights, and one security faculty member's local stack. UVU is not named anywhere. That absence is consistent with the guide's first-mover finding.
The X master table — 85 tools with momentum, incidents, and what people say they are best and worst at
| # | Tool | One-line | Category | Open / closed | Company / backer | Funding (stated only) | X momentum, last 30 days | Stability signals | Scores people cite | Best at / worst at |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | OpenCode | Model-agnostic CLI/TUI coding agent; Zen is the hosted model router (Go paid, Zen/contributor-free routes). | Coding CLI | Open (MIT claimed in promo posts) | Anomaly / @thdxr | unknown in this X sample | High. Distinct handles in Zen/Spark posts today; JetBrains 7% cited. Method: 8+ unique handles in Zen/Spark queries plus survey cite. Sentiment lean: positive-as-escape-hatch. | fact: Zen Spark 1.3 free is live; some 500s unless /v1/responses (Codex-compat) is used (@hppzyy). | unknown official harness score. People compare Spark-on-OpenCode vs Gemini 3.8 by feel. | Best: keep working when Claude/Codex/Grok die; cheap/free models. Worst: not for non-programmers (@Anurag_barl). |
| 2 | Claude Code | Anthropic’s CLI/cloud coding agent. Current default in the JetBrains workplace survey. | Coding CLI | Closed client | Anthropic | unknown amount on X this window | Very high. Outage + survey + switch posts. 10/10 latest “switched” hits involved it or Codex. Sentiment: default, plus reliability anger. | fact: partial outage 2026-09-03 covering Claude Code, API, claude.ai (@cjav_dev). Practitioner tally: “3 major outages + 6 degraded last month” (@alexgetmancom) — est, not vendor. | JetBrains 39% workplace use. Terminal-Bench cited as model+harness, not harness-only. | Best: default agentic work. Worst: capacity/outages; cost (“overpriced P.C. sludge” in one SaaS-stack post). |
| 3 | Codex CLI | OpenAI coding agent (CLI + app). Fastest-rising in the cited JetBrains numbers. | Coding CLI | Open CLI (Apache claimed in other lanes; not re-stated on X here) | OpenAI | unknown on X | Very high. Paired with Claude Code in every outage thread. Sentiment: preferred by some switchers; limit/outage complaints. | fact: ChatGPT/Codex down 2026-09-03; OpenAI posted during the incident. First dual Claude+Codex outage noted (@LexnLin). | JetBrains 16% (from 3%). Terminal-Bench 4.0 runs paired with gpt-5.6-luna. | Best: concise, rule-following, easy switch from Claude projects (@JamesonCamp). Worst: limits; app vs CLI quality split (@scottzirkel). |
| 4 | Gemini CLI | Google’s open coding CLI; now discussed next to Antigravity and account-ban risk. | Coding CLI | Open (Apache claimed in listicles) | Google-backed; no round on X | Medium-high. Less “I use Gemini CLI daily,” more ban/ToS adjacency. | fact: @GergelyOrosz says Gemini CLI access was hit when Antigravity bans landed. Consumer CLI path already in transition (prior lane; not re-confirmed here). | unknown harness-only. Gemini 3.8 Flash cost-per-task up on Artificial Analysis (@agentfred_ai). | Best: huge Flash limits on $20 plans (@championswimmer). Worst: Google-account blast radius. | |
| 5 | Antigravity | Google DeepMind agentic IDE. ToS currently the story, not the product. | Coding IDE | Closed preview | Google / @\_mohansolo | Google-backed; Windsurf-team origin is contested on X | High (negative). Theo 793k views; Gergely 72k. Sentiment lean: avoid. | fact: ToS cited as allowing Google-account suspension for third-party/OpenClaw use (@GergelyOrosz). Counter: Varun Mohan said bans were product-only (quoted). Shadow-ban support complaints exist. | JetBrains 6% in prior lane; not restated in this X sample. | Best: Gemini 3.8 Flash inside Google surfaces (@DjowDj0w). Worst: ToS vs Google-account risk; “ugly/hard” vs Windsurf-team expectations (@LinearUncle). |
| 6 | Cursor | Closed AI editor, now a SpaceX/xAI asset. | Coding IDE | Closed | Cursor → SpaceX (fact on X) | Acquisition price unknown on X | Very high. Survey drop + SpaceX narrative + today’s outage. Sentiment: mixed; still in every stack poll. | fact: outage 2026-09-03 with Claude/Codex. fact: ambassadors renamed SpaceXAI (@KevinhoMorales). est: OpenAI access cut talk (28 Aug notice / 12 Nov shutoff) is circulating, not confirmed here from OpenAI’s own account. | JetBrains 12% (from 18%). CursorBench cited by @GavinSBaker. | Best: Composer for simpler tasks; GUI. Worst: vendor concentration after SpaceX; model-provider cutoff risk. |
| 7 | Windsurf | AI editor; Cognition-owned remainder after the Google/Cognition split. | Coding IDE | Closed | Cognition (product); Google hired founders (X summaries) | No new round in this sample. Historical $2.4B Google license / Cognition remainder is summarized by @grok — treat as est until primary post is fetched. | Low-medium. Mostly listicle residue. Few first-person “I switched to Windsurf” hits. | fact: “Anthropic pulled the plug on Windsurf” as a 2025-era event still cited (@kevinsxu). est: declining mindshare vs Claude Code/Codex/Cursor. | unknown current. | Best: still named as a paid option. Worst: not the conversation; Antigravity comparison is unkind. |
| 8 | Cline | Open-source IDE/CLI coding agent; enterprise Spec Driven with LG CNS. | Coding IDE/CLI | Open | Cline | unknown amount on X this window | Medium. Official posts + OSS listicles + one enterprise inbound. Not in today’s outage pantheon. | fact: Cline Spec Driven enterprise platform with LG CNS, 2026-06-08 (@cline). fact: inbound “enterprise license for our team” 2026-09-03 (@shettygirish75). npm-compromise history not discussed in this 30d sample. | Marketplace ratings not cited on X here. | Best: editor+terminal+browser agent; BYO model. Worst: quieter than Claude/Codex; enterprise story not viral. |
| 9 | Kilo Code | Open-source agent for VS Code, JetBrains, CLI; posting as acquired by Anaconda. | Coding IDE/CLI | Open | Kilo / Anaconda | Acquisition fact; price unknown | Medium. Official @kilocode volume; third-party mentions are listicles. Sentiment: vendor-positive, community-thin. | fact: handle/bio “Kilo (acq. by Anaconda)”; grouped with Cursor/OpenRouter deals in last 30 days (@GargEtisha). Shipping JetBrains multi-agent control room (2026-09-01). | unknown independent. Demo: Fable 5.1 recreate-X for $13 (@kilocode). | Best: OSS hedge vs closed SF stack (their blog). Worst: little independent “I switched to Kilo after Anaconda” testimony in this sample. |
| 10 | Goose | Block-origin local agent; now under AAIF / Linux Foundation orbit. | Coding/general agent | Open | Block → AAIF (@goose_oss, @AgenticAIFdn) | Corporate/foundation; no round | Medium. Security + OSS-list + MCP demos. Not a daily-driver meme. | fact: GitSpawn fsmonitor patched in goose 1.44.0 (@mkeremturhan). fact: AAIF ambassador shipped goose-extension-audit (@AgenticAIFdn). | unknown. | Best: MCP recipes, local/multi-provider, foundation governance. Worst: lower X volume than Claude/Codex; still a “list of 10” item. |
| 11 | OpenHands | Autonomous coding platform (inspect, edit, test, PR). Agent Canvas as control plane. | Coding agent platform | Open (MIT claimed) | All Hands / @OpenHandsDev | unknown on X | Medium. Star-count listicles (85.7k) plus one faculty-adjacent local stack. | fact: OpenHands SDK now defaults Browser Use (@mamagnus00). Databricks event with OpenHands on security agents (Sep 1). | Star counts cited (78.5k–85.7k). Not harness-only benches. | Best: autonomous issue→PR; self-host; local+vLLM stack (@valdzone). Worst: Docker/setup friction; promo-list noise. |
| 12 | Aider | Terminal pair-programmer with git-aware edits. | Coding CLI | Open | Independent | unknown | Low-medium. Named in “what’s your harness” polls and OSS lists more than first-person 30d diaries. | unknown incidents this window. | Aider polyglot not cited in this 30d sample. | Best: terminal+git, any model. Worst: mindshare lost to Claude Code/Codex/OpenCode. |
| 13 | LibreChat | Self-hosted multi-model chat UI. | Chat front door | Open | LibreChat | unknown | Low. Thin 90d X. One MCP success vs Open WebUI failure. | unknown license-change posts this window. | unknown. | Best: Playwright MCP over SSE after Open WebUI failed (@mfaisal_khatri). Worst: almost invisible vs Open WebUI in this scan. |
| 14 | Open WebUI | Self-hosted local/cloud model UI. Branding/enterprise-license friction. | Chat front door | Source-available (branding constraint cited) | Open WebUI | unknown | Medium-low. License complaints + enterprise fork talk, not daily coding-agent chatter. | fact: users drop it because “the license is stupid. It's not a real open source license” (@libertypenguin0). Branding/white-label needs enterprise license (Grok summary 2026-06-20). Forks of v0.6.5 for MIT-friendlier stacks. | unknown. | Best: local RAG/gateway. Worst: license/branding for white-label; MCP gaps vs LibreChat in one practitioner thread. |
| 15 | n8n | Workflow automation with AI-agent nodes; still the default “no-code agent” on X. | Workflow/agent automation | Open core + cloud | n8n | unknown on X | High (automation lane), medium (coding-agent lane). Distinct from CLI coding; huge tutorial volume. | fact: agent silently deleted 92% of AI nodes in a dataset (HN, cited 2026-09-02) — governance miss (@echo_vic). | unknown coding benches. | Best: Trigger→Brain→Memory→Tools without code. Worst: change-control; some builders now prefer Python/agent scripts over n8n (@uglyrobot). |
| 16 | Dify | Visual agentic-workflow builder. Official account still posting “build your first agent” guides. | Workflow/agent platform | Open + cloud | Dify | unknown on X | Medium-low. Official posts; little practitioner debate vs n8n/Langflow. | unknown incidents this window. | unknown. | Best: start-from-a-goal onboarding. Worst: not in the coding-agent default stack. |
| 17 | Langflow | Visual agent/RAG builder (IBM-associated in security writeups). | Workflow/agent builder | Open | IBM / Langflow | unknown | Medium (security). Mentions are CVE/RCE, not adoption. | fact: CISA-confirmed unauth RCE, default no login (@SnowCrashLabs). Repeat exec() sink (CVE-2025-3248 and CVE-2026-33017) (@BRuteLogic). “11 more CVEs” cited 2026-09-02. | unknown. | Best: visual RAG/agent graphs. Worst: internet-exposed instances; secrets next to exec(). |
| 18 | CrewAI | Multi-agent “crews” framework. | Agent framework | Open + platform | CrewAI | unknown on X | Medium-low. Listicle staple; one telemetry critique. | est: “78 observe-only hooks… that’s telemetry” (@_kvnloo). | unknown. | Best: named multi-agent starter. Worst: bloat/observe-only; skipped by some greenfield builders. |
| 19 | Mastra | TypeScript agent framework; shipping sandbox computer-use and Render deploy. | Agent framework | Open core + commercial EE (prior lane; not re-stated on X) | Mastra / @calcsam | unknown amount on X | Medium. Founder launch posts + Render partnership. Sentiment: builder-positive. | fact: sandbox computer-use launch 2026-09-02 (@calcsam). Suspended-run recovery across servers (@mastra). | unknown harness benches. | Best: TS agents, long-running workflows, computer-use in sandbox. Worst: not a coding CLI; framework not daily IDE. |
| 20 | Pydantic AI | Typed Python agent framework from the Pydantic team. | Agent framework | Open | Pydantic | unknown on X | Medium-low. CopilotKit channel glue + “type-safe outputs” listicles. | unknown incidents. | unknown. | Best: schema-valid agent outputs. Worst: not a coding harness; library not product. |
| 21 | Browser Use | OSS + cloud “agents that use the browser”; payments and cookie-sync this week. | Browser agent | Open + cloud | Browser Use | unknown amount on X | High. Official product + staff departures to SpaceXAI. | fact: Link one-time cards for agent checkout 2026-09-03 (@browser_use). fact: engineer #2 left after “20k to over a million monthly agent runs” (@Alezander9). Staff → SpaceXAI (@larsencc). | Odyssey #1 claimed in promo posts. | Best: real-browser loop; now in OpenHands SDK default. Worst: action/payment risk; telemetry not discussed on X here. |
| 22 | Stagehand | Browser-agent SDK by Browserbase; v4 self-healing + WebMCP. | Browser SDK | Open | Browserbase | unknown (parent Series B not restated on X) | Medium. Official v4 (2026-08-10) still circulating. | fact: v4 caching up to 80% faster claimed (@Stagehanddev). WebMCP first-class. | unknown independent. | Best: Playwright-for-agents; iframe/self-heal. Worst: Browserbase cloud coupling. |
| 23 | Browserbase | Hosted browser runtime for agents. | Browser infra | Closed | Browserbase | unknown on X | Medium. Identity/payments/Ramp, not coding-IDE talk. | fact: Ramp agent-identity partnership (@browserbase). | unknown. | Best: authenticated web + contexts. Worst: hosted-browser data plane. |
| 24 | Muse Code + Spark 1.3 | Meta’s coding agent (Muse Code) and model (Spark 1.3). Free contributor route via OpenCode Zen. | Coding IDE + model | Closed | Meta | Meta-backed; no round | Very high (model), medium (Muse Code product). Spark free is today’s meme. | fact: contributor-free may train (@CDGalpha). fact: Muse Code $15 plan limit complaint (@dalu_hey). | AA cited; Meta DeepSWE/Terminal-Bench in Grok replies. Independent 3D-animal loop: Spark cheaper than Gemini 3.8 Flash (@thehypedotnews). | Best: free/cheap coding while Claude/Codex down. Worst: training terms; Muse Code limits; not local weights. |
| 25 | DeepSeek Harness | Plugin-everything coding harness; 100k GitHub stars in <48h claimed. | Coding harness | Open (MIT claimed) | unknown org on X | unknown | High (hype), low (verified use). Promo videos + trending bots. | fact as claimed: 100k stars <48h, faster than OpenClaw (@kubesimplify). In FrontierHarness set as “DSH.” | FrontierHarness: wall-clock play, knobs. Star counts 135k-in-days in promo (@JulianGoldieSEO) — treat as est/promo. | Best: swappable loop/plugins. Worst: star-spike + affiliate SOP-for-DM pattern; pin versions. Provenance: suspect-until-proven. |
| 26 | Ponytail | Skill/harness that forces reuse-before-code. Trending GitHub. | Coding skill/harness | Open | unknown | unknown | Medium (trending), low (use). | fact: GitHub trending +1.3k stars 2026-09-02 (@soresearcher). Claimed 54% less code / 20% lower cost / 27% faster on 12 tasks (@gitgoats) — vendor/promo, not replicated here. | unknown independent. | Best: “do less” ladder. Worst: star velocity vs first-person production posts. Provenance: mixed/thin users. |
| 27 | Paperclip | OSS control panel for a team of agents (skills, adapters, operators). | Agent orchestrator | Open | paperclipai / @papercliping | unknown | Medium. Official releases + orchestrator roundups. Name collision with “paperclip problem.” | fact: v2026.831.0, Kimi Code adapter, 175 commits / 13 contributors (@papercliping). Telemetry/observability/run-log naming split in v2026.824.0. | unknown. | Best: agent-as-employees dashboard. Worst: early, issue-heavy (prior lane); not a coding CLI. Users: real but niche. |
| 28 | OpenClaw | Personal/general autonomous agent (Claw). Named in Antigravity ToS. | General agent | unknown license on X | @steipete / OpenClaw | unknown | High. ToS villain, real users, and abandonment in the same week. | fact: Every “killed the last of our Claws” — logins expired, integrations broke (@every). fact: ToS example for Google bans. Peter Steinberger: “We all use it.” | unknown benches. | Best: personal always-on agent. Worst: maintenance; account-ban coupling; governance. Users: real, not only stars. |
| 29 | ECC | Repeatable plan/build/test/review toolkit for coding agents (skills, hooks, scans). | Coding toolkit | Open (MIT in prior lane) | unknown | unknown | Low-medium. GitHub-trending explainers, almost no “I run ECC daily.” | unknown incidents. | unknown. | Best: shared process across Claude Code/Codex/OpenCode. Worst: star-without-users pattern. Provenance: unverified. |
| 30 | Graphify | Local-folder → knowledge graph for coding agents (AST, no vectors claimed). | Context harness | Open | Graphify-Labs / @safishamsii | unknown | Medium-high (viral demo). One 208k-view ES video. | unknown security. | 100k GitHub stars claimed in viral post (@SofiaSici). | Best: stop agents rereading files (@i_mika_el). Worst: star-spike; few independent production writeups. Provenance: mixed — demo real, scale unverified. |
| 31 | Grok Build | xAI coding agent (CLI/client). In today’s outage set and GitSpawn unpatched list. | Coding CLI | Source published (prior lane) | xAI / SpaceX | xAI-backed | High. Outage screenshots + release notes + stack posts. | fact: down 2026-09-03 with Claude/Codex. fact: GitSpawn still open on retest (@mkeremturhan). v1.0.18 managed MCP policy (@XFreeze). | unknown independent harness score. | Best: daily coding with Grok 4.6 (@elshayib_). Worst: unpatched GitSpawn; issues-disabled governance (prior lane, not re-litigated here). |
| 32 | GitHub Copilot | Microsoft/GitHub IDE+cloud coding agent. Survey share falling; Slack/Teams agent shipping. | Coding IDE/cloud | Closed | Microsoft/GitHub | Microsoft-backed | High (installed base), mixed sentiment. Still #2 in JetBrains cite. | fact: Copilot in Slack and Teams, Aug ship (@github). Practitioner: GitHub-hosted issue→PR agent (@Equinox_enx). | JetBrains 21% (from 29% a year ago). | Best: GitHub-native issue agent; student free. Worst: losing agent mindshare to Claude Code/Codex. |
| 33 | Amp | Sourcegraph coding agent (TUI/desktop/web/mobile). | Coding agent | Closed | Sourcegraph / @ampcode | unknown on X | Medium. Enthusiastic niche, not survey-scale. | fact: intelligent diff-order button (@beyang). | unknown. | Best: cross-surface harness (@anthonywu). Worst: paid closed; quieter than Claude/Cursor. |
| 34 | Devin | Cognition cloud software agent; Windsurf remainder lives here. | Cloud coding agent | Closed | Cognition | fact as posted: ~$1B round at $47B, ARR >$900M, Bloomberg via @wallstengine. Independent filing unknown. | Medium-high (capital), medium (practitioners). | Fusion pricing claims vs Fable 5.1 (@notjazii) — vendor-adj. SWE-1.7 on Cerebras 1,000 t/s. | unknown independent harness-only. | Best: unattended cloud agent; cheap Fusion routing. Worst: another vendor desktop; not the CLI default. |
| 35 | Crush | Charmbracelet terminal coding agent; FSL license. | Coding TUI | Source-available (FSL) | Charmbracelet | unknown | Low-medium. 27k-star listicles; sandbox mention. | smolBSD sandbox for crush (@iMilnb). | unknown. | Best: pretty terminal agent. Worst: FSL; not in JetBrains top slice. |
| 36 | Qwen Code | Alibaba/Qwen terminal coding agent. | Coding CLI | Open | Alibaba/Qwen | Alibaba-backed | Medium-low. GitSpawn unpatched; Kilo collab. | fact: GitSpawn still open on 0.19.6/0.22.3 (@mkeremturhan). Qwen3.8-Max #4 Arena frontend, demoed on Kilo (@Alibaba_Qwen). | Arena frontend #4 (model, not harness). | Best: OSS CLI in Qwen ecosystem. Worst: unpatched GitSpawn; jurisdiction. |
| 37 | Hermes Agent | Nous-orbit general/coding agent; 24/7 worker meme. | General + coding agent | Open | Nous / community | unknown | High in OSS-agent lane. | fact: GitSpawn still open 0.18.2/0.21.0. Consulting talk of warehouse/Slack/CTO paths (@tonysimons_). | unknown. | Best: local 24/7 with cheap/open models. Worst: unpatched GitSpawn; “trips over itself” in promo comparisons. |
| 38 | Pi | Minimal coding-agent toolkit; FrontierHarness winner on cost. | Coding harness | Open | Independent (named in eval) | unknown | Medium. Eval-famous more than viral. | fact: 90 turns / $2.50 vs Claude Code 381 / $64.36 on same DeepSWE fix (@guanlan). | FrontierHarness: use if the job repeats. | Best: cheap pass rate. Worst: knobs; less “just works” than Codex. |
| 39 | Exo | Harness in FrontierHarness; cheapest $/pass. | Coding harness | unknown | unknown | unknown | Low-medium. Eval-only in this scan. | unknown ops. | FrontierHarness: $1.05/pass; quit-early. | Best: retries cheap. Worst: almost no practitioner diary. |
| 40 | LangChain | Python agent/LLM framework + LangSmith. | Agent framework | Open + cloud | LangChain Inc. | unknown amount on X this window | Medium. CEO talks, OpenWiki, LangSmith deploy. | fact: LangChain coding agent 52.8%→66.5% Terminal Bench 2.0 without changing model (@nykdotdev). Podium migrated to LangSmith deployments (@LangChain). | Terminal Bench 2.0 harness-not-model. | Best: production harness + traces. Worst: bloat; some skip it greenfield. |
| 41 | LangGraph | Stateful graph agent runtime (LangChain). | Agent framework | Open | LangChain Inc. | parent | Medium. DeerFlow built on it; Stanford treats it as homework parts. | unknown new incidents. | unknown. | Best: multi-agent graphs. Worst: “implementation part not research” (@momiji_fullmoon). |
| 42 | Google ADK | Google agent SDK; workshop circuit. | Agent SDK | Open | Google-backed | Medium. Tutorial volume, not coding-CLI default. | unknown. | unknown. | Best: GCP multi-agent patterns. Worst: Google-account adjacency (see Antigravity). | |
| 43 | AutoGen | Microsoft multi-agent; maintenance mode. | Agent framework | Open | Microsoft | Microsoft-backed | Low. Listicles note successor. | fact as listed: maintenance; new work → Agent Framework (@kv1nsiii). | unknown. | Best: tutorial reference. Worst: not for new builds. |
| 44 | Continue | OSS IDE coding assistant. | Coding IDE | Open | Continue | unknown | Low. Local+Ollama questions; listicles. | unknown shutdown posts in this sample (prior lane said maintenance retreat — not re-confirmed on X here). | DeepSWE 63.4 claimed in one promo. | Best: local VS Code. Worst: not the 2026 conversation. |
| 45 | Roo Code | Cline-family IDE agent. | Coding IDE | Open | Roo Code Inc. | unknown | Low. History extractors still name it; almost no 30d first-person. | unknown archive posts this window (prior lane archived 2026-05-15 — not re-confirmed here). | unknown. | Best: historical Cline fork. Worst: declining/dead on X. |
| 46 | Factory / Droid | Model-agnostic software agent; public-sector push. | Cloud/IDE coding agent | Closed | Factory | unknown amount on X | Medium. Readiness criteria; Carahsoft. | fact: Factory + Carahsoft public sector (@FactoryAI). Long unattended runs 1–8h (@garrettwinderrr). | unknown. | Best: agent-readiness + autonomy. Worst: not CLI-default. |
| 47 | Augment / Cosmos | Enterprise coding agent → “OS for agentic software development.” | Coding platform | Closed | Augment | unknown on X | Low-medium. Official Cosmos Advisor posts. | unknown. | unknown. | Best: factory-scale prompt box. Worst: little independent X. |
| 48 | Lovable | Prompt-to-app builder; Slack @Lovable. | App builder | Closed | Lovable | unknown on X | Medium (builders), low (harness nerds). | fact: Slack tagging (@Lovable). Hult curriculum names it. | unknown. | Best: idea→app. Worst: not a repo agent; token-middleman complaints vs Claude-direct. |
| 49 | Replit Agent | Browser IDE + agent + host. | Cloud IDE agent | Closed | Replit | unknown on X | Medium-low in this coding-agent sample; still in vibe stacks. | unknown incidents. | unknown. | Best: agent+host $20. Worst: not Claude Code/Codex discourse. |
| 50 | qm (Quartermaster) | YC-open-sourced multiplayer company harness (Slack+web). | Agent orchestrator | Open (MIT claimed) | YC / yc-software | YC-backed; no round | Medium at launch (Jul–Aug), low in last 30d. | fact as posted: MIT, Slack+web, harness-agnostic (Pi/OpenCode/Codex/Claude Code), used internally for accounting/legal/eng. Stars 1.6k–13k in days across posts. | unknown benches. | Best: company-wide rooms. Worst: launch-spike, few Sep first-person ops. Provenance: real org (YC), star-spike still marketing-heavy. |
| 51 | Warp | Agentic terminal. Stanford OSS partner list. | Agentic terminal | Closed client | Warp | unknown on X | Low-medium. Named in course partners more than diaries. | unknown. | unknown. | Best: terminal+agent. Worst: quiet vs CLI agents. |
| 52 | Kiro | AWS coding IDE/CLI path (Q successor in prior lane). | Coding IDE | Closed | AWS | Amazon-backed | Low. Giveaway spam + one OSS “Kiro Crew” workspace (may be unrelated). | unknown. | unknown. | Best: AWS-native. Worst: not the X default. |
| 53 | v0 | Vercel UI generator. | App builder | Closed | Vercel | Vercel-backed | Low-medium in vibe stacks. | unknown. | unknown. | Best: UI. Worst: not a repo agent. |
| 54 | Bolt | Prompt-to-hosted-site. | App builder | Closed | StackBlitz | unknown | Low-medium. | unknown. | unknown. | Best: hosted site fast. Worst: not enterprise coding. |
| 55 | Ollama | Local model runner; pairs with OpenCode/Continue/Open WebUI. | Local inference | Open | Ollama | unknown on X | Medium as the local floor, not a harness. | unknown. | unknown. | Best: local weights. Worst: not an agent. |
| 56 | Zed | Fast editor; AI mentioned as respected. | Editor | Open core | Zed | unknown | Low in agent discourse. | unknown. | unknown. | Best: editor. Worst: not agent-default. |
| 57 | SWE-agent | Research GitHub-issue agent; mini-SWE successor path (prior). | Research coding agent | Open | Princeton/community | Academic | Low. OSS lists only. | unknown. | SWE-bench is the point. | Best: research. Worst: not daily driver. |
| 58 | Plandex | Long-running coding engine. | Coding agent | Open | Plandex | unknown | Low. OSS lists. | unknown. | unknown. | Best: long tasks. Worst: thin X. |
| 59 | DeerFlow | ByteDance LangGraph harness; 81.3k stars cited. | Agent harness | Open (MIT cited) | ByteDance | ByteDance-backed | Medium as listicle star. | unknown. | Stars 81.3k cited. | Best: research+chat apps. Worst: jurisdiction; listicle. |
| 60 | Microsoft Agent Framework | AutoGen/SK successor. | Agent SDK | Open | Microsoft | Microsoft-backed | Low-medium. Named as the new path. | fact: AutoGen maintenance → this (@kv1nsiii). | unknown. | Best: new MSFT builds. Worst: early. |
| 61 | OpenAI Agents SDK | Official Python agents SDK. | Agent SDK | Open | OpenAI | OpenAI-backed | Low vs Codex CLI. | unknown. | unknown. | Best: OpenAI-shaped apps. Worst: not the coding CLI. |
| 62 | Strands | AWS agent SDK; DynamoDB memory post. | Agent SDK | Open | AWS | Amazon-backed | Low. | fact: strands-dynamodb-storage blog cited (@VKazulkin). | unknown. | Best: AWS memory. Worst: thin X. |
| 63 | LlamaIndex | Data/agent framework. | Framework | Open + cloud | LlamaIndex | unknown on X | Low this window (listicles). | unknown. | unknown. | Best: private data. Worst: not coding CLI. |
| 64 | Haystack | RAG/agent framework. | Framework | Open | deepset | unknown | Low. | unknown. | unknown. | Best: RAG. Worst: quiet. |
| 65 | Flowise | Visual LLM builder. | Workflow | Open | Flowise | unknown | Low. Not in 30d coding-agent hits. | unknown (prior lane had sunset risk — not X-confirmed here). | unknown. | Best: visual. Worst: miss. |
| 66 | Promptfoo | Eval/red-team for prompts/agents. | Eval | Open | Promptfoo | unknown | Low. Not retrieved this window. | unknown. | unknown. | Best: eval. Worst: miss. |
| 67 | CodeRabbit | PR review agent. | Review agent | Closed | CodeRabbit | unknown on X (prior $143M not restated) | Low this window. | unknown. | unknown. | Best: review. Worst: miss. |
| 68 | Chrome DevTools MCP / Playwright MCP | Browser debugging MCP for agents. | Browser MCP | Open | Google / Microsoft | vendor-backed | Medium. One 2026-09-03 explainer. | fact: Chrome DevTools MCP for Antigravity, Claude, Cursor, Copilot (@FReza1984). | unknown. | Best: observable browser debug. Worst: still MCP, not a harness. |
| 69 | Open Interpreter | Local computer-use / code agent. | Computer-use | Open | Open Interpreter | unknown | Low. OSS lists. | unknown rewrite (prior lane). | unknown. | Best: talk-to-computer. Worst: thin 30d. |
| 70 | Manus | Cloud general agent; Meta-owned (prior). | General agent | Closed | Meta | unknown price | Low this window vs coding CLIs. | unknown. | unknown. | Best: general agent. Worst: miss vs Claude/Codex. |
| 71 | ChatGPT | Consumer/enterprise assistant; outage today. | Assistant | Closed | OpenAI | unknown on X | Very high as infra, not harness. | fact: down 2026-09-03. | unknown. | Best: chat. Worst: coding-agent users bounce to Codex CLI. |
| 72 | Claude.ai | Consumer/enterprise assistant. | Assistant | Closed | Anthropic | unknown | Very high. | fact: login/API/Code outage. | unknown. | Best: chat+code. Worst: capacity. |
| 73 | Gemini | Google assistant. | Assistant | Closed | High. 3.8 Flash cost-per-task debate. | ToS/account-ban adjacency. | AA cost/task up 40% same sticker price. | Best: cheap Flash. Worst: account blast radius. | ||
| 74 | Perplexity | Search assistant. | Assistant | Closed | Perplexity | unknown | Low in coding-agent lane. | unknown. | unknown. | Best: research. Worst: not a coding harness. |
| 75 | M365 Copilot | Office agent. | Enterprise assistant | Closed | Microsoft | Microsoft | Medium office, low coding. | fact: 2026-08-31 M365 auth incident cited in one roundup. | unknown. | Best: Outlook/Teams. Worst: not SWE. |
| 76 | AQ | “Harness of harnesses” — multiplayer over Claude Code/Codex/OpenCode/Cursor/Devin. | Meta-harness | unknown | @aqdotdev | unknown | Low-medium. Launch 2026-09-02. | unknown. | unknown. | Best: team + many harnesses. Worst: brand-new. |
| 77 | Caveman | Token-saving “talk like caveman” skill. | Agent skill | Open | JuliusBrussee (list URL) | unknown | Low. Skill lists, not platform. | unknown. | unknown. | Best: fewer tokens. Worst: not the 100k-star platform in this window. |
| 78 | Tabby | Self-hosted completion. | Coding assistant | Open | TabbyML | unknown | Low. OSS lists. | unknown. | unknown. | Best: self-host complete. Worst: agent era passed it. |
| 79 | Herdr | CLI to manage multiple coding-agent sessions. | Orchestrator | Open | herdrdev | unknown | Low-medium. Claude/Hermes/Grok Bot plugins. | unknown. | unknown. | Best: many terminals. Worst: another layer. |
| 80 | oh-my-hermes | Skills/routing pack for Hermes Agent. | Harness pack | Open | @rlaope | unknown | Low-medium. Author-heavy. | unknown. | unknown. | Best: Hermes coding power-up. Worst: one-maintainer. |
| 81 | LiteLLM | Model gateway. | Gateway | Open | BerriAI | unknown | Low this window except security roundups. | Named next to Langflow RCE targets. | unknown. | Best: campus router. Worst: attack surface. |
| 82 | MCP (protocol) | Tool protocol under AAIF/LF. | Protocol | Open | AAIF / Anthropic origin | foundation | High as plumbing. | A2A also under AAIF (2026-08-20 cited). | unknown. | Best: common tools. Worst: not a product. |
| 83 | Kimi Code | Moonshot coding agent; Paperclip adapter. | Coding agent | Open claimed | Moonshot | unknown | Low-medium. | Paperclip adapter fact. | unknown. | Best: Kimi quota. Worst: thin global X. |
| 84 | Zite | App builder that uses your Claude/ChatGPT/Cursor sub. | App builder | Closed | Zite | unknown | Low. Launch today vs Lovable/Replit double-pay. | unknown. | unknown. | Best: no extra credits. Worst: new. |
| 85 | Funes (HF) | Session memory across Claude Code/Codex/Pi/Hermes. | Memory layer | unknown | Hugging Face | NVIDIA deal today | Low. One Chinese explainer. | unknown. | unknown. | Best: don’t re-archaeology sessions. Worst: new; HF acquisition noise. |
Source: Grok Build X scan, work id uvu-x-harness-trends-v2, September 3, 2026 — 85 tools, source log of hits and misses in the working papers. Quotes are limited to 25 words; a miss is not proof of absence.
How this sweep repeats
The method is written down step by step — discovery sweep, repository snapshot, adoption signals, funding and ownership, stability and controls, benchmark handling — so it can be re-run monthly and the deltas tracked. It lives with the plan's working papers; the standing rule is that any landscape question gets the trending, funding, growth, and social-signal dimensions by default, never adoption alone.
Survey: working paper 35, September 3, 2026 (171 entries, source log with hits and misses). GitHub data: API snapshot, same day. X signal: Grok Build scan, same day (its first run died on the provider's capacity errors after forty minutes; the rerun wrote as it went and completed).
Adoption, audited
Appendix N in five lines
- QuestionHow can Utah Valley University get teachers to use artificial intelligence at low cost and prove it helped?
- AnswerThe program has the right kinds of support, but it must test finished work and repeat use instead of counting sign-ups or classes.
- Deciding numbers$20,000 first waveest110–150 verified adoptersest$133–$182 per verified adopterest
- What the plan doesChange the current program to run the cross-campus first wave, pay peer helpers and part-time teachers, start at course setup, and defer the coach bot.
- Still unknownNo local test has run: the cost per verified adopter, the 30-day repeat rate, and adjunct parity are all estimates until the first wave measures them.
The sponsor's first worry is that professors won't use it; the university's standing constraint is money. So the adoption program in §07 was put through a second, adversarial review: a fresh reader audited it source by source against Influencer's six sources of influence, then went looking for cases across public health, hospitals, government, agriculture, schools, workplaces, consumer products, and higher education where mass adoption was achieved at low cost — and graded each one for evidence quality (A causal or systematic review, B strong but non-causal, C descriptive or vendor, D anecdote). The verdicts below changed the guide. Every claim keeps its grade: fact est unknown.
The short version
- fact The current plan covers all six sources and is unusually strong on faculty choice, adjunct access, real tasks, and privacy.
- fact Its weak point is proof: activation, training, artifact completion, repeat use, teaching use, and learning results are often treated as if they were the same outcome.
- fact The strongest evidence favors action placed inside normal work, hands-on practice, local opinion leaders, fast help, and timely prompts.
- fact Messages and social norms usually produce modest gains; mandatory checklists can reach near-universal reported compliance without changing outcomes.
- est Change the first vital behavior from “finish a demo” to “apply, check, and use or reject one result within seven days.”
- est Replace “repeat weekly” with “repeat at the next natural occurrence within 30 days.” Faculty work is often episodic.
- est Make course-shell creation, not the first week of class, the main adoption moment.
- est Add a peer behavior: each nominated catalyst supports several colleagues and shares both a useful result and a rejected result.
- fact The lowest directly relevant faculty-AI payment with a measured completion result found here was $500; its causal effect is still unknown.
- est Use a $20,000 first wave to learn UVU’s real cost per verified adopter, then scale only if the result, trust, and adjunct-parity gates pass.
- est STOP making the custom coach bot a launch dependency. Start with three templates, existing Power Hours, and human help.
- unknown No public evidence supports a promise of one-term majority faculty adoption at UVU for $0, $20,000, or $60,000.
The audit — source by source
| Source of influence | What the plan does | What it assumes | Where the evidence is thin |
|---|---|---|---|
| Personal motivation | plan Uses a real faculty task, discipline examples, skeptic stories, integrity framing, personal choice, and a no-AI route. | est Immediate usefulness and a credible failure example will overcome risk, workload, and identity concerns. | fact Kantar’s 85% daily-use claim is vendor evidence. unknown Ithaka does not establish that discipline workshops draw more faculty than generic workshops. The cited concern percentages do not prove that the proposed framing changes behavior. |
| Personal ability | plan Teaches a short draft-ground-check cycle, gives starter tasks, provides feedback, and proposes a coach bot. | est A short session teaches transferable verification skills and the bot will be safe, accurate, and easier than current support. | unknown The bot has no adoption precedent. Completing a practice artifact does not prove the faculty member used it in real work. Temple’s Canvas transition was hands-on but eventually mandatory and lacks a causal comparison. |
| Social motivation | plan Uses confidential peer nomination, includes skeptics and adjuncts, publishes only truthful norms, and avoids selecting champions solely by enthusiasm. | est The nomination questions find task-specific influence and visible participation will not feel coercive. | fact Opinion-leader trials support the broad idea, with a median 10.8-point practice gain, but do not establish the best leader-selection method or its cost. Social-comparison messages have also produced null and boomerang effects. |
| Social ability | plan Uses micro-cohorts, Power Hours, adjunct-friendly formats, a shared artifact bank, and human escalation. | est Existing staff and peers can supply fast, discipline-specific help across seven colleges. | unknown There is no UVU capacity model for review, accessibility, curation, or one-business-day help. An adjunct-first program can become unpaid adjunct labor if participation happens outside compensated work. |
| Structural motivation | plan Pays for a completed artifact, proposes prompt payment, recognizes useful contributions, and rejects login-based rewards. | est A $500 award produces adoption and repeat use rather than only course completion. | fact CSUB had 37 of 43 spring participants complete a $500 course, but there was no unpaid comparison. Hawai‘i awarded 50 $1,000 incentives but has not published a completion rate. OER programs show that faculty development work can greatly exceed the value of a small stipend. |
| Structural ability | plan Proposes Canvas placement, one-click sign-on, three starter tasks, safe defaults, local workflow saving, and help where faculty already work. | est Canvas changes, identity integration, privacy review, accessibility work, and maintenance will have little incremental cost. | unknown None of those UVU systems were inspected. Forced LMS rollouts do not prove voluntary AI adoption. A removable policy block is reasonable, but any default affecting teaching needs prior faculty-governance approval. |
The vital behaviors — keep, change, add
- V1 — Modify.
- V2 — Tighten.
- V3 — Change the cadence.
- V4 — Add a peer diffusion behavior.
The moments that matter — re-timed
| Moment | Judgment |
|---|---|
| Course-shell creation or copy, 4–6 weeks before term | est Make this the main moment. Faculty can still change assignments, the syllabus, and course structure. Place the policy block, three task cards, and support link here. |
| Adjunct contract and onboarding | est Correct high-reach moment only if it happens before syllabus deadlines, is asynchronous, and is paid when outside normal duties. Late onboarding is a poor learning window. |
| 7–14 days before the first graded assignment | est Better than the day the assignment opens. Prompt the assignment-level rule and one safe-use or non-use example. |
| First failed or uncertain attempt | est Offer human help within the person’s work window. A bad first result can end adoption unless recovery is easy. |
| Next recurrence within 30 days | est This is the repeat-use test. The cue comes from the work, not from a platform streak. |
| First week of term | est Use only for confirmation and student communication. It is too late and too busy to be the principal faculty-adoption moment. |
| Post-term/course-copy period | est Ask faculty to reuse, revise, or retire the workflow and optionally share the artifact. |
The devil's-advocate case against the plan
- est The plan may recruit the same 120 enthusiasts repeatedly while the overloaded middle remains untouched.
- fact Its strongest examples usually prove seats, course completion, or activity—not sustained faculty practice or better student learning.
- est A custom bot and full Canvas integration could turn a cheap behavior program into an expensive software project.
- est “Voluntary” use can still feel mandatory when senior leaders, default course blocks, public norms, and stipends all point one way.
- est Discipline-specific support is valuable but becomes costly if every department creates and maintains separate material.
- unknown The plan has no tested method for moving from a small paid cohort to a majority of 2,112 instructors.
- unknown Board-reported employee AI use rising from 61% to 76% has no published denominator, method, faculty-only result, or causal tie to this service.
- fact The CSU wording needs correction: the public record says at least 250,000 activations by spring 2026, not “at most half.” The 0.7% figure is student training completion; the reported faculty figure was 16%. Neither is repeated use.
Cross-domain cases — what moved adoption, at what cost, with what evidence
Codes for the sources of influence: PM personal motivation · PA personal ability · SM social motivation · SA social ability · XM structural motivation · XA structural ability. The mappings are our reading, not claims made by the studies.
| Case | Domain | What moved adoption | Sources | Cost | Result | Grade | Source · date |
|---|---|---|---|---|---|---|---|
| Delancey Street | Residential rehabilitation | fact Long immersion, resident-run work and education, shared responsibility, and peer teaching. | est All six | unknown No complete economic cost or cost/graduate; resident labor and businesses are real inputs. | fact Foundation reports more than 18,000 graduates; entrant count, attrition, comparison group, and independently reproducible outcome are absent. | D | Delancey Street, current institutional account |
| Guinea-worm campaign | Public health | fact Village volunteers, filters, water treatment, containment, surveillance, and reporting rewards. | est All six | unknown Forty-year total and cost/adopter not published. | fact Estimated human cases fell from 3.5 million in 1986 to 10 reported in 2025. | B; bundled long-run trend | WHO, 1986–2025 |
| Thailand 100% Condom Program | Public health | fact Uniform establishment policy, free supply, peer coordination, STI services, media, monitoring, and sanctions. | est All six | unknown | fact Commercial-sex condom use rose from 14% in early 1989 to over 90% from 1992; surveillance also recorded a large STI decline. | B; no causal isolation | Program review, 1989 onward |
| Geneva hand hygiene | Hospital safety | fact Bedside alcohol rub, posters, training, observation, feedback, and leadership. | est PM, PA, SM, SA, XA | fact Crude three-year cost under SFr380,000; cost/new compliant worker unknown. | fact More than 20,000 observed opportunities; compliance rose 48%→66%, while infections fell 16.9%→9.9%. | B; uncontrolled bundle | Lancet record, 2000 |
| WHO surgical checklist | Hospital safety | fact Oral team pause, local adaptation, training, champions, and feedback. | est PA, SM, SA, XM, XA | unknown Original eight-site cost. | fact Complications fell 11%→7% and death 1.5%→0.8% among 7,688 before/after patients. | B | NEJM, 2009 |
| Ontario checklist rollout | Hospital safety; important null | fact Required checklist use and reporting, without the same intensive team implementation. | est PA, SM, XM, XA | unknown | fact Across 101 hospitals, adjusted death and complication changes were not significant despite very high reported checklist use. | B, null | NEJM, 2014 |
| UK tax letters | Public administration | fact One truthful local social-norm sentence added to an existing collection letter. | est SM, XA | fact The text change had no reported added implementation cost; total program cost unknown. | fact In a cleaner 1,400-person trial, payment rose 38.7%→45.5%. | A | BIT debt trials, 2011–12 |
| UK organ-donor prompt | Public administration | fact A short prompt appeared immediately after an online transaction. | est PM, SM, XA | unknown | fact In 1,085,322 allocations, reciprocity increased registration from 2.3% to 3.1%; a norm-plus-photo version underperformed the control. | A with allocation caveat | Trial report, trial 2013; published 2018 |
| Rajasthan immunization | Vaccination | fact Reliable local camps plus small food incentives that offset travel and time. | est All six | fact Corrected cost per fully immunized child: $27.94 with incentives versus $55.83 with reliable camps alone. | fact Full immunization: 39% incentive, 18% reliable-camp-only, 6% control; 134 villages and 1,640 children. | A | BMJ, 2010; cost correction, 2016 |
| Rutgers opt-out appointments | Vaccination | fact Employees received a prebooked vaccination appointment they could freely change or cancel. | est XA | unknown Incremental scheduling cost. | fact Vaccination was 45% with opt-out scheduling versus 33% with opt-in scheduling among 478 employees. | A | JAMA, 2010 |
| Local clinical opinion leaders | Professional practice | fact Locally recognized clinicians educated or influenced peers. | est PA, SM, SA | unknown No included study reported cost-effectiveness. | fact Twenty-four randomized studies; median adjusted practice-compliance gain 10.8 points across 18 studies. | A, systematic review | Cochrane, 2019 |
| Haryana “information spreader” nominations | Vaccination/network diffusion | fact Residents nominated people good at spreading information; nominees received calls and texts about camps. | est SM, SA, XA | unknown | fact In 521 villages, nominated spreaders produced about 4.9 more vaccinated children per village-month than random seeds’ 18.11. General “trust” nominations were not clearly better. | A | Review of Economic Studies, 2019 |
| Malawi lead farmers | Agricultural extension | fact Two network-positioned farmers learned and demonstrated pit planting; multiple exposures mattered. | est PA, SM, SA, XM, XA | fact $8 in-kind gift per seed farmer; network census and downstream cost unknown. | fact In 200 villages and roughly 5,600 households, non-seed adoption gained 3.6 points over a 3.8% benchmark in year two. | A | American Economic Review, 2021 |
| Rogers and “trigger the middle” | Diffusion theory | fact Relative advantage, compatibility, trialability, visibility, networks, and organizational readiness help explain diffusion. | est All six | unknown | unknown No universal tipping percentage or rule says a fixed “middle” segment will trigger mass adoption. Field trials support task-specific spreaders and repeated exposure, not a magic adopter share. | C as a framework | Rogers; health diffusion review |
| George Mason Canvas | Higher-ed technology | fact Opt-in migration, 20 faculty mentors, course help, training, office hours, and later mandatory cutover. | est PA, SM, SA, XM, XA | fact Mentor stipend $2,000/semester; total cost unknown. | fact 892 instructors and 19% of sections used Canvas in fall 2024; 1,502 instructors and 38% in spring 2025. | B | Official final report, 2025 |
| AUT and Temple Canvas | Higher-ed technology | fact Templates, champions, hands-on course building, drop-ins, communities, and migration help. | est PA, SM, SA, XA | unknown | fact AUT moved 1,837 courses; Temple reported more than 1,000 pilot participants and later majority use, without a denominator. Both transitions were headed toward required institutional use. | B/C | AUT, 2023; Temple vendor case, rollout began 2017 |
| Google Classroom in 2020 | School technology | fact No-charge core service embedded with existing Google tools during emergency remote schooling. | est PM, PA, XA | unknown Devices, support, and rollout cost. | fact Google reported 50 million students and educators in March 2020, 100 million by January 2021, and 150 million by May 2021. No teacher-only or repeat-use denominator. | C | Google, Jan. 2021; Google, May 2021 |
| Turnitin and clickers | Teaching technology | fact Common systems, training, one-to-one help, question design, and peer discussion. | est PA, SA, XA | unknown | fact One Turnitin case had 454 accounts but only 24 of 47 departments using it. At Colorado, 70 clicker faculty—3% of faculty—reached 44% of undergraduates through large courses. | B | Turnitin case, 2010–13; clicker study, 2007 |
| OER programs | Higher-ed teaching practice | fact Release time, grants, library/design help, reusable resources, and no-cost course tags. | est PA, SA, XM, XA | fact One multi-college review valued average course development at about 180 hours and $12,600, versus an average $1,500 stipend or release award. | fact Affordable Learning Georgia reports more than $143 million in student savings over 1.1 million enrollments; causal faculty-adoption effect remains unknown. | B | ALG 2022 report; SRI OER evaluation, 2022 |
| GitHub Copilot at Accenture | Developer tools | fact Suggestions appeared inside the IDE, with almost immediate first value. | est PM, PA, XA | unknown | fact GitHub reported 81.4% same-day installation and 96% same-day first acceptance among installers; one acceptance is activation, not useful adoption. | C | GitHub/Accenture, 2024 |
| Embedded customer-service AI | Workplace AI | fact Three-hour onboarding, suggestions inside live chats, scheduled access, and ordinary coaching. | est PA, SA, XA | unknown | fact In 5,172 agents and more than 3 million chats, access raised resolved issues/hour by 15%; effects were much larger for less-skilled workers. | A−, staggered rollout | Quarterly Journal of Economics, 2025 |
| UK government Copilot | Workplace AI | fact Office integration, short training, workshops, central resources, and department support. | est PA, SA, XA | unknown | fact Of 20,000 licenses, “active” meant one interaction in 30 days; adoption reached 83% and remained near 80%. Application use varied widely and some use fell after its peak. | B | GOV.UK report, 2025 |
| Slack and enterprise InnerSource | Workplace/open source | fact Free trialability, familiar workflow placement, integrations, open backlogs, contribution guides, maintainers, and peer help. | est PA, SM, SA, XA | unknown Free licenses did not remove support, maintenance, or contribution costs. | fact Slack reported more than 10 million daily users in 2019. Ericsson reported over 1,000 internal reuse instances, but also few early contributions where work time was not funded. | B/C | Slack S-1, 2019; Ericsson InnerSource study, 2024 |
| Duolingo, Strava, Peloton | Consumer habits | fact Private progress, reminders, finite challenges, instructors, streaks, and social cues. | est PM, SM, SA, XM, XA | unknown Cost/adopter and causal contribution of each feature. | fact These services report high repeat activity, but their users are self-selected and most evidence does not show that streaks or rankings caused retention. A student RCT found personalized reminders improved first use more than streak messages. | A for reminder experiment; B/C products | Reminder RCT, 2026; Duolingo SEC filing; Peloton filing |
| California State University | Higher-ed AI | fact Systemwide licenses, training, local projects, and $3 million for faculty proposals. | est PA, SA, XM, XA | fact Initial contract $17 million/18 months; reported renewal $13 million/year. | fact More than 93,000 activations by June 2025 and at least 250,000 by spring 2026. Voluntary training completion was reported as 0.7% of students and 16% of faculty. Activation and training are different measures. | B/C | CSU launch, 2025; CalMatters, 2026 |
| CSUB and Hawai‘i | Higher-ed AI incentives | fact Payment was attached to bounded training or an implemented, shared assignment. | est PM, PA, SA, XM | fact CSUB: $500/completer. Hawai‘i: 50 awards of $1,000; actual payout unknown. | fact CSUB spring 2026 completion was 37/43, with a waitlist. Hawai‘i has no published cohort completion rate. Neither isolates the incentive’s effect. | B/C | CSUB Senate report, 2026; Hawai‘i program, 2025 |
| Virginia Tech, Manchester, Cedarville | Higher-ed AI use | fact Training or supported cohorts plus institutionally approved access. | est PA, SA, XA | unknown Comparable total cost/adopter. | fact Virginia Tech recorded 78% typical weekly activity among 425 measured pilot users. Manchester reported 90% 30-day adoption but did not publish its exact definition. Cedarville reported 82% activation and average weekly activity among 75% of active accounts. | B/B/C | Virginia Tech, 2025; Manchester, 2026; Cedarville, 2026 |
| ASU, Michigan, Arizona | Higher-ed AI | fact Proposal challenges, locally built tools, Canvas placement, workshops, and course-specific assistants. | est PM, PA, SM, SA, XA | unknown Comparable faculty cost/adopter. | fact ASU reported more than 500 completed or active projects; Michigan reported 43,800 U-M GPT users and more than 500 courses; Arizona’s AI-VERDE pilot recorded 78 users and 97,658 calls. Only Michigan published an adjacent historical course comparison. | B/C | ASU, 2025; Michigan, 2025; Arizona paper, 2025 |
| Miami Dade, Florida, UT Austin | Higher-ed AI | fact Faculty training, institution-wide curriculum, course tools, vetted activities, and local AI platforms. | est All six in varying mixes | fact Miami Dade later received a $2 million expansion grant; UF disclosed an $800,000 recurring annual QEP budget and charges $500 for external Academy enrollment; UT cost is unknown. | fact Miami Dade reports 1,056 faculty trained but its outcome claims lack denominators. UF reports 200+ AI courses and 300+ AI-focused faculty. UT Sage reported 366 tutors in 90 Canvas courses; Copilot training reached 2,400+ users. | B/C | Miami Dade; UF; UT Austin, 2024–26 |
| Ivy Tech and Middlebury | Higher-ed AI | fact Ivy Tech used workshops, departmental guides, and a 60-user pilot. Middlebury provides faculty choice, syllabus templates, workshops, consultations, and $1,000 AI mini-grants. | est PM, PA, SA, XM, XA | fact Middlebury mini-grants up to $1,000; other adoption costs unknown. | unknown Ivy Tech published no activity or outcome result. Middlebury has useful student studies but no measured faculty-rollout denominator or cost/adopter. | C | Ivy Tech account, 2024; Middlebury resources, 2025–26 |
The low-cost design for a constrained university
Behaviors
- est Apply one checked result. Finish a real, bounded task and use, edit, or reject the result within seven days.
- est Publish one clear course rule. Add an assignment-level permission, disclosure example, and checking rule before students begin.
- est Repeat at the next real occurrence. Reuse or revise the workflow within 30 days.
- est Help peers perform the behavior. Each catalyst supports three to five colleagues and shares a worked success plus a failure.
- est Retire what does not help. A faculty member may record “not useful,” “not safe,” or “not suitable” without being counted as resistant.
Moments
- est Primary: course-shell copy, 4–6 weeks before term.
- est Equity gate: adjunct contract/onboarding, with paid asynchronous completion.
- est Teaching gate: 7–14 days before the first graded assignment.
- est Recovery: immediately after the first failed or uncertain result.
- est Repeat: the next occurrence of the same task, within 30 days.
- est Reuse: post-term course-copy and artifact revision.
- est Avoid launching tools during finals or using the first week as the main training period.
People — finding the real spreaders for $0
- est Add two questions to an existing department, senate, or teaching-center pulse: Who spreads useful teaching practices? Who do you consult when a teaching tool may not be ready?
- fact Haryana’s trial suggests “information spreader” nominations may be more useful than broad trust or title.
- est Nominate at least two people per teaching cluster because complex behaviors may require more than one credible exposure.
- est Include adjuncts, skeptics, high-enrollment instructors, and people with strong peer connections—not only public AI enthusiasts.
- est Use the 120 Academy completers as a candidate pool, not as an automatic champion list.
- est This can have $0 incremental cash cost when inserted into existing processes. unknown Staff administration and faculty time still have economic cost.
Structures that cost nothing
- est Add a removable Canvas block when a new or copied shell is created: active choice among permitted, limited, or prohibited use; one example; one support link.
- est Put three discipline task cards beside the normal work: course preparation, assessment/rubric work, and student feedback.
- est Default visibility on but keep tool launch, Academy enrollment, reminders, telemetry, and public sharing opt-in.
- est Reuse Canvas Commons, existing Power Hours, approved sign-on, and existing department meetings before building a new portal.
- est Keep a small reviewed artifact library. Adapt a shared pattern rather than creating a separate program for every discipline.
- est Put human help beside the task. Do not require a faculty member to leave Canvas, search a catalog, or create another account.
- est Do not build the coach bot until support logs show a repeated problem that templates and people cannot solve cheaply.
Incentives — what the evidence supports
- fact CSUB’s $500-on-completion course produced 37 completions among 43 enrolled participants, but there was no unpaid control.
- fact Hawai‘i offered $1,000 for building, using, and sharing an assignment, but published no overall completion rate.
- fact George Mason paid Canvas mentors $2,000 per semester; its total adoption cost was not reported.
- unknown The minimum effective faculty-AI stipend is not known. No credible public test here shows that $100, $200, recognition, or early access alone changes sustained faculty behavior.
- est Pay catalysts for a reusable artifact plus peer help, not for attendance or message volume.
- est Use completion awards when the work is outside ordinary duties. Recognition, early access, and voluntary public commitment can supplement pay but must not replace compensation for adjunct labor.
- est Avoid leaderboards, streaks, prizes for frequent use, and public lists of “non-adopters.”
Measures that don't surveil
A verified adopter is a faculty member who completes a checked real-work cycle, applies or edits or rejects the result within seven days, publishes one course-use rule (or completes another approved applied workflow), and repeats a related workflow within 30 days. The funnel is measured separately:
| Measure | Privacy-preserving method |
|---|---|
| Reach | est Aggregate invitations, Canvas card views, and clinic seats. |
| First applied use | est Content-free completion receipt plus apply/edit/reject choice. |
| Course adoption | est Voluntarily submitted policy or assignment artifact count; do not scrape course content. |
| Repeat use | est Opt-in random evaluation token with no SSO/HR link and short retention. If governance rejects any longitudinal token, report repeat use UNKNOWN. |
| Peer diffusion | est Aggregate unique clinic participants and voluntarily reported referral source. |
| Quality | est Small faculty review sample using a published rubric; never feed results into evaluation or rehire. |
| Equity | est PT/FT and college participation/completion only in cells of at least 20. |
| Trust | est Anonymous usefulness, pressure, autonomy, and surveillance-concern pulse. |
| Safety | est Aggregate privacy, access, accuracy, and accessibility incidents. |
| Economics | est Cash spend and estimated staff/faculty hours divided by verified adopters; publish both. |
Cost per verified adopter at three budgets
With 2,112 instructors, a strict majority is 1,057. These are planning-capacity estimates, not adoption forecasts; they assume the AI service itself is already funded, and faculty time is not free.
| Cash budget | Proposed use | Supported capacity | Verified-adopter assumption | Cash cost per verified adopter | What it can honestly prove |
|---|---|---|---|---|---|
| $0 | est Ask each of the 120 existing Academy completers to help one new colleague during already-paid work; reuse Power Hours and templates. | est 120 new peer-support places. | est 60–84 new adopters, using a 50–70% planning completion range. | est $0 cash; full economic CPA UNKNOWN. | est Whether existing social capacity produces a measurable second wave. It is not a credible one-term majority plan. |
| ~$20k | est Ten catalysts × $1,000 for an artifact and two clinics; twenty adjunct/faculty artifact awards × $500. | est Ten catalysts plus 200 unique clinic places. | est 110–150 verified adopters: ten catalysts plus 50–70% of clinic places. | est $133–$182. | est A cross-campus first wave large enough to compare colleges, PT/FT participation, repeat use, and local CPA. |
| ~$60k | est Thirty catalysts × $1,000; sixty artifact awards × $500. | est Thirty catalysts plus 600 unique clinic places. | est 330–450 verified adopters. | est $133–$182. | est A material campus wave—about 16–21% of instructors—but not a one-term majority. |
At the modeled $133–$182 per verified adopter, reaching a majority directly would take roughly $141,000–$192,000; peer diffusion could lower later waves, but that has to be measured, not assumed. Recommendation carried into the guide: run the $20,000 first wave, and continue to the $60,000 design only if cost per verified adopter stays at or below $200, 30-day repeat is credible, adjunct completion is not materially lower than full-time, reported pressure or surveillance concern is not rising, and artifacts pass faculty quality review. est
Ranked changes — what the guide now does differently
| # | Change | Evidence and reason | Incremental cash |
|---|---|---|---|
| 1 | Modify V1 to require applied use or reasoned rejection within seven days. | fact Checklist and account cases show that formal completion can exist without faithful practice or useful results. | est $0 |
| 2 | STOP making the custom coach bot a launch dependency. | unknown No precedent here shows that a bot improves faculty adoption. fact Michigan’s locally built tools show possible scale, but their institutional cost is unknown. | est Saves or defers build cost |
| 3 | Replace weekly repetition with the next natural recurrence within 30 days. | fact Embedded workplace tools work at the task; faculty work is not uniformly daily or weekly. | est $0 |
| 4 | Move the main trigger to course-shell copy and assignment construction. | fact BIT, vaccination, and teacher-nudge evidence supports prompts at a natural action point. | est $0 cash if the Canvas template path already exists; staff time unknown |
| 5 | Add the catalyst peer behavior and task-specific nomination questions. | fact Opinion-leader review: median 10.8-point practice gain. Haryana: information-spreader nominations outperformed random seeds. | est $0 for nomination; paid catalyst budget as selected |
| 6 | Use the $20,000 budget for catalysts, adjunct artifacts, and clinics—not 30 isolated $500 completions. | est The same cash creates more supported opportunities and reusable capacity. | est Budget-neutral reallocation |
| 7 | Separate policy clarity from AI adoption. | est A clear prohibition can satisfy V2’s current wording without any adoption. Record clarity, safe AI use, and justified non-use separately. | est $0 |
| 8 | Run a stepped or randomized invitation test. | fact The strongest cheap-adoption evidence comes from randomized rollout. est Compare early versus later invitation at department or course-cluster level. | est Near-zero cash if built into rollout |
| 9 | Resolve the telemetry contradiction. | plan The bot is supposed to beat self-service on 30-day repeat, while the current hard rule forbids durable identification. est Use opt-in, pseudonymous, short-retention evaluation tokens—or report repeat use UNKNOWN. | unknown Governance and evaluation time |
| 10 | Correct the evidence claims before leadership use. | fact CSU’s “at most half” is unsupported; its 0.7% number concerns students, not faculty. Hawai‘i announced awards, not verified payouts. Manchester did not prove cohort support caused its 90% figure. | est $0 |
| 11 | Make adjunct compensation a launch gate. | fact OER evidence shows major hidden faculty workload. est Do not call unpaid extra work “free adoption.” | est Included in the $20k/$60k designs |
| 12 | Retire nudges that show no incremental effect. | fact Large education and social-comparison nudge trials show that low-cost messages can be null or backfire. | est Saves staff attention |
Risks specific to influence programs — and how the cases handled them
| Risk | How it fails | What the cases show | UVU control |
|---|---|---|---|
| Manipulation perception | Defaults, norms, and “trusted peers” can look like concealed pressure. | fact The UK organ-donor norm/photo prompt underperformed control. | est Disclose the program owner, purpose, options, and measure. Keep refusal easy. Ask anonymously whether people felt pressured. |
| Faculty autonomy and senate reaction | Central purchasing or template defaults may be read as a teaching mandate. | fact CSU’s rollout produced reported consultation and governance objections. | est Faculty Senate approves the choice architecture before launch. Faculty actively choose permitted, limited, or prohibited use. |
| Union concerns | New required work may be added without workload recognition or bargaining. | fact Large OER projects found substantial hidden development time. | est Distinguish voluntary exploration from required work. Bargain or approve workload terms before requiring training or artifacts. |
| Adjunct exploitation | Evening clinics and artifact production become uncompensated labor tied to perceived rehire risk. | fact Small stipends often cover only a fraction of course-development work. | est State pay, time estimate, ownership terms, and no-rehire consequence before enrollment. Offer asynchronous paid routes. |
| Surveillance fears | Usage logs become a list of “resistant” faculty or enter evaluation. | fact Workplace and university platforms can expose named activity even when only aggregate reporting is promised. | est Separate evaluation from SSO, HR, tenure, discipline, and rehire. Store no prompts or outputs. Suppress cells under 20. |
| AI-mandate backlash | Access, default visibility, repeated messages, and public norms combine into a soft mandate. | fact Ontario showed that mandatory reported compliance did not guarantee outcomes. | est Default visibility only. Keep participation, reminders, telemetry, testimonials, and public sharing opt-in. |
| Norm boomerang | Low participation normalizes non-use; high-user comparisons discourage or shame others. | fact Teacher social-comparison evidence found no overall gain and possible downward movement among above-average groups. | est Use a norm only when true, local, recent, clearly defined, and helpful. Otherwise omit it. |
| Incentive gaming or crowd-out | Participants optimize for a certificate or payout rather than useful practice. | fact CSUB proves completion, not repeat use. | est Pay for a reviewed artifact and bounded peer contribution; measure later repeat separately. Do not pay for prompts or logins. |
| Champion burnout or elite capture | The same visible enthusiasts receive every role and central support becomes a bottleneck. | fact InnerSource evidence reports bottlenecks and weak contribution when work time is not funded. | est Cap cohort size, pay defined work, rotate roles, include skeptics, and publish help capacity. |
| Accuracy, privacy, and safety | A smooth first experience can normalize unchecked or sensitive use. | fact UK Copilot users reported limits with nuanced, complex, and sensitive work. | est Require data classification, verification, and a safe exit. Keep clinical and other high-risk tasks in separate reviewed lanes. |
| Gamification cringe | Streaks, badges, and rankings feel infantilizing or punitive. | fact Consumer products show repeat activity, but seldom prove that public rankings caused it. | est Transfer only private progress, self-chosen reminders, finite tasks, and a lapse-recovery path. No public streaks or leaderboards. |
| Equity and accessibility | Full-time faculty capture the support while adjuncts and disabled faculty face time or access barriers. | fact Workplace AI reports accessibility benefits, but also learning curves and role differences. | est Test accessibility before launch; provide asynchronous and human routes; report PT/FT parity only in protected aggregate cells. |
Review: working paper 37, September 3, 2026 — sources dated and graded; nothing here contacted UVU faculty, inspected UVU systems, or accessed person-level data. The three budget ranges are planning assumptions; only a local staged test turns them into evidence.
The two Mac lineups, side by side
Appendix O in five lines
- QuestionShould the plan buy the new Mac computers or the older line?
- AnswerBuy the new line by default because matched old stock is not proved and its lower price does not fully pay for its slower speed.
- Deciding numbersSeptember 22, 2026 availabilityfact14.4% more costest10% more total answer-writing speedest
- What the plan doesPrice the new computers, test four after September 22, and buy old ones only with a written quote for the exact set and date.
- Still unknownNew desktop tests, university prices, stock for 4 or 26 matching old units, and exact warranty prices are NOT_RUN.
A fair question from the first read-through: the guide prices the machines on Apple's newest lineup, announced August 25, 2026 — so what did the generation before it cost, what does it cost today, and how much faster should the new one really be? This appendix answers that with dated sources, every number graded. Research cutoff September 3, 2026; prices are US dollars before tax. fact means a dated source states or displays it; est means we derived it and show the basis; unknown means the proof was not available.
What this changes in the plan
1. The plan stays priced on the newest lineup. The previous machines are no longer sold new — Apple's own store and its education store have dropped them, and the big education resellers show backorders. Single refurbished units exist at prices near or above what they cost new, and nobody can prove a matched set of 4 or 26 is available until it is reserved. 2. One price corrected. The 256GB Ultra needs Apple's top chip, so it is about $10,799 retail (about $9,900 education, an estimate), not the $9,499 an earlier pass carried; it only affects the later, optional heavyweight node, and the simulator and configurator now use the corrected figure. 3. Priced both ways everywhere. The budget simulator has a generation switch and the configurator shows the previous generation beside every configuration you build, using the refurbished prices below and speed credits deliberately below Apple's claims (+10% per box for the M5 Pro and M5 Max, +40% for the M5 Ultra). 4. An acceptance test before scale. A four-unit run after September 22 must re-prove the speed credit before any purchase beyond the validation. 5. Old units only against paper. Procurement may take previous-generation units only with a written quote for the exact quantity, drive size, and delivery date (Apple Education/NASPO Utah, CDW-G, or SHI). est
The bottom line
- fact Apple is taking orders for the new Macs; normal availability starts September 22, 2026. The 512GB M5 Ultra is due in late October.
- unknown No new Mac mini or Mac Studio serving benchmark can exist before delivery. Laptop M5 Pro/Max results are same-chip proxies only.
- est For large-model decode, budget +10% throughput for M5 Pro/Max and +40% for M5 Ultra. These credits are intentionally below Apple’s headline AI claims.
- fact Previous models are absent from Apple’s normal retail and consumer education catalogs.
- fact Apple Refurbished currently shows individual previous-generation units, but unknown whether 4 or 26 matched units can be reserved.
- fact CDW-G lists prior units as backordered; SHI reports zero stock/backorder. Current institutional fulfillment is unknown.
- A 48GB fleet priced from Apple education costs est 14.4% more than the visible M4 Pro refurb alternative and is expected to provide 10% more aggregate decode.
- Recommendation: price the plan on the newest generation, require a four-node acceptance test after September 22, and treat old inventory only as a quote-backed fallback.
The newest lineup — every configuration, priced
Apple’s Mac mini announcement and Mac Studio announcement give the starting prices and dates. Memory and networking come from the current Mac mini specifications and Mac Studio specifications.
Minimum storage is 256GB for M6, 512GB for M5 Pro/Max, and 1TB for M5 Ultra unless noted.
| Family | CPU/GPU tier | Memory | Bandwidth | Retail price | Education-store price | Availability | 10GbE | AppleCare+ premium |
|---|---|---|---|---|---|---|---|---|
| Mac mini M6 | 12/12 | 16GB | fact 153 GB/s | fact $899 | fact $799 | fact 2026-09-22 | fact +$100 retail; +$90 edu | unknown |
| Mac mini M6 | 12/12 | 24GB | fact 170 GB/s | fact $1,099 | unknown | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M6 | 12/12 | 32GB | fact 170 GB/s | est $1,299 = $899 + $400 memory | unknown | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M5 Pro | 15/16 | 24GB | fact 307 GB/s | fact $1,699 | fact $1,599 | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M5 Pro | 15/16 | 48GB | fact 307 GB/s | est $2,299 = $1,699 + $600 | est $2,139 = $1,599 + $540 | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M5 Pro | 15/16 | 64GB | fact 307 GB/s | est $2,699 = $1,699 + $1,000 | est $2,499 = $1,599 + $900 | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M5 Pro | 18/20 | 24GB | fact 307 GB/s | est $1,899 = $1,699 + $200 chip | est $1,779 = $1,599 + $180 | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M5 Pro | 18/20 | 48GB | fact 307 GB/s | est $2,499 | est $2,319 | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac mini M5 Pro | 18/20 | 64GB | fact 307 GB/s | est $2,899 | est $2,679 | fact 2026-09-22 | fact +$100; +$90 | unknown |
| Mac Studio M5 Max | 18/32 | 36GB | fact 460 GB/s | fact $2,499 | fact $2,299 | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Max | 18/40 | 48GB | fact 614 GB/s | est $3,099 | est $2,839 | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Max | 18/40 | 64GB | fact 614 GB/s | est $3,499 | est $3,199 | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Max | 18/40 | 128GB | fact 614 GB/s | est $5,099 at 512GB SSD; fact $5,399 at 1TB | est $4,639 at 512GB SSD | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Ultra | 30/64 | 96GB | fact 1.2 TB/s | fact $5,499 | fact $5,099 | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Ultra | 30/64 | 256GB | fact not configurable | N/A | N/A | N/A | fact included | N/A |
| Mac Studio M5 Ultra | 30/64 | 512GB | fact not configurable | N/A | N/A | N/A | fact included | N/A |
| Mac Studio M5 Ultra | 36/80 | 96GB | fact 1.2 TB/s | est $6,799 = $5,499 + $1,300 chip | unknown | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Ultra | 36/80 | 256GB | fact 1.2 TB/s | est $10,799 at 1TB; fact $11,299 at 2TB | unknown | fact 2026-09-22 | fact included | unknown |
| Mac Studio M5 Ultra | 36/80 | 512GB | fact 1.2 TB/s | unknown | unknown | fact late October 2026 | fact included | unknown |
Price notes:
- est values add Apple’s displayed option deltas to a sourced base price. They are not institutional quotes.
- AppleCare+ eligibility is displayed, but an exact new-desktop premium was not exposed in the retrievable store pages. Apple says education buyers may save up to 10%; the actual premium remains unknown.
- Configured-to-order delivery can be later than the family’s September 22 availability date.
The previous lineup — what it cost, and whether it can still be bought
Launch facts come from Apple’s 2024 Mac mini announcement and 2025 Mac Studio announcement.
| Previous family | Chip tier(s) | Memory | Bandwidth | Launch price: retail / education | Apple retail and consumer edu today | Apple Refurbished seen 2026-09-03 | Institutional/reseller state |
|---|---|---|---|---|---|---|---|
| Mac mini M4 | 10-core CPU/GPU | 16GB | fact 120 GB/s | fact $599 / $499 | fact not listed | fact $679, 16GB/256GB | Older Apple institution list: fact $899 for 16GB/512GB; present fulfillment unknown |
| Mac mini M4 | 10-core CPU/GPU | 24GB | fact 120 GB/s | unknown exact CTO; family starts $599/$499 | fact not listed | fact $1,019, 24GB/512GB | Older institution list: fact $1,099 for 24GB/512GB; fulfillment unknown |
| Mac mini M4 | 10-core CPU/GPU | 32GB | fact 120 GB/s | unknown exact CTO | fact not listed | fact $1,269, 32GB/512GB/10GbE | Current large-quantity price unknown |
| Mac mini M4 Pro | 12/16 or 14/20 | 24GB | fact 273 GB/s | Base 12/16: fact $1,399 / $1,299; top-tier exact unknown | fact not listed | fact $2,889 with 4TB; CDW-G base listing $1,599 backordered | Older institution list: fact $1,499 base; current fulfillment unknown |
| Mac mini M4 Pro | 12/16 or 14/20 | 48GB | fact 273 GB/s | unknown exact CTO | fact not listed | fact $1,949, 12/16, 48GB/512GB/10GbE | No four- or 26-unit stock proof |
| Mac mini M4 Pro | 12/16 or 14/20 | 64GB | fact 273 GB/s | unknown exact CTO | fact not listed | fact $2,379, 14/20, 64GB/512GB/1GbE | No four- or 26-unit stock proof |
| Mac Studio M4 Max | 14/32 | 36GB | fact 410 GB/s | fact $1,999 / $1,799 | fact not listed | fact $1,949, 36GB/512GB | CDW-G: fact $2,499, backordered; older institution list $2,299; fulfillment unknown |
| Mac Studio M4 Max | 16/40 | 48GB | fact 546 GB/s | unknown exact CTO | fact not listed | fact $2,459, 48GB/512GB | SHI stock zero/backorder; quote unknown |
| Mac Studio M4 Max | 16/40 | 64GB | fact 546 GB/s | unknown exact CTO | fact not listed | fact $6,029 with 8TB SSD | Quantity and useful-storage price unknown |
| Mac Studio M4 Max | 16/40 | 128GB | fact 546 GB/s | unknown exact CTO | fact not listed | fact $4,159, 128GB/512GB | Quantity unknown |
| Mac Studio M3 Ultra | 28/60 | 96GB | fact 819 GB/s | fact $3,999 retail; fact $3,599 institution education | fact not listed | fact $6,029 with 2TB | CDW-G: fact $5,299, backordered; older institution list $4,899 |
| Mac Studio M3 Ultra | 32/80 | 96GB | fact 819 GB/s | est $5,499 retail = $3,999 + $1,500 chip; edu unknown | fact not listed | Exact current match unknown | Fulfillment unknown |
| Mac Studio M3 Ultra | 28/60 or 32/80 | 256GB | fact 819 GB/s | Top tier: est $7,099 = $3,999 + $1,500 chip + $1,600 memory; edu unknown | fact not listed | fact $8,149, 28/60, 256GB/2TB | CDW-G 256GB listing: fact $11,639.99, backordered |
| Mac Studio M3 Ultra | 32/80 | 512GB | fact 819 GB/s | est $9,499 = $3,999 + $1,500 chip + $4,000 memory; edu unknown | fact not listed | fact $17,419, 512GB/8TB | Quantity unknown |
University purchasing status
| Channel | Current result | Usable price band | University order conclusion |
|---|---|---|---|
| Apple retail | fact previous models not listed | N/A | New prior-generation units unavailable through the ordinary public catalog |
| Apple consumer education store | fact previous models not listed | N/A | No public education checkout path found |
| Apple Certified Refurbished | fact individual Add to Bag listings | Serving-relevant examples: $1,949–$17,419 | A single unit may be purchasable; matched quantity is unknown until reserved |
| Apple institution/NASPO documents | fact old SKUs remain in pre-refresh documents | Examples: $899–$4,899 | Documents prove catalog pricing, not September acceptance, stock, or delivery |
| CDW-G | fact public listings; backordered | $1,599–$11,639.99 | Order entry may exist; delivery date and UVU contract price are unknown |
| SHI | fact search results show backorder/zero stock | Conflicting public prices | Require a written UVU quote; current fulfillment is unknown |
| Apple Education/NASPO Utah | fact Utah participation path exists | Contract-specific | Previous-generation availability and UVU price require a current quote |
Apple warns refurbished supply is limited. An “Add to Bag” page is not evidence that 4 or 26 matched computers can be delivered.
Speed evidence, and the credits we use for planning
Apple's own claims — the ceiling, not the plan
These are selective “up to” claims, not guaranteed LLM serving gains. “Peak GPU AI” refers to GPU Neural Accelerators; it is not a Neural Engine percentage.
| Upgrade | CPU claim | GPU claim | Neural Engine claim | Bandwidth change | Apple LM Studio claim |
|---|---|---|---|---|---|
| M6 vs M4 | fact “up to 40 percent faster CPU performance” | fact “up to 2x faster graphics” | fact “up to 2x faster Neural Engine” | fact 120 → 153/170 GB/s; +27.5%/+41.7% | fact up to 4.8x prompt processing |
| M5 Pro vs M4 Pro | fact up to 30% | fact up to 20%; ray tracing up to 35%; peak GPU AI over 4x | fact Apple says faster; numeric percentage unknown | fact 273 → 307 GB/s; +12.45% | fact up to 4x prompt processing |
| M5 Max vs M4 Max | fact up to 15% | fact up to 20%; ray tracing up to 30%; selected Studio GPU workload up to 50% | Numeric percentage unknown | fact 546 → 614 GB/s; +12.45% | fact up to 3.9x prompt processing |
| M5 Ultra vs M3 Ultra | fact multi-core up to 30%; single-core up to 25% | fact selected graphics up to 1.8x; generic chip comparison up to 40%; peak GPU AI 4.3–4.5x | Numeric percentage unknown | fact 819 → 1,200 GB/s; +46.52% | fact up to 4x prompt processing |
Sources: Apple’s M6/M5 Ultra chip announcement, M5 Pro/Max announcement, and the two product announcements above.
Independent measurements
PP is prompt processing or prefill. TG is token generation. Results are only comparable when model, quantization, software build, prompt, and batching match.
| Hardware and test | Model/stack | Prefill | Single-stream generation | Concurrent aggregate | Evidence state |
|---|---|---|---|---|---|
| M4 Pro 16-GPU → M5 Pro 16-GPU laptop; same llama.cpp build 8e672ef | Same Q4 llama-bench row | 364.06 → 403.19 t/s; +10.75% | 49.64 → 60.04 t/s; +20.95% | Not reported | fact same-build chip proxy |
| M4 Max 40-GPU → M5 Max 40-GPU laptop; same build | Same Q4 llama-bench row | 885.68 → 990.53; +11.84% | 83.06 → 104.03; +25.25% | Not reported | fact same-build chip proxy |
| M4 Pro mini 64GB | gpt-oss-20B MXFP4, llama.cpp build 6280 | 700.78 @2K; 618.59 @8K; 534.95 @16K; 419.74 @32K | 63.34 | Not reported | fact |
| M4 Pro mini 64GB | Qwen3.8-27B-oQ4e, oMLX 0.6.x, MTP | Generation test at 1K/4K/8K contexts | 24.9 / 20.5 / 20.8 | 120 t/s at 8-wide | fact field report |
| M4 Max 128GB | Qwen3.5-35B-A3B 4-bit, oMLX | 1,216 | 122.2 | 205.3 batch-2; 283.6 batch-4 | fact public benchmark |
| M3 Ultra 256GB, 60-GPU | gpt-oss-120B oQ8, oMLX | 868.1 | 68.0 | 94.3 / 129.3 / 152.5 at batch 2/4/8 | fact public benchmark |
| M3 Ultra 512GB, 60-GPU | gpt-oss-120B MXFP4, oMLX (verified result) | 600.7 @32k | 68.4 @1k · 57.8 @8k · 35.2 @32k | 187.1 at batch 8 (short context) | fact public benchmark — corrected 2026-09-03: the earlier row (1,405.9 / 79.2) cited a Qwen3.5-122B-A10B result |
| M3 Ultra 512GB | Qwen3.8-27B-oQ4e, oMLX 0.6.1, MTP | 420–442 | 65–75 at 8K | 259–274 at 8-wide | fact four-node field report |
| New Mac mini/Studio chassis | Any serving stack | unknown | unknown | unknown | Not deliverable until 2026-09-22+ |
The same-build llama.cpp evidence is in the project’s Apple Silicon benchmark discussion. The larger-model M4 Pro result is in the gpt-oss llama.cpp guide. The Qwen field comparison is documented in oMLX discussion #2811.
The planning credits we use
Large-model decode is commonly limited by memory bandwidth. Prefill uses more compute and can benefit more from CPU/GPU changes and software-specific accelerators.
| New tier | Decode expectation | Planning credit | Prefill expectation | Concurrent serving expectation | Reason |
|---|---|---|---|---|---|
| M6 16GB | est +20–30% | est +25% | est +20–80%; absolute speed unknown | Aggregate capacity est +25%; actual streams unknown | Bandwidth rises 27.5%; Apple’s 4.8x LM Studio claim is not treated as general serving speed |
| M6 24/32GB | est +30–40% | est +35% | est +20–100%; absolute speed unknown | Aggregate capacity est +35%; actual streams unknown | Bandwidth rises 41.7%; more RAM also permits larger models/KV caches |
| M5 Pro | est +10–15% | est +10% | est +10–20% | Aggregate capacity est +10% | Bandwidth +12.45%; same-build small-model PP +10.75%, TG +20.95% |
| M5 Max | est +10–15% | est +10% | est +10–20% | Aggregate capacity est +10% | Bandwidth +12.45%; same-build PP +11.84%, TG +25.25% |
| M5 Ultra | est +35–50% | est +40% | est +25–50% | Aggregate capacity est +40% | Bandwidth +46.52%; larger Apple GPU-AI claims are not assumed to transfer to ordinary decode |
“Concurrent capacity” means aggregate token throughput. It does not promise an equal percentage increase in integer sessions because scheduling, context length, KV cache, and latency targets still matter.
Side by side, role by role
| Serving role | Previous generation: model, memory, bandwidth, launch/current price, availability, measured serving | Newest generation: model, memory, bandwidth, price, date, expected serving | Price delta | Performance delta | University choice |
|---|---|---|---|---|---|
| Workhorse mini 48GB | M4 Pro 12/16, 48GB, 273 GB/s. Launch exact unknown; family base $1,399/$1,299. Current Apple refurb: fact $1,949, 512GB SSD and 10GbE; quantity unknown. M4 Pro 64GB proxy: Qwen3.8 20.8 t/s @8K, 120 aggregate @8-wide. | M5 Pro 15/16, 48GB, 307 GB/s. est $2,299/$2,139; with 10GbE $2,399/$2,229. Ships 9/22. est 22.9 t/s, 132 aggregate @8-wide-equivalent; actual chassis unknown. | Versus refurb with 10GbE: edu +$280, +14.4%; retail +$450, +23.1% | est +10% | Buy new by default. It gives supported supply, warranty, and a modest decode gain. Use old only with a written matched-quantity quote. |
| 64GB mini | M4 Pro 14/20, 64GB, 273 GB/s. Launch exact unknown. Current refurb: fact $2,379, 512GB, only 1GbE. Measured 64GB mini: 20.8 t/s @8K, 120 batch-8. | M5 Pro 15/16, 64GB, 307 GB/s. est $2,699/$2,499; 10GbE adds $100/$90. Ships 9/22. est 22.9 t/s, 132 batch-8-equivalent. | Without network parity: edu +5.0%; retail +13.5%. With new 10GbE: +8.8%/+17.7%, but old remains 1GbE. | est +10% decode; prefill est +10–20% | Buy new. Choose the base CPU/GPU tier for bandwidth-bound decode; pay for 18/20 only when prefill testing proves value. |
| Heavyweight Studio 128GB | M4 Max 16/40, 128GB, 546 GB/s. Launch CTO unknown. Current refurb fact $4,159. Qwen3.5-35B: 122.2 TG, 283.6 aggregate batch-4; same-build llama row 83.06 TG. | M5 Max 18/40, 128GB, 614 GB/s. est $5,099/$4,639 at 512GB SSD; ships 9/22. Planning proxy: est 134.4 TG, 312.0 batch-4 aggregate; chassis result unknown. | Edu vs refurb: +$480, +11.5%; retail: +$940, +22.6% | Fleet credit est +10%; small-model same-build evidence showed +25.25% TG | Buy new for planned deployments. Old refurb is attractive only if exact quantity and storage are confirmed. |
| 256GB Ultra | M3 Ultra up to 32/80, 256GB, 819 GB/s. Top-tier launch est $7,099, edu unknown. Current refurb: $8,149 for lower 28/60 tier with 2TB. gpt-oss-120B on 60-GPU: 868.1 PP, 68 TG, 152.5 batch-8. | M5 Ultra 36/80, 256GB, 1.2 TB/s. est $10,799 at 1TB; fact $11,299 at 2TB; edu unknown. Ships 9/22. est 95.2 TG, 213.5 batch-8 aggregate. | Like-for-like 2TB new vs lower-tier refurb: +38.7%. Top-tier launch comparison: +52.1% | est +40% | Buy new only when models exceed 128GB. The bandwidth gain is material, but obtain an institutional price and validate the exact model first. |
| 512GB Ultra | M3 Ultra 32/80, 512GB, 819 GB/s. Launch est $9,499; current 8TB refurb fact $17,419. Qwen3.8: 65–75 TG, 259–274 batch-8; gpt-oss-120B: 68.4 TG at 1k, 35.2 at 32k (corrected). | M5 Ultra 36/80, 512GB, 1.2 TB/s. Price unknown; late October. Qwen proxy est 91–105 TG, 363–384 batch-8 aggregate. | unknown | est +40% | Wait. Do not budget a purchase until Apple posts the price and a post-delivery serving run confirms thermals, speed, and memory behavior. |
The fleet priced both ways
This comparison uses the workhorse configuration:
- Previous: Apple-refurbished M4 Pro 48GB/512GB/10GbE at fact $1,949.
- New education: M5 Pro 15/16, 48GB/512GB/10GbE at est $2,229.
- New retail: the same configuration at est $2,399.
- Performance basis: measured previous aggregate 120 t/s at 8-wide; new planning estimate 132 t/s, or +10%.
- AppleCare, tax, racks, switches, storage upgrades, and spares are excluded.
| Fleet | Previous refurb cost | New education cost | New retail cost | Edu delta | Retail delta | Previous aggregate throughput | New aggregate throughput EST |
|---|---|---|---|---|---|---|---|
| 4-mini validation | fact (list price) $7,796 = 4×$1,949 | est $8,916 = 4×$2,229 | est $9,596 = 4×$2,399 | +$1,120; +14.4% | +$1,800; +23.1% | est 480 t/s = 4×120 | est 528 t/s = 4×132 |
| 26-mini faculty fleet | fact (list price) $50,674 = 26×$1,949 | est $57,954 = 26×$2,229 | est $62,374 = 26×$2,399 | +$7,280; +14.4% | +$11,700; +23.1% | est 3,120 t/s = 26×120 | est 3,432 t/s = 26×132 |
The previous totals are counterfactual list-price calculations. Availability of 4 or 26 matched refurbished units is unknown.
Cost per stream-capacity unit
A “stream-capacity unit” is 15 aggregate tokens/s. It is useful for cost comparison, but 8.8 units does not prove nine latency-compliant simultaneous sessions.
| Option | Aggregate throughput per node | 15-t/s capacity units | Hardware cost per capacity unit | Change from previous |
|---|---|---|---|---|
| Previous M4 Pro refurb | fact (proxy) 120 t/s | est 8.0 | est $243.63 = $1,949 / 8 | Baseline |
| New M5 Pro education | est 132 t/s | est 8.8 | est $253.30 = $2,229 / 8.8 | est +4.0% |
| New M5 Pro retail | est 132 t/s | est 8.8 | est $272.61 = $2,399 / 8.8 | est +11.9% |
Plan-ready statement: At education pricing, the newest 48GB fleet costs about 14.4% more and is expected to deliver about 10% more aggregate decode, raising hardware cost per stream-capacity unit by about 4.0%.
Power and cooling, both ways
Apple maximum power is an electrical envelope, not expected LLM draw. The only directly measured wall result found for the target mini class was 46.2W on an M4 Pro 64GB while generating a 70B model.
Formula: BTU/h = watts × 3.412142.
| Hardware | Per-node power basis | Evidence | 26-unit demand | 26-unit heat |
|---|---|---|---|---|
| M4 mini | 65W | fact Apple maximum wall power | 1.69 kW | 5,767 BTU/h |
| M4 Pro mini | 46.2W | fact measured wall power under Llama 70B generation | 1.201 kW | 4,099 BTU/h |
| M4 Pro mini | 140W | fact Apple maximum wall power | 3.64 kW | 12,420 BTU/h |
| M6 or M5 Pro mini | Actual LLM load unknown; 155W | fact Apple maximum continuous power | 4.03 kW maximum | 13,751 BTU/h maximum |
| M4 Max Studio | 145W | fact Apple maximum wall power | 3.77 kW | 12,864 BTU/h |
| M3 Ultra Studio | 270W | fact Apple maximum wall power | 7.02 kW | 23,953 BTU/h |
| M5 Max or M5 Ultra Studio | Actual LLM load unknown; 480W | fact Apple maximum continuous power | 12.48 kW maximum | 42,584 BTU/h maximum |
For electrical planning, the 26-mini maximum envelope rises from 3.64kW to 4.03kW, or est +10.7%. Do not compare the new 155W maximum directly with the old 46.2W measured workload result.
Sources: Apple’s Mac mini power-consumption page, Mac Studio power-consumption page, current technical specifications, and the measured M4 Pro LLM report.
Source log
All dynamic catalogs and store pages were accessed 2026-09-03 MDT.
| ID | Source | Published or document date | Used for |
|---|---|---|---|
| A01 | Apple: M6/M5 Pro Mac mini announcement | 2026-08-25 | Starting prices, availability, Apple performance claims |
| A02 | Apple: M5 Max/M5 Ultra Mac Studio announcement | 2026-08-25 | Prices, dates, workload claims, 512GB timing |
| A03 | Apple: M6 and M5 Ultra chip announcement | 2026-08-25 | CPU, GPU, AI and bandwidth claims |
| A04 | Apple: M5 Pro and M5 Max announcement | 2026-03 | M5 Pro/Max percentage claims |
| A05 | Apple: current Mac mini specifications | Current 2026-09-03 | Chips, memory, bandwidth, networking, maximum power |
| A06 | Apple: current Mac Studio specifications | Current 2026-09-03 | Chips, allowed memory tiers, bandwidth, 10GbE, maximum power |
| A07 | Apple retail Mac mini store | Current 2026-09-03 | Retail bases, option prices, date |
| A08 | Apple education Mac mini store | Current 2026-09-03 | Education bases and option deltas |
| A09 | Apple retail Mac Studio store | Current 2026-09-03 | Retail configurations and availability |
| A10 | Apple education Mac Studio store | Current 2026-09-03 | Education configurations |
| A11 | Apple: M4/M4 Pro Mac mini launch | 2024-10-29 | Previous starting prices and launch availability |
| A12 | Apple Support: 2024 Mac mini specifications | 2024 model | Previous memory and bandwidth |
| A13 | Apple: M4 Max/M3 Ultra Mac Studio launch | 2025-03-05 | Previous Studio launch prices and date |
| A14 | Apple: M3 Ultra announcement | 2025-03-05 | M3 Ultra specifications |
| A15 | Apple Support: 2025 Mac Studio specifications | 2025 model | Previous chip tiers, memory and bandwidth |
| A16 | Apple US Education Institution Price List | 2026-07-15 | Pre-refresh institutional catalog prices |
| A17 | Apple NASPO PSS catalog | 2026-06 | Cooperative-contract catalog |
| A18 | Apple education contracts: Utah | Accessed 2026-09-03 | Utah purchasing path |
| A19 | Apple refurb M4 Pro 48GB/10GbE | Accessed 2026-09-03 | $1,949 current listing |
| A20 | Apple refurb M4 Pro 64GB | Accessed 2026-09-03 | $2,379 current listing |
| A21 | Apple refurb M4 Max 128GB | Accessed 2026-09-03 | $4,159 current listing |
| A22 | Apple refurb M3 Ultra 256GB | Accessed 2026-09-03 | $8,149 current listing |
| A23 | Apple refurb M3 Ultra 512GB | Accessed 2026-09-03 | $17,419 current listing |
| A24 | CDW-G M4 Pro 24GB listing | Accessed 2026-09-03 | $1,599, backordered |
| A25 | CDW-G M4 Max 36GB listing | Accessed 2026-09-03 | $2,499, backordered |
| A26 | CDW-G M3 Ultra 96GB listing | Accessed 2026-09-03 | $5,299, backordered |
| A27 | CDW-G M3 Ultra 256GB listing | Accessed 2026-09-03 | $11,639.99, backordered |
| A28 | SHI Mac mini search | Accessed 2026-09-03 | Backorder/zero-stock evidence |
| A29 | SHI Mac Studio search | Accessed 2026-09-03 | Backorder/zero-stock evidence |
| A30 | MacRumors: maximum M3 Ultra configuration | 2025-03-05 | Dated M3 Ultra chip/memory option prices |
| A31 | Tom’s Hardware: M3 Ultra memory-upgrade price history | 2026-03-06 | Prior 256GB upgrade price |
| P01 | llama.cpp Apple Silicon performance table | Living discussion; accessed 2026-09-03 | Same-build M4/M5 Pro and Max PP/TG comparisons |
| P02 | llama.cpp gpt-oss guide | Living discussion; accessed 2026-09-03 | M4 Pro 64GB gpt-oss-20B benchmark |
| P03 | oMLX Qwen3.8 field report | 2026-08-18 and 2026-08-20 | M4 Pro 64GB and M3 Ultra 512GB serving |
| P04 | oMLX M4 Max 128GB benchmark | Accessed 2026-09-03 | Qwen3.5-35B serving |
| P05 | oMLX M3 Ultra 256GB benchmark | Accessed 2026-09-03 | gpt-oss-120B serving |
| P06 | oMLX M3 Ultra 512GB benchmark | Accessed 2026-09-03 | gpt-oss-120B serving |
| W01 | Apple: Mac mini power consumption | Accessed 2026-09-03 | Previous maximum wall power |
| W02 | Apple: Mac Studio power consumption | Accessed 2026-09-03 | Previous maximum wall power |
| W03 | Eastkode M4 Pro LLM wall-power measurement | 2026; accessed 2026-09-03 | 46.2W Llama 70B generation measurement |
Not run — the checks that need the machines in hand
- not run Physical serving benchmarks on M6, M5 Pro mini, M5 Max Studio, or M5 Ultra Studio; hardware is not available until September 22 or later.
- not run UVU-specific Apple Education, NASPO, CDW-G, or SHI quote.
- not run Stock reservation for 4 or 26 identical previous-generation units.
- not run Exact AppleCare+ premium quote for each new configuration.
- not run Matched end-to-end benchmark using the same model, quantization, serving version, context, batch, and power meter across both chassis generations.
- not run Validation of the owner’s private M3 Ultra benchmark; no numerical receipt was supplied in this lane.
- not run Purchase, order, deployment, or other live-system action.
Review: working paper 36, September 3, 2026 — every price and date checked against the source listed; nothing here contacted UVU or any vendor on UVU's behalf. Configured-to-order prices are Apple's displayed option deltas added to a sourced base price, not institutional quotes. The guide's §04 carries the plain-language summary and the fleet priced both ways.
Every angle, and where it lives
Appendix P in five lines
- QuestionHas the plan checked every issue a university leader may raise, and where is each answer?
- AnswerAlmost; the current table has 37 rows marked COVERED and 1 marked OPEN.
- Deciding numbers38 anglesest37 COVERED rows, a derived countest1 OPEN row, a derived countest
- What the plan doesUse this matrix as the plan’s index, keep the open quality issue as a test gate, and grade it again after each wave.
- Still unknownQuality loss from compressed models stays open until the delivered machines are benchmarked.
The owner's test for this plan is not whether each number is right but whether every angle a university leader would raise has been looked at. This page is the answer, kept honest: each angle graded COVERED (analyzed with evidence), THIN (mentioned, not analyzed), or OPEN (not yet workable), with the place it lives. First graded September 3, 2026 by searching the shipped guide and appendices for each angle; regraded the same afternoon after the five engineering, safety, law, and money reviews (Appendices Q–U) landed. It is regraded every wave; the rows that were THIN at the first grading became lanes of their own and are now COVERED, which leaves 37 COVERED and one OPEN.
The owner's test for this plan is not whether each number is right but whether every angle a university leader would raise has been looked at. This page is the answer, kept honest: each angle is graded COVERED (analyzed with evidence), THIN (mentioned, not analyzed), or OPEN (not yet worked), with the place it lives. First graded on 2026-09-03 by searching the guide and appendices for each angle, then regraded twice the same day as Appendices Q–U and V–Z were added. It is regraded whenever new evidence lands.
| Angle | Grade | Where it lives |
|---|---|---|
| Model choice per memory tier; the most intelligent open models | COVERED | §11, Appendix L, Appendix Q |
| Capacity: conversations per box per model; intelligence per dollar | COVERED | §04, Appendix Q, configurator (by model and by workload) |
| Two or more models resident on one machine; swap vs resident | COVERED | §04, Appendix Q |
| 512GB frontier node: when it earns its place | COVERED | §04 three-path decision, Appendix Q |
| Workload classes: chat, coding, documents, agents, speech, images | COVERED | §06, Appendix R |
| Prompt-reading (prefill) limits on Apple Silicon; long documents | COVERED | §06, §08, Appendix R |
| Batching, prefix-cache reuse, speculative decoding, quantization loss | COVERED | Appendix R |
| Service levels (time to first token, p95) and a real queue model | COVERED | §06, Appendix R |
| Coding-agent loads: their own queue, cap, overflow | COVERED | §06, §08, §10, Appendix R |
| Minors: the 18,163 concurrent-enrollment students | COVERED | §08, §10, Appendix S (eleven gates) |
| Health data: HIPAA vs FERPA, placements, business associates | COVERED | Appendix S |
| Crisis disclosures, duty of care, human handoff | COVERED | §08, §10, Appendix S |
| Prompt injection, exfiltration, model supply chain, abuse limits | COVERED | §08, Appendix S (twenty graded controls) |
| Failover, single points of failure, backups, recovery targets, incident runbook | COVERED | §08, §10, Appendix S |
| Utah AI disclosure law and the Office of AI Policy | COVERED | §08, §09, Appendix T |
| ADA Title II web accessibility rule (WCAG 2.1 AA by April 26, 2027) | COVERED | §09, §10, Appendix T |
| FERPA and open-records status of chat logs; retention; legal holds | COVERED | §08, Appendix T |
| State higher-education policy, USHE task force and credential, legislature, governor | COVERED | §07, §09, Appendix T |
| Model license fitness for a public university, by exact checkpoint | COVERED | §11, Appendix T |
| Utah peers and similar public universities elsewhere | COVERED | §12, Appendix T |
| Five-year cost, refresh cycle, resale value | COVERED | §04, §12, Appendix U, simulator |
| Lease vs buy | COVERED | Appendix U |
| Who pays: central, colleges, chargeback, student fees | COVERED | §09, Appendix U |
| Student operations staff | COVERED | §09, Appendix U, simulator and configurator switch |
| Grants and partnerships | COVERED | §09, Appendices F and U |
| Electricity at UVU's rate | COVERED (proxy; UVU rate unknown) | Appendix U |
| Adoption, behavior, catalysts | COVERED | §07, Appendix N |
| Procurement path, UVU policies 445/447/452 | COVERED | §09, Appendix H |
| Contingencies (28 branches), failure-modes table | COVERED | §10, Appendix G |
| Free-tier cloud, harness market | COVERED | Appendices L, M |
| Hardware generations, prices, availability | COVERED | §04, Appendix O |
| Campus map, department demand | COVERED | map, Appendices D, I |
| Pedagogy: assessment redesign and integrity | COVERED | §07, §10, Appendix V (160-plus sources) |
| Pedagogy: learning outcomes and how to measure them | COVERED | §07, Appendix W (learning-outcomes contract) |
| Human factors: trust, sources, error handling in the interface | COVERED | §08, Appendix X (trust contract, test plan) |
| Accessibility and language: screen readers, WCAG for chat, Spanish, neurodivergent, mobile | COVERED | §08, §09, Appendix Y |
| Roadmap: Apple cadence, open-model trajectory, buy-vs-wait, waves | COVERED | §09, Appendix Z (24-month watchlist) |
| Quantized-model quality vs published scores; delivered-unit benchmarks | OPEN until hardware ships | Appendix Q and R acceptance-test lists |
Model × machine: the matrix
Appendix Q in five lines
- QuestionShould the university buy a 512-gigabyte computer for the smartest open model, and how many people can each model-and-machine pair serve?
- AnswerKeep smaller Mac minis as the default and keep the 512-gigabyte box as a research choice only after its price and tests pass.
- Deciding numbers2 × 60 = 120 score-conversations on the large boxest40 × 52 = 2,080 on five minisestabout $100,000 and no more than 15% of hardware spending as the buying gateest
- What the plan doesKeep minis as the campus default, use two-model pairs that fit, and hold the large research box behind price and delivered-machine tests.
- Still unknownThe 512-gigabyte price, new-Mac tests, full 32,000-token batch tests, and compressed-model quality tests are NOT_RUN.
The owner's question was direct: why not run the most intelligent models on 512GB Studios, and did we do the multivariate analysis — every model against every machine, capacity included, with more than one model per box? The plan had one capacity number ("four conversations per box") applied to everything. This appendix replaces it with the matrix: ten open models × six machines, each cell with memory fit, single-conversation speed, prompt-reading speed, conversations at ≥10 and ≥20 tokens/s, cost per conversation, and an intelligence-per-dollar index; then the two-models-per-box analysis, the same-money comparison, where an eight-point score gap actually matters, and the three-path decision on a 512GB frontier node. Research cutoff September 3, 2026. fact dated measurements · est derived, basis shown · unknown no defensible number.
What this changes in the plan
1. "Four per box" is retired as a fleet-wide constant. It was conservative for the daily 27B model on a mini (about eight conversations at ≥10 tokens/s by the measured proxies) and for the sparse 35B fast model, fair for coding, and optimistic for the biggest models: on a 512GB node the best open model (GLM-5.3, score 60) fits only two 32k-context conversations with the preferred quantization, at about 18 tokens/s each. The engine keeps 4 as the blended admission cap for mixed campus traffic (Appendix R explains why), and the configurator now shows capacity by workload and by model. 2. Same money, both ways. The 256GB Ultra's price buys five 48GB minis: five minis carry about 40 conversations of a score-52 model; one 512GB node carries two conversations of a score-60 model — 2,080 versus 120 "score-conversations". The eight points matter for advanced coding (Terminal-Bench 88.2 vs 73.0; DeepSWE 66.9 vs 42.2) and complex professional artifacts (a 217-Elo lead on GDPval); for tutoring chat, summaries, and ordinary drafts no matched result shows a difference. 3. The 512GB node becomes a conditional research node (Path 1): kept as an unpriced option behind the router, bought only when the hardware pool reaches about $100,000, it takes no more than 15% of hardware spend, Apple has posted the price, and a delivered-unit test passes. What breaks first on it: long-document reading (a 131k-token prompt took about 23 minutes on the previous generation) and a single point of failure. 4. Two models on one machine is real and now designed in: a 48GB mini holds the 27B daily model plus a 9B fast model comfortably; the 64GB mini is the one that holds the 27B plus the sparse 35B fast model with production margin — that, not speed, is what the extra $360 buys; a 128GB Studio holds a 120B-class model plus the 27B; a 256GB Studio holds GLM-5.3-Flash plus the 27B; a 512GB node running GLM-5.3 at the preferred quantization has no room for a companion. Swapping a 27B takes about three seconds; a 428GB model about 75 seconds — not interactive. llama-server (router mode), oMLX, LM Studio, and vLLM-MLX all support keeping several models resident today. 5. Names fixed. The 256GB anchor is GLM-5.3-Flash (320B total, 18B active, score 57, MIT), not an unnamed "235B-class"; the Appendix O row that credited gpt-oss-120B with 1,405.9 prompt / 79.2 writing tokens/s belonged to a different model and is corrected below (the verified gpt-oss-120B result on the previous 512GB Ultra: 68.4 tokens/s at 1k context, 35.2 at 32k, 187 aggregate at batch eight). est
Research cutoff: 2026-09-03 MDT.
fact means a dated source states or measures the number. est means arithmetic or a proxy, with its basis stated. unknown means no defensible number was available.
Executive verdict
- The current “four conversations per box at ≥10 t/s” rule is not valid as one fleet-wide constant.
- est It is conservative for Qwen3.8-27B and Qwen3.6-35B-A3B workhorses.
- est It is too optimistic for Qwen3.8-Flash-Next on 128GB when four 32k contexts are required.
- est It is too optimistic for the quality-preferred 427.7GB GLM-5.3 build: only two 32k contexts fit under the production memory rule.
- FACT/EST A 512GB GLM node buys real gains on advanced coding and complex professional artifacts, but minis buy far more concurrent capacity per dollar.
- Recommendation: retain minis as the campus default. Keep a 512GB node as a conditional research lane only after Apple posts the price and UVU verifies the exact model on delivered hardware.
1. Candidate models
Practical-memory rule
est I reserve 15% of unified memory for macOS, Metal/MLX buffers, the server, allocator variance, and safety. Model weights plus every reserved 32k KV cache must fit inside the remaining 85%.
Quantized-model scores remain a caveat: Artificial Analysis scores the model, not each community Apple quantization. Quality retained by a given quant is unknown until tested.
Model ledger
| Model | Total / active | AA Intelligence Index | License and inputs | Practical Apple weights | BF16 KV for one 32k conversation | Weights + one cache |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | fact 9.7B dense | fact 22, accessed 2026-09-03 | Apache-2.0; text/image/video | fact 5.98GB MLX 4-bit | est 1.074GB | est 7.05GB |
| Qwen3.8-27B | fact 27B dense | fact 52 | Apache-2.0; text/image/video | fact 16.1GB MLX 4-bit | fact (math) 2.147GB | est 18.25GB |
| Qwen3.6-35B-A3B | fact 35B / 3B; AA rounds total to 36B | fact 32 | Apache-2.0; maker supports text/image/video; AA records text/image | fact 20.7GB MLX DWQ | fact (math) 0.671GB | est 21.37GB |
| Qwen3.8-Flash-Next | fact 180B stored / 6B active | fact 56 | Qwen Community 1.0; text/image/video | fact 106.2GB mixed MLX | est 0.830GB | est 107.03GB |
| gpt-oss-120b | fact 117B / 5.1B | fact 24 | Apache-2.0; text | fact 62.4GB MLX MXFP4-Q4 | est 1.213GB | est 63.61GB |
| Mistral Small 4 | fact 119B / 6.5B | fact 20 | Apache-2.0; text/image | fact 67.8GB MLX 4-bit | fact (math) 0.755GB | est 68.55GB |
| GLM-5.3-Flash | fact 320B / 18B | fact 57 | MIT; maker supports text/image/video/files | fact 181.9GB preferred mixed MLX; uniform is 177.6GB | est 0.392GB | est 182.29GB |
| GLM-5.3 | fact 753B checkpoint / 40B active | fact 60 | Custom GLM-5.3 license; text | fact 427.7GB preferred mixed MLX; uniform is 418.6GB | FACT/EST 3.121GB | est 430.82GB |
| Kimi K3 | fact 2.8T / 104B | fact 60 | Custom Kimi K3 license; maker supports text/image/video | fact 1.56TB official MXFP4; fact 928.6GB quality GGUF | fact (math) 0.906GB | est 929.51GB using GGUF |
| DeepSeek V4 Pro 0813 | fact about 1.57T / 48B; AA rounds to 1.6T/49B | fact 53 | MIT; text | fact about 850GB Q4_K_XL GGUF | fact (math) 0.331GB optimized hybrid cache | est 850.33GB |
Model facts and scores come from the current Artificial Analysis model pages and the makers’ model cards: Qwen3.8-27B, Qwen3.6-35B-A3B, Qwen3.8-Flash-Next, gpt-oss-120b, Mistral Small 4, GLM-5.3-Flash, GLM-5.3, Kimi K3, and DeepSeek V4 Pro.
KV-cache math
All arithmetic uses 32,768 tokens and two-byte BF16 cache values.
Qwen3.5-9B: 32,768 × 8 full layers × 2 K/V × 4 KV heads × 256 × 2 bytes = 1,073,741,824 bytes Qwen3.8-27B: 32,768 × 16 full layers × 2 × 4 × 256 × 2 = 2,147,483,648 bytes Qwen3.6-35B-A3B: 32,768 × 10 full layers × 2 × 2 × 256 × 2 = 671,088,640 bytes Qwen3.8-Flash-Next: core = 32,768 × 12 full layers × 2 × 2 × 256 × 2 index estimate = (32,768 / 4) × 12 × 128 × 2 total = 830,472,192 bytes gpt-oss-120b: [(32,768 × 18 full layers) + (128 × 18 sliding layers)] × 2 K/V × 8 KV heads × 64 × 2 = 1,212,678,144 bytes Mistral Small 4 optimized MLA: 32,768 × 36 × (256 latent + 64 rope) × 2 = 754,974,720 bytes GLM-5.3-Flash: main = 32,768 × 11 MLA layers × 512 latent × 2 plus estimated compressed index = about 23MB total = about 392MB GLM-5.3: main = 32,768 × 78 × (512 latent + 64 rope) × 2 indexers = 32,768 × 21 × 128 × 2 total = 3,120,562,176 bytes Kimi K3: 32,768 × 24 MLA layers × (512 latent + 64 rope) × 2 = 905,969,664 bytes DeepSeek V4 Pro: 30 c4 layers × 10,616,832 bytes + 31 c128 layers × 393,216 bytes = 330,694,656 bytes
Hybrid linear-attention models also have fixed recurrent state. That state is not fully exposed by model configs; it is covered by the 15% reserve.
Releases since August 15
- fact Qwen3.8-Flash-Next arrived August 26 and changes the 128GB quality ceiling.
- fact GLM-5.3-Flash received its formal open-weight launch September 2 and changes the 256GB choice.
- fact K2 Horizon launched September 3. Its 375B-A23B flagship scores 47, but the interesting 32B and 36B-A4B variants have no current independent score or proven Apple quant. est It is a watch item, not a matrix winner today. IFM announcement, AA flagship analysis.
2. Machines
| Machine | Physical / usable memory | Bandwidth | Price used | Availability |
|---|---|---|---|---|
| Mac mini M5 Pro 48GB | fact 48GB; est 40.8GB usable | fact 307GB/s | est $2,139 education | fact 2026-09-22 |
| Mac mini M5 Pro 64GB | fact 64GB; est 54.4GB usable | fact 307GB/s | est $2,499 education | fact 2026-09-22 |
| Mac Studio M5 Max 128GB | fact 128GB; est 108.8GB usable | fact 614GB/s | est $4,639 education, 512GB SSD | fact 2026-09-22 |
| Mac Studio M5 Ultra 256GB | fact 256GB; est 217.6GB usable | fact 1.2TB/s | est $10,799 retail, 1TB SSD; education unknown | fact 2026-09-22 |
| Mac Studio M5 Ultra 512GB | fact 512GB; est 435.2GB usable | fact 1.2TB/s | unknown retail and education | fact late October 2026 |
| Mac Studio M3 Ultra 512GB proxy | fact 512GB; est 435.2GB usable | fact 819GB/s | fact $17,419 current 8TB refurb listing; launch-equivalent est $9,499 | Available quantity unknown |
Hardware and dates are from Apple’s Mac mini announcement, Mac Studio announcement, mini specifications, and Studio specifications.
The configured price sums are est, not UVU quotes. In particular, data.js treats some derived prices as facts and includes a $9,899 Ultra education estimate that is not supported by a current Apple quote.
3. Model × machine matrix
Estimation method
For measured or closely matched models, I use the published oMLX/MLX result.
For unmeasured MoE models:
EST single decode = memory bandwidth ÷ active 4-bit bytes per token × 30% efficiency
est 30% is bracketed by approximately 34% for the measured Qwen 3B-active MoE and approximately 23% for measured gpt-oss MXFP4.
New-generation planning factors are:
- est ×1.10 for M5 Pro/Max versus comparable M4 measurements.
- est ×1.40 for M5 Ultra versus M3 Ultra.
S10/S20 means estimated simultaneous conversations at at least 10/20 generated tokens per second. Counts reserve one 32k cache each, but published batch tests generally use shorter active prompts. Therefore 32k memory fit is stronger evidence than 32k speed.
8 means “at least eight under the batch proxy”; no claim above eight is made.
Score-streams/$10k = AA score × S10 × 10,000 ÷ hardware price. It is only a capacity-quality index.
M5 Pro mini 48GB
| Model | Fit | Single decode | Prefill | S10 / S20 | Cost per S10 / S20 | Score-streams per $10k |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | est yes | est 62 t/s | unknown | est 8 / 8 | est $267 / $267 | est 823 |
| Qwen3.8-27B | est yes | est 22.9 t/s @8k | est 110–141 t/s @4k | est 8 / 1 | est $267 / $2,139 | est 1,945 |
| Qwen3.6-35B-A3B | est yes | est 62 t/s; same-chip 20-GPU upper proxy is 68.7 | est about 1,600 @16k | est 8 / 8 | est $267 / $267 | est 1,197 |
| Other seven models | FACT/EST no | — | — | 0 / 0 | — | — |
The 48GB machine can fit Qwen3.8-27B and Qwen3.6 together only with narrow headroom; see co-residency.
M5 Pro mini 64GB
Decode speed is the same as 48GB because bandwidth is unchanged.
| Model | Fit | Single decode | Prefill | S10 / S20 | Cost per S10 / S20 | Score-streams per $10k |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | est yes | est 62 | unknown | est 8 / 8 | est $312 / $312 | est 704 |
| Qwen3.8-27B | est yes | est 22.9 @8k | est 110–141 @4k | est 8 / 1 | est $312 / $2,499 | est 1,665 |
| Qwen3.6-35B-A3B | est yes | est 62 | est about 1,600 @16k | est 8 / 8 | est $312 / $312 | est 1,024 |
| Other seven models | FACT/EST no | — | — | 0 / 0 | — | — |
The extra $360 buys co-residency and cache space, not more decode bandwidth.
M5 Max Studio 128GB
| Model | Fit | Single decode | Prefill | S10 / S20 | Cost per S10 / S20 | Score-streams per $10k |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | est yes | est 135 | unknown | est 8 / 8 | est $580 / $580 | est 379 |
| Qwen3.8-27B | est yes | est 52 | unknown | est 8 / 8 | est $580 / $580 | est 897 |
| Qwen3.6-35B-A3B | est yes | est 110 @8k | est 1,635 @8k, same-chip proxy | est 8 / 8 | est $580 / $580 | est 552 |
| Qwen3.8-Flash-Next | est yes, tight | est 61 | unknown | est 3 / 3, memory-limited | est $1,546 / $1,546 | est 362 |
| gpt-oss-120b | est yes | est 49 short-context | unknown | est 8 / 6 short-context | est $580 / $773 | est 414 |
| Mistral Small 4 | est yes | est 57 | unknown | est 8 / 7 | est $580 / $663 | est 345 |
| Larger four models | FACT/EST no | — | — | 0 / 0 | — | — |
Qwen3.8-Flash-Next permits three 32k caches by arithmetic, not four. Its current custom Apple artifact also excludes MTP and operates text-only despite carrying vision weights.
M5 Ultra Studio 256GB
| Model | Fit | Single decode | Prefill | S10 / S20 | Cost per S10 / S20 | Score-streams per $10k |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | est yes | est 244 | unknown | est 8 / 8 | est $1,350 / $1,350 | est 163 |
| Qwen3.8-27B | est yes | est 98, range 91–105 | est about 603 @8k | est 8 / 8 | est $1,350 / $1,350 | est 385 |
| Qwen3.6-35B-A3B | est yes | est 205 | unknown | est 8 / 8 | est $1,350 / $1,350 | est 237 |
| Qwen3.8-Flash-Next | est yes | est 120 | unknown | est 8 / 8 | est $1,350 / $1,350 | est 415 |
| gpt-oss-120b | est yes | est 96 short-context | est about 841 @32k | est 8 / 8 short-context | est $1,350 / $1,350 | est 178 |
| Mistral Small 4 | est yes | est 111 | unknown | est 8 / 8 | est $1,350 / $1,350 | est 148 |
| GLM-5.3-Flash | est yes | est 40 | unknown | est 8 / 4 | est $1,350 / $2,700 | est 422 |
| GLM-5.3, Kimi K3, DeepSeek V4 Pro | FACT/EST no | — | — | 0 / 0 | — | — |
The plan’s unnamed “235B-class” anchor remains unknown: a parameter count does not specify active weights, artifact size, cache design, or batch behavior. GLM-5.3-Flash is the stronger named 256GB choice.
M5 Ultra Studio 512GB
Price-based columns are unknown until Apple posts the 512GB price, denoted P.
| Model | Fit | Single decode | Prefill | S10 / S20 | Cost per S10 / S20 | Score-streams per $10k |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | est yes | est 244 | unknown | est 8 / 8 | est P/8 / P/8 | unknown |
| Qwen3.8-27B | est yes | est 98 | est about 603 @8k | est 8 / 8 | est P/8 / P/8 | unknown |
| Qwen3.6-35B-A3B | est yes | est 205 | unknown | est 8 / 8 | est P/8 / P/8 | unknown |
| Qwen3.8-Flash-Next | est yes | est 120 | unknown | est 8 / 8 | est P/8 / P/8 | unknown |
| gpt-oss-120b | est yes | est 96 short-context | est about 841 @32k | est 8 / 8 short-context | est P/8 / P/8 | unknown |
| Mistral Small 4 | est yes | est 111 | unknown | est 8 / 8 | est P/8 / P/8 | unknown |
| GLM-5.3-Flash | est yes | est 40 | unknown | est 8 / 4 | est P/8 / P/4 | unknown |
| GLM-5.3 mixed 4/8 | est yes, very tight | est 18 | est about 132; GLM-5.2 131k proxy | est 2 / 0, memory-limited | est P/2 / N/A | unknown |
| Kimi K3, DeepSeek V4 Pro | fact no | — | — | 0 / 0 | — | — |
The lower-quality 418.6GB uniform GLM build has room for five calculated 32k caches and could reach est three ≥10 t/s streams. The preferred 427.7GB mixed build fits only est two.
M3 Ultra Studio 512GB measured proxy
| Model | Fit | Single decode | Prefill | S10 / S20 | Cost per S10 / S20 | Score-streams per $10k |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | est yes | est 174 | unknown | est 8 / 8 | est $2,177 / $2,177 | est 101 |
| Qwen3.8-27B | fact yes | fact 65–79; central 70 | fact 420–442; central 431 | FACT/EST 8 / 8 | est $2,177 / $2,177 | est 239 |
| Qwen3.6-35B-A3B | est yes | est 141 | unknown | est 8 / 8 | est $2,177 / $2,177 | est 147 |
| Qwen3.8-Flash-Next | est yes | est 82 | unknown | est 8 / 8 | est $2,177 / $2,177 | est 257 |
| gpt-oss-120b | fact yes | fact 68.4 @1k; 57.8 @8k; 35.2 @32k | fact 600.7 @32k | FACT/EST 8 / 8 short-context | est $2,177 / $2,177 | est 110 |
| Mistral Small 4 | est yes | est 76 | unknown | est 8 / 8 | est $2,177 / $2,177 | est 92 |
| GLM-5.3-Flash | est yes | est 27 | unknown | est 7 / 2 | est $2,488 / $8,710 | est 229 |
| GLM-5.3 mixed 4/8 | est yes, tight | est 12 | est about 95; GLM-5.2 proxy | est 1 / 0 | est $17,419 / N/A | est 34 |
| Kimi K3, DeepSeek V4 Pro | fact no | — | — | 0 / 0 | — | — |
Important source correction
findings-36 associates oMLX result mrqg3jhh with gpt-oss-120b, but that result is for Qwen3.5-122B-A10B.
The verified gpt-oss-120b M3 Ultra result is:
- fact 68.4 t/s at 1k.
- fact 57.8 t/s at 8k.
- fact 35.2 t/s at 32k.
- fact 187.1 t/s aggregate at batch eight in the short-context batch test.
The earlier 1,405.9 prefill / 79.2 decode claim should be removed.
Audit of “four conversations per box”
| Tier and model | Verdict | Why |
|---|---|---|
| 48GB or 64GB, Qwen3.8-27B | est conservative at ≥10 | Old 64GB proxy reached 120–134 aggregate at batch eight. Full 32k decode is still unknown. |
| 48GB or 64GB, Qwen3.6-35B-A3B | est conservative | Sparse active weights and same-chip results leave substantial throughput room. |
| 128GB, Mistral or gpt-oss | est conservative at normal context | Both have memory for many caches and estimated S10 of eight. |
| 128GB, Qwen3.8-Flash-Next | est optimistic | Only three 32k caches fit under the rule. |
| 256GB, GLM-5.3-Flash | est conservative at ≥10; fair near ≥20 | Estimated eight S10 but only four S20. |
| 256GB, unnamed “235B-class” | unknown | No exact model or quant was specified. |
| 512GB, preferred GLM-5.3 | est optimistic | Mixed quant permits only two 32k contexts; uniform quant permits more but has lower known quant quality. |
| 512GB, GLM-5.3-Flash | est conservative | Eight S10 is plausible with comfortable memory. |
4. More than one model on one machine
Pairs that fit
Figures include one 32k cache for each model.
| Machine | Pair | Memory | Verdict |
|---|---|---|---|
| 48GB mini | Qwen3.8-27B + Qwen3.5-9B | est 25.30GB | Comfortable |
| 48GB mini | Qwen3.8-27B + Qwen3.6-35B | est 39.62GB of 40.8GB usable | Fits by arithmetic; too little operating margin for production |
| 64GB mini | Qwen3.8-27B + Qwen3.6-35B | est 39.62GB of 54.4GB | Good dual-resident case |
| 128GB Studio | gpt-oss-120b + Qwen3.8-27B | est 81.86GB | Comfortable |
| 128GB Studio | Mistral Small 4 + Qwen3.8-27B | est 86.80GB | Comfortable |
| 128GB Studio | Qwen3.8-Flash-Next + another useful model | More than 108.8GB | Does not fit under production rule |
| 256GB Studio | GLM-5.3-Flash + Qwen3.8-27B | est 200.54GB | Fits with about 17GB usable headroom |
| 512GB Studio | preferred mixed GLM-5.3 + Qwen3.5-9B | est 437.87GB | Does not fit under 85% rule |
| 512GB Studio | uniform GLM-5.3 + Qwen3.5-9B | est 428.77GB | Fits, but uses the lower-quality GLM quant |
| 512GB Studio | GLM-5.3 + Qwen3.6-35B | At least est 452GB using preferred GLM | Does not fit |
The proposed “GLM-5.3 plus a fast MoE lane” is therefore not valid with the preferred mixed quant and the production reserve. A small companion requires the uniform GLM build or a looser safety rule.
Speed cost when both models are active
Memory bandwidth is shared. A simple first-order estimate is:
active speed ≈ solo speed × assigned bandwidth share
| Machine and pair | Equal 50/50 split | Better routing choice |
|---|---|---|
| 48GB: Qwen27 + Qwen9 | est 11.5 + 31 = 42.5 aggregate | Keep Qwen27 above 10; send short/simple work to Qwen9 |
| 64GB: Qwen27 + Qwen35 | est 11.5 + 31 = 42.5 | Qwen35 handles high-volume chat; Qwen27 gets quality requests |
| 128GB: gpt-oss + Qwen27 | est 24.5 + 26 = 50.5 | Both remain responsive |
| 256GB: GLM-Flash + Qwen27 | est 20 + 49 = 69 | Strong coding/research lane plus workhorse |
| 512GB: uniform GLM + Qwen9 | est 9 + 122 | Give GLM at least 56% of bandwidth: est 10 GLM + 107 Qwen9 |
These are ideal bandwidth splits, not scheduler measurements.
Serving-stack support today
- fact llama-server has router mode, model-specific child processes, dynamic load/unload, autoload, and a default maximum of four loaded models. A current LRU issue can unload an active model, so production validation is required.
- fact oMLX has a multi-model engine pool with LRU, TTL, and explicit load/unload. An M4 Pro 64GB field run kept about 37.73GB of two models resident and switched between them in fact 0.23–0.9 seconds.
- fact LM Studio can hold multiple models when JIT auto-evict is off or models are loaded manually. Its default auto-evict behavior keeps at most one JIT model.
- fact vLLM-MLX provides a named multi-model registry, lazy loading, preloading, LRU, and wait/fail/preempt policies. Its memory budget counts weights, so KV, activations, and OS reserve must be added by the operator.
- fact Native mlx-lm server replaces its primary model when a different model key is loaded; it is not a multi-resident router.
- fact Ollama can keep multiple models resident when memory permits, controlled separately from per-model parallelism.
Swap times
Prior internal-SSD measurements are fact 5.1–5.8GB/s. New M5 desktop SSD speed is unknown.
| Model | Raw-read floor | Operational meaning |
|---|---|---|
| Qwen3.5-9B | est 1.0–1.2s | Easy to swap |
| Qwen3.8-27B | est 2.8–3.2s | Often acceptable on demand |
| Qwen3.6-35B | est 3.6–4.1s | Often acceptable |
| gpt-oss-120b | est 10.8–12.2s | Noticeable |
| Mistral Small 4 | est 11.7–13.3s | Noticeable |
| Qwen3.8-Flash-Next | est 18.3–20.8s | Keep resident if used regularly |
| GLM-5.3-Flash | est 31.4–35.7s | Keep resident during active periods |
| GLM-5.3 preferred mixed | est 73.7–83.9s | Swapping is poor interactive service |
An oMLX incident loaded fact 403.96GB in 73 seconds, or 5.53GB/s effective, although the storage medium was not proved to be the internal SSD.
Residency verdict
- Tutoring, summaries, and first drafts: pin Qwen3.6-35B or Qwen3.5-9B.
- General faculty work: pin Qwen3.8-27B.
- Advanced coding and complex document work: route to Qwen3.8-Flash-Next, GLM-Flash, or the metered frontier pool.
- Keep two models resident when both receive frequent, alternating traffic.
- For rare large-model requests, swapping a 60–70GB model can be acceptable. Swapping a 428GB model is not.
5. Frontier questions
Same money: 512GB GLM versus 48GB minis
Let the unknown 512GB M5 Ultra price be P.
Same-money line: est GLM preferred quant: 2 conversations × score 60 = 120 score-conversations, cost P/2 each; Qwen27 minis: 8 × floor(P / $2,139) conversations × score 52 = 416 × floor(P / $2,139) score-conversations, cost about $267 each.
For illustration only, using the cheaper 256GB Ultra’s est $10,799 price—not a 512GB quote—buys five 48GB minis:
GLM node: 2 × 60 = 120 score-conversations Five minis: 40 × 52 = 2,080 score-conversations
This multiplication is a routing-capacity heuristic, not a measure of completed academic work.
Where the eight-point gap matters
The comparison is AA 60 versus 52. The AA methodology weights agentic work, coding, scientific reasoning, and general knowledge. Eight index points do not mean eight percentage points on every task.
| UVU work | Evidence | Verdict |
|---|---|---|
| Advanced coding | fact Official cards report GLM versus Qwen: Terminal-Bench 2.1 88.2 vs 73.0, DeepSWE 66.9 vs 42.2, NL2Repo 58.0 vs 42.3 | Gap matters |
| Complex research and professional artifacts | fact Independent GDPval-AA v2 reports 1763 vs 1546 Elo, a 217-Elo GLM lead across work products such as memos, slides, and spreadsheets | Gap matters |
| Math and scientific reasoning | unknown Qwen reports GPQA 89.2, while GLM-5.3 does not publish a directly comparable current GPQA figure | Do not claim a proven advantage |
| Tutoring chat | unknown No AA component directly measures ordinary campus tutoring | Route to cheaper model unless a UVU test shows a difference |
| Basic summaries | unknown No matched task result proves the GLM advantage | Qwen-tier model is sufficient until tested |
| Ordinary drafts | unknown GDPval supports complex professional work, not simple prose | Qwen-tier model is the default |
Official task-level results: GLM-5.3 card, Qwen3.8-27B card, and GDPval-AA.
When a 512GB node earns a place
Path 1 — Recommended: conditional shared research node
- Result: keep it in the plan as an unpriced option behind the router.
- Budget trigger: est total hardware pool of about $100,000 or more, assuming an eventual node price near $15,000 and a policy that one frontier node consumes no more than 15% of hardware spend.
- Benefit: local advanced coding and professional-artifact work.
- Risk: price, quant quality, concurrency, thermals, and long-context behavior remain unproved.
- Undo: remove the option before purchase if price or acceptance results fail.
Path 2 — Ring-fenced research purchase
- Result: one department or grant buys it once P is posted.
- Benefit: earlier access for research that has demonstrated need.
- Risk: low utilization and no campus-service redundancy.
- Undo: repurpose it for batch research, but the purchase itself is hard to undo.
Path 3 — No 512GB node
- Result: spend the same money on minis, 128GB nodes, and the metered frontier pool.
- Benefit: much higher concurrency and multiple failure domains.
- Risk: some advanced local work remains cloud-routed.
- Undo: easy; add a later model generation when evidence improves.
Approve Path 1, the conditional shared research-node option? Yes or no.
What breaks first
Long-document prefill breaks first.
The closest measured proxy is GLM-5.2, which shares GLM-5.3’s base architecture:
- fact A 131k-token cold prefill took about 1,385 seconds, roughly 23 minutes.
- fact A 196k attempt was rejected at an estimated 494.81GB peak.
- fact Its steady KV allocation was only about 20GiB.
- est M5 Ultra’s planning gain does not turn a multi-minute 131k prefill into an interactive request.
That is transient prefill memory and latency failing before steady 32k KV capacity. One node is still a structural single point of failure: a hardware or service fault removes fact 100% of the local frontier lane.
Source: oMLX GLM long-context report.
What the plan should change
- 1. Replace the blanket four-stream assumption. Store streams10 and streams20 by exact model, quant, context, runtime, and machine.
- 2. Keep Qwen3.8-27B as the 48GB default. It has the best estimated score-capacity value in that tier.
- 3. Buy 64GB minis only for co-residency or longer contexts. They have the same 307GB/s bandwidth and expected single-stream speed as 48GB.
- 4. Use Mistral Small 4 or gpt-oss as the stable 128GB shared service. Treat Qwen3.8-Flash-Next as a three-context pilot, not a four-stream default.
- 5. Name GLM-5.3-Flash as the 256GB candidate. Remove the unspecified “235B-class” capacity claim.
- 6. Do not count four GLM-5.3 conversations. Use est two for the preferred mixed quant or est three for the smaller uniform quant, pending measurement.
- 7. Make the 512GB node conditional and research-oriented. Do not use it as the campus workhorse.
- 8. Correct the gpt-oss benchmark citation and price evidence states in findings-36 and data.js.
- 9. Pin small daily models and route advanced requests. Avoid making two large resident models generate simultaneously unless latency is not important.
- 10. Run acceptance tests after delivery. The minimum matrix is 1k/8k/32k prefill and decode at batch 1/2/4/8, cold-load time, peak memory, and a 24-hour mixed-traffic soak.
Source log
- 1. Apple Mac mini announcement, 2026-08-25 — price bases and availability.
- 2. Apple Mac Studio announcement, 2026-08-25 — Ultra tiers and late-October timing.
- 3. Apple Mac mini specifications, accessed 2026-09-03 — memory and bandwidth.
- 4. Apple Mac Studio specifications, accessed 2026-09-03 — memory, chip restrictions, and bandwidth.
- 5. Artificial Analysis open-model index, accessed 2026-09-03 — current Intelligence Index scores.
- 6. Artificial Analysis methodology, accessed 2026-09-03 — index composition and limits.
- 7. Qwen3.8-27B model card, accessed 2026-09-03 — architecture, modalities, license, and task scores.
- 8. Qwen3.6-35B-A3B model card, accessed 2026-09-03 — architecture and license.
- 9. Qwen3.8-Flash-Next announcement, 2026-08-26 — component and active-parameter counts.
- 10. OpenAI gpt-oss announcement, 2025-08-05 — model release and license.
- 11. Mistral Small 4 announcement, 2026-03-16 — architecture and modalities.
- 12. GLM-5.3-Flash launch, 2026-09-02 — formal launch, multimodality, and benchmarks.
- 13. GLM-5.3 model card, accessed 2026-09-03 — parameters, license, and task results.
- 14. Kimi K3 announcement, 2026-07-16 — model facts and modalities.
- 15. DeepSeek V4 Pro release, 2026-08-13 — release and model identity.
- 16. oMLX Qwen3.8 field report, 2026-08-18 onward — mini and M3 Ultra decode, batching, and co-residency.
- 17. Verified oMLX gpt-oss-120b result, 2026-03-08 — corrected M3 Ultra prefill, decode, and batch values.
- 18. oMLX Qwen3.5-35B M5 Max proxy, 2026-03-09 — same-chip prefill, decode, and batching.
- 19. oMLX GLM long-context report, 2026 — long-document latency and peak memory.
- 20. llama-server documentation, accessed 2026-09-03 — router mode.
- 21. oMLX repository, accessed 2026-09-03 — multi-model engine pool.
- 22. LM Studio auto-evict documentation, accessed 2026-09-03 — multi-resident behavior.
- 23. vLLM-MLX model registry, accessed 2026-09-03 — multi-model registry and memory warnings.
- 24. K2 Horizon announcement, 2026-09-03 — post-August-15 release check.
- 25. Current local-model landscape, 2026-09-02 — shipped quant and tier evidence.
- 26. Mac generations side-by-side, 2026-09-03 — shipped Apple price and performance evidence, with the P06 correction noted above.
NOT_RUN
- not run Any M5 Pro mini, M5 Max Studio, or M5 Ultra Studio benchmark; the desktops are not delivered yet.
- not run GLM-5.3 or GLM-5.3-Flash on the target Apple machines.
- not run Quantized-model quality comparison against the full-precision Artificial Analysis scores.
- not run Full 32k batch 1/2/4/8 tests for every model.
- not run New-machine SSD load tests.
- not run UVU institutional price quote.
- not run 512GB M5 Ultra price; Apple has not posted it.
- not run Power, thermal, failure-rate, or 24-hour mixed-traffic tests.
- not run Contact with UVU, Apple, a model maker, or any vendor.
- not run Purchase, deployment, publication, or live-system change.
The data behind this appendix
In the underlying file, memGB is the practical weight of a model plus one reserved 32,000-token cache, and a missing prefill figure means unknown. Every conversation count in it is a planning estimate at the benchmark settings described above.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/38-model-machine-matrix.md in the research pack.
Service engineering
Appendix R in five lines
- QuestionHow should the service handle seven kinds of work, and what replaces four conversations per box?
- AnswerCold documents and coding agents fail first, so the service needs separate lines, saved document work, strict agent limits, and overflow.
- Deciding numbers8 chat conversations per 48-gigabyte miniest2 cold or 4 warm document conversations per miniest1 agent session per miniest
- What the plan doesMake separate queues, index each document once, reuse saved work on the same box, limit agents, and test four new nodes after September 22.
- Still unknownActual university traffic, saved-work reuse, exact M5 capacity, the four-node test, speech and image tests, and cloud overflow are NOT_RUN or UNKNOWN.
A campus AI service does not carry one kind of traffic. This appendix builds the engineering view the plan lacked: seven workload classes with measured token profiles, which ones are bound by writing speed and which by prompt reading (Apple's weak spot), the capacity multipliers available in today's Apple serving software and what they actually measured, proposed service levels per class, a queue model that replaces the plan's fixed 30-second service time, what coding agents do to a fleet, and a replacement for "four conversations per box". Research cutoff September 3, 2026. fact measured · est derived · unknown not yet measurable on the unshipped M5 desktops.
What this changes in the plan
1. Documents are the workload that breaks first. A 30-page upload read cold by the 27B daily model takes about 69–167 seconds to the first word on a mini (36–76 on a 128GB Studio); the sparse 35B fast model does it in 10–25 seconds. So the service indexes every upload once, retrieves 2,000–6,000 relevant tokens per question instead of replaying the file, caches exact course prefixes, keeps a student's later questions on the box holding that cache, and sends cold large prompts to the fast lane — never the dense model. 2. The multipliers are real but must not be multiplied together. Continuous batching: 2.4–2.6× aggregate throughput at four to eight conversations (measured); reusing a cached prefix: a 115k-token prefix went from about 20 minutes cold to 43 seconds warm (measured); speculative decoding: 2.3–2.6× for one conversation on the exact model but only 1.06× at batch four; 4-bit versus 8-bit compression: 1.75 quality points lost for 36% more speed and 46% less memory — use 4-bit for the daily lane and 6-bit where writing or code correctness matters. 3. The fixed 30-second service time understates a finals-week mix by about 70% (weighted mean 51 seconds; 93 occupied conversations instead of 54 for the same arrivals). That sits inside the plan's 9× stress envelope, so the fleet sizes stand — but the reason is now stated honestly. 4. Coding agents change everything if admitted freely. One agent turn consumes about 580,000 prompt tokens (94% cached) — roughly 1,200× a chat; if 5% of users start agents in the peak hour, a 26-mini fleet runs at 133% utilization and p95 waits reach 16 minutes. Agents therefore get their own queue, a metered allowance (illustrative: 64K in + 8K out per agent-active user-day, about 0.84× the whole chat baseline), session affinity so caches survive, and overflow to the metered frontier pool before utilization reaches 85%. 5. "Four per box" becomes a table: chat 8 / coding 4 / documents 2 cold or 4 warm / agents 1 / research writing 2 per 48GB mini (Studios about double on the daily model); speech and image work are measured in seconds per job, not conversations. Every number is an admission cap pending the four-node acceptance test after September 22: cold and warm prompts, 1/2/4/8 concurrency, p50/p95 time-to-first-token, per-user and aggregate speed, cache reuse, memory, thermals. est
Workload classes
“Request” means one external user turn unless noted. These are planning inputs, not measured UVU usage.
| Workload | Input per request | Output per request | Typical context | Requests per user-day | Evidence |
|---|---|---|---|---|---|
| Tutoring/chat | est 3,000 processed tokens | fact 1,115 mean tokens | est 6,400 tokens | fact (derived) 7.0 per weekly-active-user calendar day | Cedarville reported fact 49 interactions/week among active users. ShareChat reported fact 135-token user messages, fact 1,115-token assistant messages and fact 4.62 turns/conversation. Cedarville, Summer 2026, ShareChat, revised 2026-05-17 |
| Coding help | est 4,000 tokens | est 800 tokens | est 8,000 tokens | est 8 per coding-active day | StudyChat measured fact 2,214 conversations, fact 16,851 student utterances, and fact 7.6 utterances/conversation from fact 203 students. Token sizes were unknown. StudyChat, revised 2026-03-04 |
| Document questions | est 4,000 retrieved tokens/question; cold file est 6,667–66,667 tokens for fact supplied 10–100 pages at est 500 words/page | est 500 tokens | est 16,000 tokens | est 4, rounded from fact (derived) 3.70 | Prof. Leodar recorded fact 12,330 queries from fact 34 consenting students over fact 14 weeks. Frontiers, 2026-06-16 |
| Agentic sessions | fact 582,500 cumulative prompt tokens/external turn, including fact 545,300 cached | fact 4,000 cumulative tokens/turn | fact about 68,000 prompt tokens/internal call | fact (derived) 4.2455 external turns/calendar day | GitHub’s production trace covered fact 3.2M users, fact 13.5M sessions, fact 95.1M turns, fact 760.5M model calls, and fact 774.7M tool calls in fact one week. Agentic Coding in the Wild, 2026-07-30 |
| Research writing | est 2,000 tokens | est 1,000 tokens | est 16,000 tokens | fact 8.5 prompts per controlled writing-task day | A fact 45-minute study with fact 60 students recorded fact 512 prompts; actual API token counts were unknown. Frontiers, 2026-03-19; corrected 2026-05-11 |
| Lecture transcription | Input is audio, so text input tokens are UNKNOWN/not applicable | est 5,260–7,380 transcript tokens/lecture | fact 30-second audio windows | unknown; planning placeholder est 1 per submitting-user day | Medical lectures measured fact 30.71–45.71 minutes and fact 3,945–5,535 words. Whisper processes fact 30-second chunks. Medical lecture study, 2025-02-04, Whisper, 2022-09-21 |
| Image tasks | fact 256–1,280 visual tokens/image plus est 50 text tokens | Image generation tokens are UNKNOWN/not applicable; image understanding est 250 text tokens | est 2,000 working tokens; model maximum fact 32,768 | unknown; planning placeholder est 2 per image-active day | Qwen2.5-VL’s model card supplies visual-token geometry, not campus demand. Qwen2.5-VL, January 2025 |
The page conversion uses fact approximately 1 token per 0.75 English words, with model and language variation. OpenAI token guidance, accessed 2026-09-03
For document service, index the upload once and retrieve est 2,000–6,000 relevant tokens/question. Replaying the entire est 6,667–66,667-token file on every question is not a workable default.
Bound analysis on Apple Silicon
The planned M5 Pro mini and M5 Max Studio do not ship until fact 2026-09-22. Their desktop serving performance is therefore unknown on fact 2026-09-03. The estimates below use shipped M4 desktop results and an explicit est +10% M5 planning credit supported by same-build laptop-chip comparisons. Apple Mac mini announcement, 2026-08-25, Apple Mac Studio announcement, 2026-08-25
All bound classifications below are est engineering classifications.
| Class | 48GB mini | 64GB mini | 128GB Studio | Main risk |
|---|---|---|---|---|
| Tutoring/chat | Decode-bound | Decode-bound | Decode-bound | Long answers and too many active decoders |
| Coding help | Mixed; prefill-bound when code/repository text is pasted | Mixed, with more KV headroom | Mixed; usually decode-bound on short help | Repository-scale prompts |
| Document questions | Prefill-bound cold; mixed warm | Prefill-bound cold; extra RAM does not materially raise bandwidth | Prefill-bound cold; mixed warm | Full-file replay and cache misses |
| Agentic sessions | Prefill/cache-bound | Prefill/cache-bound | Prefill/cache-bound; decode also matters with a heavy model | Repeated internal calls and resident KV state |
| Research writing | Mixed; output-heavy after sources load | Mixed | Mixed/decode-bound | Long source packs and long output |
| Lecture transcription | Audio encoder-bound, not text decode-bound | Same | Same | Real-time factor and audio memory |
| Image tasks | Vision-prefill or diffusion-bound | Same | Same, with more model headroom | Image size, model type and GPU memory |
Cold first token for a 30-page upload
est 30 pages × 250–500 words/page ÷ 0.75 words/token = 10,000–20,000 prompt tokens. The range excludes parsing, OCR, queueing and retrieval.
| Box | Dense 27B | Sparse 35B-A3B MoE | Result |
|---|---|---|---|
| Shipped M4 Pro 48/64GB proxy | est 76–184 seconds | est 11–28 seconds | Prefill-bound |
| Planned M5 Pro 48/64GB | est 69–167 seconds | est 10–25 seconds | Direct result unknown |
| Shipped M4 Max 128GB proxy | est 39–84 seconds | est 9–17 seconds | Prefill-bound |
| Planned M5 Max 128GB | est 36–76 seconds | est 8–16 seconds | Direct result unknown |
The measurements behind those estimates include:
- M4 Pro/48GB dense Qwen3.8-27B: fact 126.1 prompt tokens/s at 16K. oMLX, 2026-08-27
- M4 Pro/48GB sparse Qwen3.5-35B-A3B: fact 890.2 prompt tokens/s at 8K, fact 830.3 at 16K, and fact 713.2 at 32K. oMLX, 2026-04-07
- M4 Max/128GB dense Qwen3.8-27B: fact 255.9 prompt tokens/s at 16K and fact 238.7 at 32K. oMLX, 2026-08-19
- M4 Max/128GB sparse Qwen3.5-35B-A3B: fact 1,160 prompt tokens/s at 32K and fact 28.247-second TTFT for 32,768 tokens. oMLX, 2026-03-17
The router should therefore:
- 1. Put short chat and short coding turns in a fast lane.
- 2. Parse and index uploads before the first question.
- 3. Preserve exact course prefixes and route later turns back to the box holding that cache.
- 4. Use the sparse MoE lane for cold large prompts.
- 5. Send unusually large, uncached or deadline-sensitive work to the metered cloud pool.
Capacity multipliers available today
Serving-stack support
| Stack | FACT support available by 2026-09-03 | Measured gain | Planning treatment |
|---|---|---|---|
| llama-server | Parallel slots, continuous batching enabled by default, prompt caching, cache reuse, quantized KV, and several speculative modes | Controlled Apple batching/cache gain unknown; one M1 Max MTP report measured fact 11%–28% slower | Use batching, but give speculation est 1.0× capacity credit until the exact model passes an A/B test. Server docs, MTP issue, 2026-05-27 |
| MLX-LM | Continuous-batch server, separate prompt/decode concurrency, chunked prefill, LRU prompt cache and draft models | Official controlled Apple batch multiplier unknown | Current code makes the draft-model path non-batchable, so batch and speculation gains must not be multiplied. MLX-LM server, accessed 2026-09-03 |
| LM Studio / MLX engine | Draft-model speculation, disk-backed KV checkpoints and continuous VLM batching | Short-chat version comparison: fact 2.2× output gain. Repeated image: fact 3.5× faster. Long-prompt extra RAM: fact 82% lower | Useful evidence for caching and batching, but not a universal stream multiplier. LM Studio, 2026-06-05 |
| oMLX | Continuous batching, prefix-cache preservation, SSD cache, Lightning MTP and DFlash | Exact-model MTP: fact 2.33×–2.62× single-stream gains; deeper MTP gave est 1.50× at one stream but only est 1.06× at batch four | Best current evidence for these models. Test model by model. oMLX releases, 2026-08 |
| vLLM-MLX | Community vLLM-style stack with continuous batching, paged KV, prefix sharing and SSD cache | M4 Max/128GB Qwen3-30B-A3B: fact 98.1 → 233.3 aggregate tokens/s, or fact 2.38×, at fact five requests | Promising pilot option; production reliability remains unknown. Guide, accessed 2026-09-03 |
What batching does to conversations per box
| Measured proxy | Single decode | Batched aggregate | Per stream | Reading of current est 4/box |
|---|---|---|---|---|
| M4 Pro/48GB, sparse 35B-A3B | fact 66.9 t/s | fact 163.8 t/s at eight-wide | est 20.5 t/s | Conservative for short decode-heavy work |
| M4 Pro/48GB, dense 27B | fact 15.4 t/s | fact 59.4 t/s at four-wide | est 14.9 t/s | Fair-to-conservative |
| M4 Pro/64GB, dense 27B with MTP | fact about 20.8 t/s at 8K | fact about 120 t/s at eight-wide | est 15.0 t/s | Conservative for short work |
| M4 Max/128GB, dense 27B | fact 29.0 t/s | fact 217.4 t/s at eight-wide | est 27.2 t/s | Very conservative when the Studio runs the daily model |
| M4 Max/128GB, sparse 35B-A3B | fact 122.2 t/s | fact 267.8 t/s at four-wide | est 67.0 t/s | Very conservative for decode-only work |
These are throughput tests, not service-level tests. A machine can show higher aggregate output while cold prompts wait too long.
Prefix/KV reuse
Caching has no cold-request gain.
- LM Studio measured fact 3.5× on a repeated image prompt with fact 3,584 cached and fact 145 uncached tokens.
- An oMLX field report reduced an approximately fact 115K-token agent prefix from fact about 20 minutes cold to fact about 43 seconds warm; the arithmetic ratio is est about 28×. oMLX field report, 2026-08-18/20
- Exact course-material capacity gain is unknown until UVU measures reusable-prefix share, hit rate, restore cost and session-affinity failures.
Quantization
Official MLX results below used the same Qwen3-30B-A3B model on M4 Max/64GB. Quality is MMLU-Pro, not the Artificial Analysis Index. MLX benchmark, build dated 2025-10-08
| Precision | Quality | Decode | Memory | Change from q8 |
|---|---|---|---|---|
| q8 | fact 72.46 | fact 83.16 t/s | fact 33.46GB | Baseline |
| q6 | fact 72.41 | fact 94.14 t/s | fact 25.82GB | est -0.05 quality point, est +13.2% decode, est -22.8% memory |
| q4 | fact 70.71 | fact 113.33 t/s | fact 18.20GB | est -1.75 quality points, est +36.3% decode, est -45.6% memory |
Recommendation: use q4 for the high-volume daily lane only after course tests show acceptable answers. Use q6 where writing, code correctness or nuance warrants the small measured quality advantage. The stream-count multiplier from bit width alone is unknown.
Service levels
All targets are est proposed service levels, not current UVU commitments.
| Class | TTFT or first-result target | Queue budget | Speed target |
|---|---|---|---|
| Tutoring/chat | est p95 ≤10 seconds | est p95 ≤2 seconds | est ≥15 output t/s |
| Coding help | est p95 ≤20 seconds | est p95 ≤3 seconds | est ≥15 output t/s |
| Document questions | est upload acknowledgment ≤2 seconds; cold answer est ≤30 seconds; warm answer est ≤10 seconds | est p95 ≤5 seconds | est ≥12 output t/s |
| Agentic session | est acknowledgment ≤2 seconds; first model action est ≤30 seconds; progress every est 15 seconds | est start wait ≤10 seconds | est ≥10 output t/s per model call |
| Research writing | est p95 ≤30 seconds | est p95 ≤5 seconds | est ≥12 output t/s |
| Lecture transcription | est acknowledgment ≤2 seconds; start est ≤30 seconds; finish est ≤0.5× recording length | Async | Text t/s not applicable |
| Image task | est acknowledgment ≤2 seconds; first result est ≤60 seconds | est start ≤15 seconds | Text t/s not applicable |
Replacement for the fixed 30-second service time
Use:
slot occupancy=prefill and scheduling time+ {output tokens{decode rate
For planning, model each class as an est lognormal service-time distribution:
| Class | EST mean | EST coefficient of variation | EST p95 |
|---|---|---|---|
| Tutoring/chat | 24 seconds | 0.60 | 51.2 seconds |
| Coding help | 66 seconds | 0.70 | 152.8 seconds |
| Chunked document QA | 94 seconds | 0.80 | 233.4 seconds |
| Agentic session | 536 seconds | 1.00 | 1,490.7 seconds |
| Research writing | 104 seconds | 0.80 | 258.3 seconds |
| Transcription scenario input | 225 seconds | 0.70 | 520.8 seconds |
| Image scenario input | 45 seconds | 0.50 | 87.5 seconds |
The supplied finals model gives est 6,534.553 requests/hour. An est finals class mix of chat 57.89%, coding 15.79%, documents 12.63%, research 8.42%, speech 2.11%, and image 3.16% produces:
- est weighted mean service = 51.11 seconds.
- est offered load = 92.76 occupied streams.
- The old est 30-second assumption produces only est 54.45 streams, understating this class mix by est 70.4%.
Small queue simulation
Method: est Poisson arrivals, class-specific lognormal service, one shared first-come queue, est 500 seeded one-hour runs. The method follows standard multi-server queue concepts. MIT queueing notes, Spring 2026
Because the brief does not define \(k\), this report defines it as est the number of Studios replacing minis in a fixed 26-node fleet. When every node runs the same approximately est 30B daily model:
EST capacity=4(26-k)+8k=104+4k
| Mix | EST capacity | EST utilization without agents | EST median-run p95 queue wait | EST jobs queued after hour |
|---|---|---|---|---|
| est k=0: 26 minis | 104 streams | 89.2% | 4.6 seconds | 0 |
| est k=4: 22 minis + 4 Studios | 120 streams | 77.3% | 0.0 seconds | 0 |
| est k=8: 18 minis + 8 Studios | 136 streams | 68.2% | 0.0 seconds | 0 |
| est k=13: 13 minis + 13 Studios | 156 streams | 59.5% | 0.0 seconds | 0 |
The est 26-mini case is stable, but its est 4.6-second p95 queue delay plus the measured-proxy fact 8.6-second short-prompt TTFT gives est 13.2 seconds, missing the proposed est 10-second chat target.
Finals with 5% starting agents in the peak hour
est 5% × 6,123 = 306.2 agent sessions/hour.
At est 536 seconds of slot occupancy, agents add est 45.59 occupied streams, taking total offered work to est 138.35 streams.
| Mix | EST utilization | EST median-run p95 wait | EST median queued after hour |
|---|---|---|---|
| est k=0 | 133.0% | 965.0 seconds or 16.1 minutes | 1,469 |
| est k=4 | 115.3% | 400.3 seconds or 6.7 minutes | 701 |
| est k=8 | 101.7% | 63.8 seconds | 100 |
| est k=13 | 88.7% | 0.0 seconds | 0 |
For est k=0, est k=4, and est k=8, utilization exceeds fact 100%, so no steady state exists; waits continue growing if the peak continues.
If est 306.2 daily agent users spread evenly over an est eight-hour day, agents add only est 5.70 streams. The est 26-mini fleet reaches est 94.7% utilization, but median-run p95 wait still rises to est 14.9 seconds.
If Studios are fixed to the plan’s est four-stream 120B role, replacing minis does not increase nominal capacity. No tested static \(k\) mix serves both the interactive queue and the concentrated agent queue. The service needs dynamic routing plus cloud overflow, not only a fixed hardware ratio.
What fails first
- 1. Cold document TTFT fails before total decode capacity: a mini proxy already measured fact about 75 seconds at an 8K prompt.
- 2. Agent admission makes the shared queue unstable.
- 3. Chat then waits behind multi-minute work.
- 4. KV memory and fresh prefill become the agent bottleneck.
- 5. A rigid Studio-only heavy lane can starve the mini chat lane.
Agentic and coding loads
A controlled SWE-bench study gives the cleanest chat comparison:
| Workload | Average total tokens/task | Input/output ratio |
|---|---|---|
| Code chat | fact 3,390 tokens | fact 1.33 |
| Agentic coding | fact 4.17M tokens | fact 153.85 |
That is fact (derived) 1,230×; the authors report approximately fact 1,200×. The same task varied by as much as fact 30× across runs, and higher token use did not reliably improve success. How Do AI Agents Spend Your Money?, revised 2026-04-29
The production GitHub trace adds:
- fact (derived) 7.04 user turns/session.
- fact mean 6.6 model calls/user turn.
- fact 582,500 prompt tokens/user turn, of which fact 545,300 were cached.
- fact (derived) 37,200 fresh prompt tokens/user turn.
- fact (derived) about 15.7× more fresh-prefill work if cache reuse is lost.
- fact mean 396.3 seconds/turn.
For the supplied est 6,123-user staff case:
- Agent users: est 306.2.
- Agent turns/day at the production rate: est 1,299.98.
- Agent token work: est 614.3M–762.4M tokens/day.
- Existing demand-model baseline: est 26.36M tokens/day.
- Added agents: est 23.3×–28.9× the entire baseline.
- Combined work: est 24.3×–29.9× baseline.
The agent paper contains conflicting aggregate and median values. Its exact median token count is therefore unknown; the range above uses its aggregate-derived lower case and Table 4 mean upper case.
Recommended policy:
- Cap locally admitted agent jobs.
- Meter cumulative input, cached input, output, model calls and elapsed time separately.
- Preserve session affinity and cache state.
- Give agents their own queue.
- Route overflow and very long jobs to the metered frontier pool.
- Start with an illustrative est one 64K-input plus 8K-output local agent allowance per agent-active user-day. At est 306.2 users, that is est 22.05M tokens/day, or est 0.84× the existing baseline. Change the cap only after measured local traces.
Audit of the plan’s number
The audited number is in data.js: est four simultaneous streams at at least 10 t/s.
Replacement counts below are est admission caps pending a delivered-M5 acceptance test. Studio counts assume the daily est 30–35B model unless noted.
| Class | Current 4/box judgment: 48GB mini | 64GB mini | 128GB Studio | Replacement |
|---|---|---|---|---|
| Tutoring/chat | Conservative | Conservative | Conservative with daily model; fair with heavy model | est 8 / 8 / 8 |
| Coding help | Fair short; optimistic at repository scale | Fair | Conservative short; fair long | est 4 / 4 / 8 |
| Document QA | Optimistic cold; fair warm | Optimistic cold; fair warm | Fair when chunked/warm | est 2 cold or 4 warm / 2 cold or 4 warm / 4 |
| Agentic sessions | Optimistic | Optimistic, but extra RAM helps KV | Optimistic or unknown with 120B model | est 1 / 1–2 / 2–4 |
| Research writing | Fair-to-optimistic | Fair | Fair | est 2 / 4 / 4 |
| Lecture transcription | Invalid unit | Invalid unit | Invalid unit | unknown; separate ASR benchmark |
| Image tasks | Invalid unit | Invalid unit | Invalid unit | unknown; separate vision/image benchmark |
Even if concurrency stays at est four, one mini’s request rate varies from:
- Chat: est 600 requests/hour.
- Coding: est 218.2/hour.
- Document QA: est 153.2/hour.
- Research writing: est 138.5/hour.
- Agent sessions: est 26.9/hour.
That is an est 22.3× span hidden by the phrase “four conversations.”
What the plan should change
- 1. Replace the single stream count with the class table above. Size the fleet from \(\sum \lambda_jE[S_j]\), p95 TTFT and per-user decode speed.
- 2. Create separate queues for interactive text, documents/research, agents/API, and speech/image. Protect chat from long work.
- 3. Cap and meter agents from launch. Spill excess jobs to the frontier pool before projected utilization reaches est 85%.
- 4. Index documents once. Use retrieval, exact-prefix caching and box affinity. Do not replay whole files.
- 5. Require a mixed-load acceptance test on est four delivered M5 nodes after fact 2026-09-22: cold and warm prompts, est 1/2/4/8 concurrency, p50/p95 TTFT, per-user and aggregate output, cache reuse, memory, thermals and failures.
- 6. Treat continuous batching as a measured curve. Do not multiply batching, cache, speculation and quantization gains.
- 7. Start with q4 for the daily lane and q6 for quality-sensitive work, subject to UVU course evaluations.
- 8. Collect per-class UVU telemetry before moving from the pilot to the est 26-node fleet.
Source log
- 1. fact accessed 2026-09-03 — Supplied capacity assumptions, data.js
- 2. fact dated 2026-09-02 — Supplied demand stress finding
- 3. fact dated 2026-09-03 — Supplied hardware comparison
- 4. fact 2026-06-23 / Summer 2026 — Cedarville campus AI results
- 5. fact revised 2026-05-17 — ShareChat telemetry
- 6. fact published 2024-07-09 — Cipherbot student-use study
- 7. fact revised 2026-03-04 — StudyChat coding-help telemetry
- 8. fact published 2026-06-16 — Prof. Leodar RAG tutor telemetry
- 9. fact published 2026-03-19; corrected 2026-05-11 — Academic-writing study
- 10. fact submitted 2026-07-30 — Agentic Coding in the Wild
- 11. fact revised 2026-04-29 — Coding-agent token-cost study
- 12. fact published 2025-02-04 — Medical lecture transcript study
- 13. fact published 2022-09-21 — Whisper architecture
- 14. fact January 2025; accessed 2026-09-03 — Qwen2.5-VL model card
- 15. fact accessed 2026-09-03 — OpenAI token-counting guidance
- 16. fact 2026-08-25 — Apple Mac mini announcement
- 17. fact 2026-08-25 — Apple Mac Studio announcement
- 18. fact 2026-03 through 2026-08 — oMLX public performance database
- 19. fact 2026-08-18/20 — oMLX Apple serving field report
- 20. fact accessed 2026-09-03 — llama-server documentation
- 21. fact accessed 2026-09-03 — MLX-LM server implementation
- 22. fact benchmark build dated 2025-10-08 — MLX-LM quantization benchmarks
- 23. fact 2026-06-05 — LM Studio MLX batching and cache measurements
- 24. fact accessed 2026-09-03 — vLLM-MLX continuous-batching guide
- 25. fact Spring 2026 — MIT queueing-model notes
NOT_RUN
- not run Physical M5 Pro mini or M5 Max Studio benchmark: hardware availability begins fact 2026-09-22.
- not run UVU trace replay: no class, token, upload, cache, queue or peak trace was supplied.
- not run Four-node acceptance test.
- not run Token-scheduler-level continuous-batching simulation; virtual streams are an approximation.
- not run ASR and image-generation benchmarks on the planned boxes.
- not run Cloud-valve quota, privacy, latency or failover test.
- not run Like-for-like Artificial Analysis quantization comparison; official task benchmarks were used instead.
- not run Academic-research helper scripts because their local requests dependency was unavailable and the workspace prohibited installation.
- not run Contact with UVU or any vendor, as instructed.
- unknown Actual UVU class mix, student-agent behavior and cache-hit distributions.
- unknown Exact M5 chassis stream counts and exact 120B Studio capacity.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/39-service-engineering.md in the research pack.
Duty of care and resilience
Appendix S in five lines
- QuestionWhat must be ready before the service handles minors, patient data, crisis messages, or a system failure?
- AnswerMinors and real patient data stay out until privacy, human help, security, and recovery checks all pass.
- Deciding numbers18,163 concurrent-enrollment studentsfacteleven named conditionsestseventeen security gapsest
- What the plan doesStart with practice health cases, block outside tools and web searches, provide a real human crisis path, and test backup routes and the incident guide.
- Still unknownStudent ages and the live system design are UNKNOWN; legal review, system checks, and attack, failover, restore, and crisis drills are NOT_RUN.
What happens when a 16-year-old concurrent-enrollment student uses the tutor, when a nursing student pastes a patient's chart, when someone types that they want to hurt themselves, when an uploaded file carries hidden instructions, or when the router dies during finals? The plan had a routing rule and an isolation test; it did not answer those questions. This appendix does, from the statutes, federal guidance, and named campus examples. Research cutoff September 3, 2026. It is planning research, not legal advice. fact dated law and guidance · est design controls and targets · unknown what only UVU's counsel or records can settle.
What this changes in the plan
A. Minors and real health data stay closed until specific gates pass. A concurrent-enrollment student controls their own UVU record even under 18 (FERPA), while survey-type questions stay under parental rights until 18 (PPRA); Utah's student-data law governs the school-district side; COPPA reaches only under-13s. Phase two opens only when eleven named conditions pass — parental permission alone is not one of them. Placement patient records never enter the service (the clinical host's HIPAA duties continue); UVU's own clinics and counseling are FERPA records; the service uses synthetic cases, and real patient data fails closed at the router. UVU's public pages disagree about whether its own health records are FERPA or HIPAA — a question for UVU, not a finding of fault. B. A crisis handoff is built before broad student use. The bot is not an employee, a therapist, or a campus security authority, so Title IX, Clery, and Utah duty-to-warn do not attach to it — they attach to the trained person who receives the alert (and Utah's child-abuse reporting duty attaches to that reviewer). Design: a separate detector, an honest message ("I'm an AI tutor, not a crisis service… call or text 988"), an acknowledged human queue rather than an unread email, a restricted case record, graduated handling for low-confidence signals, and no discipline on classifier output alone. C. The first pilot removes the dangerous features: no tools, no web fetching, no shared document retrieval, no automatic cloud spill, outbound network denied from inference and parsing workers. D. "Content logging off" becomes a verified data map covering histories, caches, retrieval stores, telemetry, backups, support access, incident capture, retention, and deletion. E–I. The control plane gets finished — redundant routers, database and queue recovery, identity-outage behavior, model mirrors, a same-class spare for the one Ultra, tested recovery targets (a failed mini: 5 minutes; total local loss for public data: 30 minutes via the metered route; the restricted route: fails closed, 4 hours), a change freeze from seven days before finals through the grade deadline, a model admission record with pinned hashes and licenses, named owners (a service owner, a cross-trained backup, central IT after hours), and an incident runbook that is a launch gate. Twenty controls are graded: three exist in the plan, seventeen were gaps, each now with a fix and a cost class. est
Research date: FACT — September 3, 2026 MDT.
Statute, policy, model, section, emergency-service, and source-log numbers are identifiers. Every measured quantity below is marked fact, est, or unknown.
Minors
Who is in scope
fact — 18,163 concurrent-enrollment students, dated September 2, 2026. This comes from the supplied engine assumptions. Their age distribution is unknown. The plan’s decision to keep them out of the first phase is sound.
What the laws require
FERPA
A student attending UVU controls their UVU education records even when the student is younger than eighteen. The parent keeps FERPA rights over the high-school record and can inspect UVU records that UVU sends back to the school. UVU may disclose records to a tax-dependent student’s parent under a FERPA exception, but it is not required to do so. A parent’s concurrent-enrollment permission is not FERPA permission to inspect the student’s UVU chat record. U.S. Department of Education dual-enrollment guidance, updated August 2026.
An AI operator receiving education records without student consent must fit FERPA’s school-official exception: it must perform a real institutional function, remain under UVU’s direct control over record use and maintenance, use the information only for the approved purpose, observe redisclosure limits, and meet UVU’s annual-notice criteria. A privacy statement alone is insufficient. FERPA school-official guidance.
PPRA
fact — federal rule current through 2026: PPRA rights normally remain with a parent until the student is eighteen or emancipated. Postsecondary attendance by itself does not transfer those rights.
A normal open-ended tutoring chat is not automatically a PPRA survey. Risk arises when UVU or a participating school requires the bot to ask, analyze, or evaluate protected subjects such as mental health, sexual behavior, illegal or self-incriminating conduct, religion, family relationships, privileged relationships, or income. Required activities can require prior written parental consent; some other school-administered activities require notice and an opt-out. Each intake form, tutor script, wellness check, evaluation, and research use must therefore receive a PPRA classification. Department of Education PPRA guidance.
Utah Student Data Privacy Act, Title 53E Chapter 9
The main “education entity” definition covers Utah’s state and local public-school bodies, not UVU. It nevertheless applies to the school-district side of concurrent enrollment. Among other things, it governs parent notice, optional-data consent, disclosure, contractor controls, and protected-topic examinations or surveys.
The Utah law requires parental consent for specified disclosures of an under-eighteen pupil’s defined educational data and creates a process for school-to-higher-education data sharing. UVU and each participating school must agree which organization owns each record and which rule controls it. Current Utah Code Title 53E Chapter 9.
UVU also has direct duties under Utah’s higher-education student-data law. Those duties cover contracted purposes, collection, storage, sharing, deletion, audits, secondary uses, affiliates, advertising, and sale. Utah Code Title 53H Chapter 14 Part 5.
COPPA
fact — federal threshold current September 3, 2026: COPPA concerns covered online operators collecting personal information from children younger than thirteen. It does not cover every minor. A commercial AI vendor may still be covered even if UVU itself is outside the FTC’s ordinary nonprofit jurisdiction.
A school may consent for a covered operator only for a school-authorized educational use, with no separate commercial purpose. The operator remains responsible for notice, security, collection limits, review, deletion, and stopping further use. FTC COPPA guidance.
What peer institutions do
fact — 2025–2026 practice: Utah’s concurrent-enrollment process uses parent permission for participation, while UVU’s FERPA-protected college records remain controlled by the student. Weber State follows the same student-owned-record approach. Existing concurrent-enrollment permission should not be treated as consent for AI prompt processing. UVU concurrent-enrollment information and Weber State parent information.
Salt Lake Community College uses a student-and-guardian agreement while separately telling families that college FERPA rights belong to the student. The University of Arizona and Worcester Polytechnic Institute likewise state that college records belong to the student regardless of age. Arizona also tells users to use vetted enterprise AI, remove personal information, and avoid putting student records into external AI tools. SLCC agreement, Arizona FERPA guidance, and WPI FERPA guidance.
Age-appropriate design
These are est design controls, based on UNICEF’s fact December 2025 and June 2026 guidance, not separate Utah legal mandates:
- Give recurring, plain-language notice that the service is AI.
- Use maximum privacy settings by default.
- Collect an age band rather than a full birth date unless exact age is necessary.
- Do not present the tutor as a therapist, friend, confidant, or exclusive relationship.
- Do not use emotional dependency, streaks, guilt, or re-engagement nudges.
- Keep sexual, violent, manipulative, and self-harm testing in the release gate.
- Give an equal non-AI route without penalty.
- Explain before use that narrow safety content may be sent to trained people.
- Let students see, correct, export, and delete records where applicable.
- Write notices separately for students and parents; do not hide them inside general terms.
UNICEF Guidance on AI and Children.
Exact conditions for opening phase two
These are est launch conditions. All must pass:
- Population gate: Block children younger than thirteen unless a documented COPPA path passes. Confirm vendor minimum-age terms.
- Rights-holder matrix: State when the student, parent, school district, or UVU controls consent and access. Do not give parents automatic access to UVU chat records.
- Separate permission: Obtain clear student opt-in and the promised parent or guardian permission. State that this permission does not transfer the student’s FERPA rights.
- Data map: List every prompt, output, identity field, history item, safety event, retrieval item, cache, log, backup, and cloud transfer.
- PPRA screen: Remove or separately approve required questions that function as protected-topic surveys, analyses, or evaluations.
- Contract pass: Prohibit model training, targeted advertising, sale, profiling, unrelated product improvement, and unauthorized redisclosure. Name subprocessors, deletion, audit, breach, access, and export duties.
- Age-safe product pass: Test the exact interface and model for manipulation, sexual content, self-harm, bias, jailbreaks, false crisis alerts, and age-gate bypass.
- Health exclusion: Prohibit real patient information, counseling notes, disability files, and real clinical cases.
- Logging proof: Prove prompt text is absent from ordinary logs, caches, backups, telemetry, and support paths while keeping useful security metadata.
- Human safety route: Prove that a serious alert reaches an acknowledged trained person rather than an unattended form or email.
- Governance: Obtain written approval from UVU counsel/privacy, Registrar, CISO, accessibility, the institutional data owner, and each participating school body.
Phase two cannot open on parental consent alone.
Health data
When HIPAA applies
Health-related words do not make a record subject to HIPAA. HIPAA applies when a covered health plan, clearinghouse, covered provider, or business associate creates, receives, maintains, or transmits protected health information in the covered role.
| Context | Governing boundary |
|---|---|
| Synthetic or properly de-identified teaching case | Usually not HIPAA PHI. FERPA, ethics, license, and re-identification controls may still apply. |
| Real patient information from a clinical-placement host | The host’s HIPAA duties normally continue. Placement access does not authorize copying the patient’s information into UVU’s AI. |
| UVU student clinic or student counseling record maintained by UVU | Generally a FERPA education or treatment record excluded from HIPAA PHI. |
| Nonstudent patient treated by a covered university clinic | Normally HIPAA. |
| University hospital treating people without regard to student status | Normally HIPAA. |
| Student disability-services record | Normally FERPA plus disability-law confidentiality and need-to-know controls. |
| Employee accommodation record | ADA confidentiality applies; HIPAA depends on where the record came from and which entity maintains it. |
HHS expressly says postsecondary student clinic and counseling records are generally on the FERPA side, even when the university is also a HIPAA covered entity. Nonstudent clinic records can remain HIPAA records. HHS FERPA/HIPAA clinic guidance.
UVU’s public pages create an unknown that must be resolved: Accessibility Services says its disability records are FERPA records, while Student Health says it operates under HIPAA. This may reflect voluntary practices, a hybrid-entity boundary, a covered transaction, or public wording; the reviewed evidence does not settle it. Do not label UVU noncompliant.
Clinical placements
Students may see patient records under the clinical host’s approved training workflow. That does not authorize transfer into a campus AI service.
Until each host approves otherwise:
- Students must not paste or upload patient information.
- UVU should use synthetic cases.
- Real patient data must fail closed at the router.
- A host-specific use would need authorization, minimum-necessary controls, security review, and possibly a business associate agreement.
HHS guidance on trainees’ access to patient information.
Business-associate question
A business-associate question arises when the service performs work for a HIPAA-covered component or clinical host and creates, receives, maintains, or transmits PHI.
- If an external AI or cloud provider processes PHI for that covered entity, it is generally a business associate or subcontractor and needs an appropriate agreement.
- A cloud provider storing encrypted PHI can still be a business associate even when it lacks the decryption key.
- If UVU components are inside one legal covered entity, the issue may instead be the university’s hybrid-component boundary and internal safeguards.
- If the record is a FERPA student record rather than HIPAA PHI, the HIPAA business-associate concept generally does not control. FERPA and Utah contract controls still do.
HHS now gives a third-party AI chatbot operating on a patient portal as an express business-associate example. HHS Business Associates, reviewed July 30, 2026.
What “content logging off” solves
It helps by reducing stored copies of conversations.
It does not solve:
- Whether the disclosure was authorized.
- Prompt and output data in memory.
- Browser history, screenshots, downloads, or device traces.
- Retrieval stores, embeddings, caches, queues, backups, and crash reports.
- Vendor telemetry, support access, or model-training use.
- Cross-user leakage.
- Role-based access, minimum necessary use, or deletion.
- Required business-associate terms.
- Security-event logging and incident evidence.
The plan should replace “content logging stays off” with a tested data-retention statement. Ordinary logs should omit raw content but retain request identifiers, pseudonymous identity, time, route, model and configuration hash, access decision, tool decision, error, safety event, and administrative action.
Crisis and duty of care
What is required
Title IX
fact — current federal position reviewed January 29, 2025: the federal government is enforcing the 2020 Title IX rule after the later rule was vacated. Federal “actual knowledge” generally requires notice to the Title IX coordinator or another university official with authority to act. An unread disclosure to a general tutoring bot is an unknown legal trigger.
UVU’s own policy is broader for people: an employee aware of covered conduct must report within fact twenty-four hours under the policy last reviewed October 9, 2025. The bot is not an employee. UVU must name the human who receives and acts on alerts. Federal Title IX status and UVU Policy 162.
Generic self-harm or violence is not automatically Title IX. The disclosure must involve covered sex discrimination, sexual harassment, or a listed offense.
Clery
Clery is not a general crisis-chat law. It applies to defined crimes, geography, and reports received by campus police, security, designated reporting offices, or officials with significant campus responsibility. Timely warnings and emergency notifications have different threat tests. Self-harm alone is not a listed Clery crime. A tutor bot is not automatically a campus security authority. Current Clery regulation.
A trained person must decide whether an alert creates a Clery crime-report, warning, emergency-notification, or no-Clery outcome.
Utah duty to warn and abuse reporting
Utah’s named duty-to-warn rule applies to specified licensed therapists when a client communicates an actual threat of physical violence against an identifiable victim. The therapist discharges the statutory duty by making reasonable efforts to warn the victim and notifying law enforcement. A tutoring bot is not a statutory therapist. Utah duty-to-warn statute.
fact — Utah law effective May 1, 2024: a person with reason to believe that a child is being or has been abused or neglected must report immediately to child protection or law enforcement, subject to narrow exceptions. A software classifier is not the reporting person. A human reviewer can acquire that duty after seeing the disclosure. Utah child-abuse reporting law.
FERPA emergency disclosure
FERPA permits disclosure to appropriate parties when necessary to address an articulable and significant threat. When UVU relies on that exception, it must record the threat basis and recipients. It is not a blanket reason to circulate every flagged chat. FERPA emergency-disclosure record rule.
Safe messaging pattern
est design based on NIMH, SAMHSA, and UVU crisis guidance:
I’m sorry you’re dealing with this. I’m an AI tutor, not a crisis service. If you or someone else may be in immediate danger, call 911 or UVU Police now. For suicide or emotional crisis, call or text 988. You do not need to repeat the details here.
The bot should:
- Stop ordinary tutoring after a serious flag.
- Use calm, direct, nonjudgmental language.
- Ask only the minimum needed to distinguish immediate danger from nonacute distress.
- Offer a trusted person and trained human help.
- Never promise confidentiality, active monitoring, dispatch, or a response time unless each promise is true.
- Say that it is sending an alert only after an acknowledged route exists.
- Avoid long disclaimers while a person may be in danger.
UVU already distinguishes immediate emergencies, crisis support, and nonemergency Behavioral Assessment Team reports. Ordinary email or an unacknowledged form is not an acute handoff. UVU crisis services.
What campus bots do today
- UNCW Sammy — FACT, checked September 3, 2026: says it is not for emergencies, directs users to emergency and crisis services, warns against sharing private information, and can connect unanswered ordinary questions to a person. Its crisis-alert backend is unknown. UNCW Sammy.
- University of Houston Wayhaven — FACT, launched July 23, 2025: is described as a bridge to care, not a clinician. The vendor says it screens messages, interrupts normal coaching, shows crisis resources, and can alert designated contacts. Houston’s exact alert settings and acknowledgement path are unknown. University of Houston Wayhaven.
- NJIT Charlie — FACT, published August 10, 2026: uses a platform whose default safety process emails designated contacts after certain risk flags. NJIT’s exact configuration and after-hours coverage are unknown. Email delivery alone is not human acknowledgement. NJIT Charlie.
Router behavior: detect → human handoff → record
Detect — EST
Use a separate safety gate, not the tutor’s ordinary judgment alone. Classify imminent self-harm, violence toward others, child or vulnerable-adult abuse, sexual misconduct, and nonacute distress separately. Combine rules with a tested classifier. Do not use keywords or one model alone.
Human handoff — EST
Send the smallest useful packet to an acknowledged human queue:
- Identity and contact details if known.
- Exact triggering text.
- Time and service route.
- Category and classifier version.
- The message shown to the student.
- Delivery receipt and acknowledgement state.
A trained person—not the model—decides clinical risk, police contact, Title IX, Clery, abuse reporting, and any outside disclosure. Immediate-risk routing needs a genuinely staffed destination. Lower-risk concerns can go to a normal student-support queue.
Record — EST
Keep a restricted safety-case record containing the trigger, response, recipient, acknowledgement, action, legal basis for any outside disclosure, and closure. Keep it separate from tutoring analytics. Do not retain the entire chat merely because one part created a case.
False-positive costs
No validated public precision or false-positive rates were found for the named campus systems. Those rates are unknown.
Likely costs are est:
- Student fear, shame, or loss of trust.
- Disclosure of private text to more people.
- Unnecessary police or emergency involvement.
- Alert fatigue that delays real emergencies.
- Unequal flagging of dialect, disability, quoted coursework, or culturally different speech.
- Students avoiding tutoring or counseling.
Use graduated handling: offer private resources for low-confidence signals, ask a short clarifying question for ambiguous risk, and reserve automatic human alerts for defined serious cases. Never use classifier output alone for discipline.
Security
Existing controls and present gaps
“Existing” below means present in the plan or supplied findings, not proven in a live deployment.
| Control | Status | What it covers | Gap and concrete fix |
|---|---|---|---|
| Isolation canary battery | Existing design; live proof not run | Cross-user leakage around caches and concurrent sessions | Extend testing through UI history, databases, retries, restarts, timeouts, node loss, RAG, and cloud failover. Drain the route on any canary leak. |
| Routing by data class | Existing design; live proof not run | Keeps sensitive and restricted requests on approved routes | Enforce the decision outside the model, show the destination before submission, use egress allowlists, and fail closed when identity or classification fails. |
| No content logging | Existing intent; actual state unknown | Reduces stored prompt and response copies | Inventory databases, histories, caches, traces, backups, crash reports, and support access. Prove raw text is absent while retaining useful security metadata. |
| Uploaded-document isolation | Gap | Malicious instructions, malware, parser exploits, and oversized archives | Quarantine files; allowlist necessary formats; verify real type; cap compressed and expanded size; scan or disarm active content; parse in a no-egress sandbox. |
| Retrieval isolation | Gap | Poisoned documents and cross-course or cross-user retrieval | Hash every source, record provenance, attach access rules to every chunk, separate indexes by data class, and recheck access at query time. |
| Tool control | Gap | Exfiltration or unauthorized action through connectors | Launch tutoring without tools. If tools are later added, use narrow read-only functions, per-user credentials, server-side authorization, parameter validation, and human confirmation for consequential actions. |
| Model supply chain | Gap | Altered weights, unsafe serialization, unsupported licenses, and compromised runtimes | Pin weights, runtime, container, prompt, and configuration by cryptographic hash. Save source, model card, license, required notices, conversion recipe, scan receipt, test receipt, and approver. |
| Abuse limits | Gap | Service denial, model extraction, queue starvation, and cloud overspend | Apply identity-based request, upload, context, output, concurrency, timeout, retry, tool-depth, and daily-cloud limits with a global breaker. Exact thresholds remain unknown until load testing. |
| Insider controls | Gap | Admin misuse, secret theft, hidden content access, or log deletion | Use unique administrator accounts, MFA, short-lived elevation, least privilege, separate approval and operating roles, protected central logs, and a second campus approver for releases and destructive administration. |
| Forensic logging | Gap | Incident detection and reconstruction | Record pseudonymous identity, time, route, data class, model/configuration hash, access decision, document hashes, tool decision, latency, error, safety event, cloud spend, and admin action. Keep raw content off by default. |
OWASP identifies indirect prompt injection, retrieval poisoning, excessive tool authority, sensitive-information disclosure, supply-chain weaknesses, and unbounded consumption as core AI-service risks. OWASP GenAI risks.
Main exfiltration paths
- Another user’s chat, cache, history, workspace, or retrieval index.
- A poisoned uploaded file or retrieved page.
- A tool using broad service credentials.
- Automatically fetched links, images, HTML, or URLs in model output.
- Outbound DNS or web traffic from the model or parser.
- Silent cloud overflow to an unapproved provider.
- Logs, traces, backups, crash reports, or administrative consoles.
- Model or system prompts containing secrets.
- A compromised weight, runtime, package, container, or conversion script.
- An insider with broad production and log access.
The smallest safe pilot has no external tools, no arbitrary web fetching, no shared retrieval uploads, and outbound access denied from inference and parsing workers.
Supply-chain gate
est required release record:
- Repository owner and immutable revision.
- Download source and file list.
- Cryptographic hashes.
- Signature status.
- Exact license and required notices.
- Model card, intended uses, and limits.
- Serialization format and conversion process.
- Scanner and evaluation receipts.
- Named approvers and approval date.
- Read-only local copy plus an independent mirror.
A matching hash proves byte identity. It does not prove safety, provenance, model quality, or legal permission. Exact license suitability remains unknown until the chosen models are reviewed.
Availability
Single points of failure
The exact deployed topology is unknown. The supplied red-team review identifies an incomplete control plane.
| Failure point | Result | Concrete fix |
|---|---|---|
| Router or ingress | Whole service becomes unreachable | Use redundant stateless router instances, independent health checks, and a tested failover address. |
| Identity provider | New sessions fail | Fail closed for new access. Consider only a short, CISO-approved grace period for already authenticated sessions. Keep a public status page outside the login path. |
| The single Ultra node | Heavy jobs stop | Keep a compatible model mirror and documented degraded route. Buy or qualify a same-class spare before promising heavy-route continuity. A mini is not an equivalent replacement. |
| Database, queue, or cache | Histories, routing state, or jobs fail | Replicate only required state, keep inference nodes stateless where practical, and test restore. |
| Network switch or uplink | Fleet becomes unreachable | Separate failure domains, use redundant network paths where available, and test a disconnected-node case. |
| Power or room | Every colocated Mac stops | Use UPS-backed circuits and a warm spare in another campus location if the recovery promise requires site resilience. |
| DNS, certificates, or secrets | Service fails despite healthy Macs | Monitor expiry, keep documented renewal and break-glass procedures, and test recovery. |
| One operator | Vacations or simultaneous incidents halt recovery | Cross-train a second campus operator and use central IT/security as the after-hours receiver. |
What the spare policy buys
fact — plan assumption dated September 2, 2026: sparePolicy is set to est fifteen percent, with a comment describing about one spare for every seven and a minimum of one per deployment. The rounding rule and whether “spare” means cold, warm, or unused live capacity are unknown. Engine assumption.
est arithmetic:
- One ready spare beside four active equal-capacity machines is twenty-five percent of active capacity.
- One machine among seven equal machines is about fourteen-point-three percent of installed capacity.
- A literal fifteen-percent margin may not equal a whole device in a small fleet.
The policy buys one-for-one protection only when the spare is the same class, configured, licensed, connected, patched, loaded with verified models, and regularly boot-tested. It does not protect against a router, identity, network, power, bad release, shared database, site, or staff failure. It also does not replace the unique Ultra node unless the spare is Ultra-class.
Fleet pattern
est recommended design:
- Run minis active-active behind redundant routers.
- Keep ordinary load low enough that one node can disappear without overload.
- Store approved weights on each compatible serving node.
- Keep a separate checksum-verified model mirror.
- Make inference nodes replaceable from an immutable manifest.
- Keep public-data cloud fallback separate from restricted local recovery.
- Never send sensitive or restricted traffic to cloud merely because a Mac failed.
- Prevent automatic replay of tool actions after a failed stream.
Backups and disaster recovery
Back up:
- Router and data-class policy.
- System prompts and safety messages.
- Model, runtime, container, and configuration manifests.
- Identity and access mappings.
- Retrieval source inventory and access rules.
- Security and incident metadata.
- Required database state.
- License and approval records.
Do not back up raw chats when the approved policy says they are not retained.
Keep an encrypted off-site copy of critical configuration and manifests, plus a separate model mirror. Test restoration into a clean machine. A backup without a successful restore test is not recovery evidence. NIST contingency guidance.
Proposed RTO and RPO
These are est targets pending a UVU business-impact review, load tests, and recovery drills.
| Failure | Proposed target and basis |
|---|---|
| One inference Mac or router instance | est RTO: five minutes. Basis: automated health removal and redundant capacity. The in-flight request may be lost. |
| Total local compute loss for public data | est RTO: thirty minutes. Basis: a pre-approved metered route with tested credentials and budget. |
| Restricted local route | Fail closed; est RTO: four hours. Basis: ready same-class spare and immutable rebuild. Current request may be lost. |
| Model or configuration compromise | est RTO: four hours; RPO: last approved manifest. Basis: clean image, pinned artifacts, and independent mirror. |
| Retrieval-index corruption | est RTO: four hours; RPO: twenty-four hours. Basis: daily protected index state while source documents remain authoritative. |
| Security-log collector | est RTO: four hours; RPO: five minutes. Basis: buffered or replicated event forwarding. |
| Campus-room loss | est RTO: one business day only if UVU maintains a separate-site warm spare and tested restore. Without those controls, RTO is unknown. |
Finals maintenance
est controls based on repair lead time, not published UVU service levels:
- Fourteen days before finals: test node loss, router loss, identity failure, model mirror, backup restore, cloud quota, and the physical spare.
- Seven days before finals through the grade deadline: freeze ordinary OS, model, prompt, retrieval, schema, router, and provider changes.
- Permit emergency security changes only through the canary and rollback path.
- Stop nonessential batch, conversion, and fine-tuning work.
- Check spare readiness, backup freshness, capacity, provider budget, and escalation contacts each day.
- Tune warning and overflow thresholds only after real load tests.
Incident response
One-page runbook outline
Detect
- Receive a canary failure, route anomaly, safety alert, user report, provider notice, capacity alarm, or administrator alert.
- Open a restricted incident record.
- Record MDT and UTC time, symptoms, request identifiers, affected accounts and nodes, route, data class, model hash, document hashes, and actions.
- Name one incident lead.
Contain
- Drain or quarantine affected nodes.
- Disable the affected upload, retrieval source, tool, model, cloud route, or account.
- Fail sensitive and restricted traffic closed.
- Stop abnormal metered spending.
- Revoke affected sessions and rotate exposed credentials.
- Preserve central logs and volatile evidence. Do not erase or casually restart a compromised system.
Notify
- Follow UVU Policy 447 by escalating suspected or actual security incidents to the CISO/CITRM path and appropriate data owner.
- Add privacy, General Counsel, Registrar, Title IX, Clery, campus safety, counseling, disability services, the school district, clinical host, or business associate only when the incident facts require them.
- Do not let the local operator decide external legal notice alone.
- For a crisis, use the acknowledged safety route immediately rather than waiting for the security-notification process.
Recover
- Rebuild compromised machines from a known-good image.
- Restore only verified configuration, models, and data.
- Re-run identity, data-class routing, cross-user isolation, upload injection, retrieval access, tool denial, egress, rate-limit, logging, and fallback tests.
- Canary the repaired route before normal traffic.
- Record actual RTO, RPO, data affected, and remaining limits.
Learn
- Complete a plain-language review.
- Identify the root cause and missed control.
- Assign an owner and due date.
- Add a regression test where it protects the failed boundary.
- Update notices, training, contracts, and runbooks.
- Preserve or delete evidence under the correct records and legal-hold rules.
This follows the lifecycle in NIST incident-response guidance, April 2025.
On-call model
est — supplied operating assumption: one to one-and-a-half operating staff cannot provide safe continuous human coverage alone.
Use:
- One named platform owner for normal operation, releases, and incident leadership.
- One cross-trained backup for leave, drills, and simultaneous work.
- Central IT/security as the after-hours technical receiver.
- Existing campus safety and crisis services for immediate danger.
- Title IX, Clery, privacy, and legal owners for decisions in their areas.
- A published service window for ordinary faults.
- Temporary pooled coverage around finals if UVU promises faster recovery.
The local operator may be pre-authorized to take reversible containment actions: drain a node, stop a tool, disable an upload, block cloud egress, and invoke tested failover. External notice and destructive recovery remain with the proper university authority.
How students are told
Use a public status page plus in-product and direct notices appropriate to the event.
Tell students:
- What service is affected.
- When it started and whether it continues.
- What information or functions may be involved.
- What UVU has done.
- What the student should do.
- Where to obtain help.
- When the next update will appear.
Do not speculate, identify a victim, disclose crisis details, or promise that no data was affected before the review supports that claim. Make notices accessible and provide language support when needed.
What the plan should change
Ranked from highest priority:
Priority A — Keep minors and real health data closed. Open concurrent-enrollment access only after every phase-two gate passes. Continue synthetic clinical cases until the specific FERPA, HIPAA, hybrid-entity, clinical-host, and business-associate boundaries are approved.
Priority B — Build the crisis handoff before broad student use. Add the separate detector, honest user message, acknowledged human route, record rules, and tests for self-harm, violence, abuse, Title IX, and false alerts.
Priority C — Remove dangerous features from the first pilot. Begin without tools, arbitrary web access, shared document retrieval, or automatic cloud spill. Add each only after its own control and test gate.
Priority D — Replace privacy slogans with a verified data map. Change “content logging off” to an exact statement covering histories, caches, retrieval, telemetry, backups, support access, incident capture, retention, and deletion.
Priority E — Finish the control plane. Add redundant routing, database and queue recovery, identity-outage behavior, network and power failure handling, model mirrors, same-class spares, and tested RTO/RPO.
Priority F — Adopt a model admission record. Pin weights and runtimes, preserve provenance and license terms, scan untrusted artifacts, test before promotion, and retain a known-good rollback copy.
Priority G — Fund real operating ownership. Name the service owner, cross-trained backup, central after-hours receiver, crisis destinations, change approvers, and incident authorities.
Priority H — Make incident response a launch gate. Attach runbooks for cross-user leakage, cloud misrouting, poisoned retrieval, compromised weights, stolen credentials, excess spend, node failure, database exposure, and safety disclosures.
Priority I — Prove the controls. Run the isolation battery, data-class failover, upload and retrieval attacks, node and router loss, backup restoration, finals load, and incident tabletop before student launch.
Source log
Source-log ordinals are reference labels, not measured quantities.
- 1. fact — September 3, 2026: Current program, supplied plan.
- 2. fact — September 2, 2026: Engine assumptions, including concurrent enrollment and spare policy.
- 3. fact — accessed September 3, 2026: Red-team findings.
- 4. fact — accessed September 3, 2026: Architecture contingencies.
- 5. fact — accessed September 3, 2026: Governance findings.
- 6. fact — updated August 2026: U.S. Department of Education dual-enrollment guidance.
- 7. fact — checked September 3, 2026: FERPA school-official guidance.
- 8. fact — Department guidance dated October 22, 2020; law checked September 3, 2026: PPRA guidance.
- 9. fact — current chapter checked September 3, 2026: Utah Student Privacy and Data Protection, Title 53E Chapter 9.
- 10. fact — effective October 14, 2025: Utah higher-education student-data law.
- 11. fact — checked September 3, 2026: FTC COPPA guidance.
- 12. fact — December 2025: UNICEF Guidance on AI and Children.
- 13. fact — December 2019: HHS and Department of Education student-health-record guidance.
- 14. fact — reviewed July 30, 2026: HHS Business Associates.
- 15. fact — checked September 3, 2026: HHS cloud-computing guidance.
- 16. fact — reviewed January 29, 2025: Department of Education Title IX enforcement status.
- 17. fact — policy last formally reviewed October 9, 2025: UVU Policy 162.
- 18. fact — regulation checked September 3, 2026: Clery regulation.
- 19. fact — Utah statute amended 2022: Utah therapist duty-to-warn law.
- 20. fact — effective May 1, 2024: Utah child-abuse reporting law.
- 21. fact — revised 2024: NIMH suicide-help guidance.
- 22. fact — checked September 3, 2026: UVU crisis services.
- 23. fact — checked September 3, 2026: UNCW Sammy.
- 24. fact — July 23, 2025: University of Houston Wayhaven.
- 25. fact — August 10, 2026: NJIT Charlie.
- 26. fact — 2025 risk set, checked September 3, 2026: OWASP GenAI risks.
- 27. fact — March 2025: NIST adversarial machine-learning taxonomy.
- 28. fact — November 27, 2023: Secure AI development guidance.
- 29. fact — April 2025: NIST incident-response guidance.
- 30. fact — May 2010: NIST contingency-planning guidance.
- 31. fact — effective August 10, 2026: UVU Policy 447.
NOT_RUN
- not run — Contact with UVU, any school, clinical host, or vendor. None was attempted.
- not run — Legal review by UVU counsel.
- not run — Review of private contracts, business-associate agreements, placement agreements, or provider terms.
- not run — Inspection of a live router, model server, identity system, database, retrieval index, logs, backups, network, power system, or cloud account.
- not run — Cross-user isolation, prompt-injection, retrieval-poisoning, tool-exfiltration, penetration, load, failover, restore, disaster-recovery, or incident-tabletop testing.
- unknown — Exact Mac count and classes, location, topology, model revisions, licenses, staffing assignments, supported hours, provider capacity, and current backup state.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/40-duty-of-care-resilience.md in the research pack.
Utah law, state policy, licenses, and peers
Appendix T in five lines
- QuestionWhich Utah and federal rules, state programs, model licenses, and peer work change the plan?
- AnswerThe service needs a clear artificial-intelligence label, a firm web-access gate, exact license checks, and proof before state aid counts.
- Deciding numbersApril 26, 2027 web-access deadlinefact36,629 potential campus accountsest$15,000,000 for a shared research data centerfact
- What the plan doesShow that the assistant is not human before and during each chat, meet web-access rules before launch, pin each model’s license, and add the state workforce course to onboarding.
- Still unknownThe university’s share of the $15,000,000 is UNKNOWN; legal opinions, private contract checks, and a live access test are NOT_RUN.
The plan mapped UVU's own policies and procurement path. This appendix adds what arrived from the state and Washington in 2024–2026, whether each model's license fits a public university, and what Utah's peers have actually deployed — with sources. Research cutoff September 3, 2026; planning research, not legal advice. fact dated law, policy, and public pages · est reasoned conclusions · unknown what only a counsel opinion or a private contract can settle.
What this changes in the plan
1. One universal label. Utah's AI disclosure law (effective May 7, 2025) requires disclosure when a person asks in a consumer transaction and before any high-risk interaction in a regulated occupation; whether a free campus assistant is a "consumer transaction" is undecided, so the service shows "UVU AI assistant — not a human" before and throughout every chat (the statute's safe harbor) and keeps counseling, clinical, legal, and financial advice out of the general assistant. Since May 6, 2026 a state-funded university can ask Utah's Office of AI Policy for a joint interpretation, or a 12-month regulatory mitigation agreement — useful for that one question, not a general approval. 2. Accessibility is a launch gate with a date. The federal ADA Title II web rule requires WCAG 2.1 AA; the 2026 interim rule moved UVU's deadline to April 26, 2027. Test the real chat, uploads, streaming answers, and exported documents with assistive technology; generated PDFs must be tagged. 3. Records get precise. A chat tied to a student and kept is a FERPA education record; retained chats and metadata are records under Utah's open-records law (classified, not "confidential"); retention follows content (advising records: five years after separation); legal holds can stop deletion. The promise becomes: "Chat content logging is off by default. UVU retains only the data listed in its records notice, except when a user saves a conversation or a scoped legal or records hold requires preservation." 4. Ride the state's programs. The Utah Board of Higher Education set statewide AI direction (December 2, 2025); its AI Task Force launched May 1, 2026 with Dr. Burns a named member; the free statewide AI Workforce Credential (available since July 1, 2026 to 50,000+ graduates) goes into onboarding instead of a competing certificate; the 2026 budget holds $15 million one-time for a shared AI research data center for Utah's public universities and $3 million ongoing for an AI workforce accelerator — UVU's share is unknown, so neither is counted as funding until an allocation exists. 5. Licenses pinned by exact checkpoint. Qwen3.8-27B (Apache-2.0), gpt-oss-120B (Apache-2.0), Mistral Small 4 (Apache-2.0), GLM-5.3-Flash (MIT), and DeepSeek V4 (MIT) fit; GLM-5.3 and Kimi K3 carry custom terms for counsel; Qwen3.8-Flash-Next is on hold — its Qwen Community License requires a separate license for commercial hosted services and student access may defeat the internal-use exception; "the Qwen3.8 family is Apache-2.0" was false and is corrected. Downloaded published weights sit outside export-control classification unless UVU fine-tunes at scale or gives foreign access. 6. The uniqueness claim narrows. All six Utah peers run cloud vendor tools (Copilot, Gemini, ChatGPT, Claude); UC Irvine, CSU Fullerton, and UC San Diego already run campus-built interfaces, a campus-hosted model, and a hybrid gateway. What no reviewed source documents is UVU's complete three-layer design with a university-owned Apple inference tier. Say that, and publish real adoption measures rather than "first". est
Research date: September 3, 2026, MDT.
This is planning research, not legal advice. Primary law and official government or university sources were used first.
est — 36,629 potential campus accounts. Basis: the supplied data.js adds 30,506 non-concurrent students and 6,123 employees, rounded in the brief to about 37,000. Student employees may create overlap.
1. Utah AI Policy Act and Office of AI Policy
Disclosure duties
Utah’s rules now have two tracks:
| Use | Current duty | Likely UVU effect |
|---|---|---|
| Consumer transaction | fact — effective May 7, 2025. A supplier using generative AI in a consumer transaction must say it is AI, not human, when the person clearly asks. | unknown. No Utah court or published Office interpretation located decides whether UVU’s no-charge campus assistant is part of the tuition-based education transaction. Public status and a zero-dollar user price are not clear exemptions. |
| Regulated occupation | fact — effective May 7, 2025. A person providing services in a Utah-regulated occupation must disclose before a high-risk AI interaction and follow every professional rule. | Applies if AI is used to provide licensed counseling, medical, mental-health, legal, financial, or similar professional services. Keep these services out of the general assistant. |
| Safe harbor | fact — current September 3, 2026. The disclosure enforcement safe harbor applies when the interface clearly identifies itself at the outset and throughout the interaction as generative AI, not human, or an AI assistant. | Put “UVU AI assistant — not a human” above the first prompt and keep a visible label during the chat. |
| Penalties | fact — current September 3, 2026. Up to $2,500 per violation and up to $5,000 per violation of an administrative or court order. | Use one universal disclosure pattern instead of trying to classify each chat in real time. |
“High-risk” includes collecting health, financial, or biometric data and providing personalized financial, legal, medical, or mental-health guidance likely to affect a significant personal decision. Current Utah disclosure chapter
Reasoned conclusion: Treat student-facing use as covered until Utah counsel obtains a contrary interpretation. Provide an obvious path to a person. Do not let the general assistant present itself as a licensed professional.
Regulatory mitigation and joint interpretation
fact — effective May 6, 2026. H.B. 320 expressly includes a state-funded higher-education institution within the Office of AI Policy program. UVU may apply for:
- A joint interpretation agreement explaining how a Utah law or rule applies to the planned service.
- A regulatory mitigation agreement temporarily adjusting an identified Utah rule through cure periods, reduced civil fines, safeguards, disclosures, or reporting.
An agreement can limit users, geography, scope, and duration. It can require safeguards, disclosures, reports, and audits. It is not state approval, and laws not expressly addressed remain in force.
fact — H.B. 320: The first term may not exceed 12 months. An extension request is due at least 30 days before the term ends, and the Office may grant up to 2 extensions. Current Office statute
Recommendation: If UVU wants certainty about “consumer transaction,” seek a joint interpretation. Use mitigation only after naming a specific Utah rule that blocks a bounded pilot. The state program cannot waive FERPA, the ADA, or other federal law.
Other 2026 change
fact — effective January 1, 2027. Utah’s Digital Content Provenance Standards Act covers a provider that produces a publicly accessible generative system with more than 1,000,000 monthly users or visitors. It requires provenance information for generated or substantially altered image, audio, and video content when technically feasible. H.B. 276
Reasoned conclusion: UVU’s planned 36,629-account EST service is below that provider threshold. The rule could still govern a large underlying model provider. Recheck if UVU later makes a public multimodal service available beyond campus.
2. ADA Title II web accessibility rule
Coverage and deadline
fact — DOJ rule published April 24, 2024. Public universities are covered, including services supplied through contractors and licensed platforms. The technical standard is WCAG 2.1 Level A and Level AA.
fact — 2026 correction. DOJ’s interim final rule, effective April 20, 2026, moved the large-entity deadline from April 24, 2026 to April 26, 2027. It moved the small-entity and special-district deadline to April 26, 2028. DOJ current fact sheet
fact — current DOJ guidance. A state university uses its state’s population, not its enrollment or city population. UVU therefore uses the large-entity deadline: April 26, 2027. DOJ first-steps guide
The extension does not suspend UVU’s existing duties to provide effective communication, reasonable changes, and equal access.
Chat interface
The production service should provide:
- Full keyboard operation and a logical focus order.
- Visible focus and clear labels, instructions, and error messages.
- Screen-reader announcements for streaming answers, upload progress, and status changes.
- Contrast, zoom, and reflow support.
- Accessible sign-in, model selection, conversation controls, and time-limit controls.
- Alternatives for visual, audio, and video information.
- A real assistive-technology test using complete conversations. An automated page scan is not enough.
Uploaded documents
The upload button, instructions, accepted-file information, progress, errors, preview, extracted text, and generated answer are part of UVU’s service.
A document independently uploaded by a user may sometimes fit the unaffiliated third-party-content exception. That exception does not cover UVU’s upload interface or the summaries, conversions, and answers produced by UVU or its contractor.
PDF outputs
A newly generated PDF is not preexisting content. After the deadline, public, course, and general-service outputs normally must meet WCAG 2.1 AA.
Use accessible HTML as the primary result. If the user requests PDF, generate a tagged file with:
- Headings and correct reading order.
- Document title and language.
- Alt text.
- Accessible tables.
- Selectable text and sufficient contrast.
A password-protected, individualized document may fit a narrow exception, but UVU must still provide the information promptly in an accessible form. A separate accessible version is not a routine escape hatch.
Section 508 overlap
fact — current September 3, 2026. Section 508 directly governs federal agencies’ information and communications technology. UVU does not become directly subject to Section 508 merely because it is public or receives federal funds. U.S. Access Board
Section 508 may enter through a federal contract or when UVU supplies technology to a federal agency. Section 504 is the broader federal-funding rule that applies to public colleges receiving federal assistance. A vendor VPAT or accessibility report is useful evidence but does not prove the complete UVU service is accessible.
3. Records
FERPA
A chat becomes a FERPA education record when it is directly related to a student and maintained by UVU or a party acting for UVU. Department of Education FERPA materials
Authenticated chats about advising, assignments, grades, progress, accommodations, discipline, financial aid, or student support are likely education records if retained.
A general anonymous chat is not automatically an education record. The answer depends on its content, whether it can be linked to a student, and whether it is maintained. Removing a name alone may not de-identify a record when the remaining facts identify the student.
Sending education-record content to a cloud model can be a FERPA disclosure even if the provider says it does not retain prompts. A provider relying on the school-official exception must perform an institutional function, remain under UVU’s direct control for record use and maintenance, observe redisclosure limits, and meet UVU’s published criteria. Department of Education school-official guidance
Local inference reduces third-party disclosure risk. It does not remove FERPA from saved history, exports, administrator access, backups, or support records.
Utah GRAMA
fact — current September 3, 2026. GRAMA covers state-funded higher-education institutions. Its record definition includes reproducible electronic data prepared, owned, received, or retained by the governmental entity. Utah GRAMA
Retained prompts, answers, account associations, saved chats, moderation events, exports, and usable logs can be GRAMA records. A vendor’s storage location does not necessarily remove them if UVU owns or controls them.
A “record” is not automatically a “public record.” It can be private, controlled, protected, or otherwise exempt. FERPA education records remain governed by FERPA. Other chats may be public unless their content supports a classification or exemption. UVU must classify, segregate, and redact them rather than promise that all chats are confidential.
Retention
unknown: No approved Utah or UVU schedule specifically titled “AI chat logs” was verified.
Retention follows the content and business purpose, not the file format. Possible current schedules include:
| Possible function | Published schedule | Retention |
|---|---|---|
| Educational advising | GRS-2042 | fact — 5 years after separation, then destroy. |
| Student discipline | GRS-1504 | fact — until the issue is resolved, then destroy. |
| Transitory correspondence | GRS-1759 | fact — until resolution, then destroy. |
| Transitory tracking, including website visitor information | GRS-1720 | fact — 1 year after final action, then destroy. |
| Routine state administrative correspondence | GRS-48 | fact — 7 years, then destroy. |
| Program and policy development | GRS-1717 | fact — retain 3 years after final action, then transfer to the archives permanently. |
These are possible mappings, not one automatic schedule for all AI data. Before launch, UVU’s records officer should map chat content, saved history, feedback, safety events, access logs, model-routing logs, exports, caches, and backups to approved series or obtain a new schedule. Utah Archives retention guidance
E-discovery
Utah civil discovery can reach electronically stored information within UVU’s possession, custody, or control. That can include retained chat content, metadata, answers, moderation records, and provider-held data when the contract gives UVU control.
Routine deletion may follow an approved schedule. Once UVU reasonably expects litigation or receives a preservation duty, it must be able to stop deletion for relevant records. Utah courts may act when a party destroys or fails to preserve electronic evidence in violation of that duty. Utah Rule of Civil Procedure 37
Meaning of “no content logging”
| Area | Required meaning |
|---|---|
| FERPA | Prompts and answers are not persisted after the session unless the user saves them or an authorized record workflow requires them. This reduces maintained education records but does not authorize cloud disclosure. |
| GRAMA | Content never retained generally cannot be reproduced as a content record. Retained metadata remains a record. |
| Retention | Saved chats, caches, backups, error traces, safety samples, exports, and provider copies must appear in the records map. |
| E-discovery | Ordinary minimization reduces existing evidence. A legal hold must preserve relevant existing and future data once a duty arises. |
Replace “we never log anything” with:
Chat content logging is off by default. UVU retains only the data listed in its records notice, except when a user saves a conversation or a scoped legal or records hold requires preservation.
4. State higher-education policy
USHE and the Utah Board of Higher Education
fact — December 2, 2025. The Utah Board set statewide direction calling for human-centered AI literacy and responsible AI use in teaching, research, student services, and operations. It is strategic direction, not a detailed technical standard. USHE announcement
fact — May 1, 2026. USHE launched its statewide AI Task Force. UVU Chief AI Officer Barclay Burns is a named member. USHE task-force announcement
Its first project is an AI Workforce Credential:
- fact — USHE dated May 1, 2026: More than 50,000 graduates from the classes of 2025, 2026, and 2027 are eligible.
- fact: The credential became available at no cost on July 1, 2026.
- fact: It uses a statewide online environment and involves USHE institutions, Talent Ready Utah, employers, and the governor’s Pro-Human AI group.
- unknown: Public sources do not yet state its completion rules, platform ownership, or exact place in UVU degree programs.
UVU’s onboarding should carry this credential instead of creating a competing general AI-literacy certificate.
UETN and UEN
fact — checked September 3, 2026. UETN supplies statewide education networking, bulk purchasing, and technology training. Its public material does not establish a UETN-operated higher-education generative-AI platform. That capacity is unknown. UETN
UEN offers free self-paced courses including “AI and Student Learning” and “AI and the Writing Process,” plus a public AI toolkit. Its published programming is aimed mainly at school educators, although parts can support teacher preparation and faculty development. UEN courses
Use UETN as a network, purchasing, and training partner. Do not count it as model-hosting capacity without written evidence.
2025 legislative session
- fact: H.B. 168 would have created an AI-in-education task force covering privacy, security, literacy, integrity, and equity. It appropriated $0 and ended “House filed,” so it did not pass. H.B. 168 record
- fact: H.B. 265, Higher Education Strategic Reinvestment, was signed. Its implementation shifted institutional funds toward high-demand programs including applied AI and computer science. It was not a dedicated campus-chat appropriation. USHE reinvestment record
- fact/UNKNOWN: A higher-education subcommittee recommended $2,000,000 one-time for UVU’s Applied AI Institute. A final matching award was not verified in this run, so UVU receipt is unknown. 2025 subcommittee recommendation
2026 legislative session
- fact — budget dated March 3, 2026: $15,000,000 one-time for an Artificial Intelligence Public-Private Partnership Ecosystem, including a shared AI research data center for Utah public universities and researchers, with spending across 3 years.
- fact — same budget: $3,000,000 ongoing for an Energy, Artificial Intelligence, and Deep Tech Workforce Accelerator.
- unknown: The budget does not name a UVU allocation, access date, service level, or cost offset. Official 2026 budget index
No enacted standalone 2026 mandate requiring a USHE campus chatbot was located.
Governor initiatives
- fact — March 10, 2025: UVU is already named in Utah’s NVIDIA education agreement. It offers teaching kits, workshops, accelerated-computing resources, an instructor-ambassador route, internships, and apprenticeships. Governor’s Office announcement
- fact — February 25, 2026: The Pro-Human AI Task Force covers academic research, education, public protection, government, industry, and workforce development.
- unknown: No open university application process was published. UVU’s practical path is through its USHE task-force seat and existing NVIDIA participation. Pro-Human AI Task Force
The shared research data center and statewide credential are the most concrete programs to pursue. Neither should be shown as UVU funding until a written allocation or agreement exists.
5. Model-license fitness
est — about 37,000 users. This is below all user-count branding thresholds found. Revenue, business-purpose, and third-party-use clauses remain separate issues.
“FIT” below covers the downloadable-weight license only. It does not approve model quality, privacy, security, accessibility, or procurement.
| Model | Commercial use and thresholds | Attribution, indemnity, and other terms | Verdict |
|---|---|---|---|
| GLM-5.3 | Custom license permits hosting, modification, derivatives, distribution, sublicensing, and sale. fact — 2026 license: A Model-as-a-Service business with affiliate revenue above $10 billion during any consecutive 12 months must pass Z.AI security review before commercial use. No user threshold. | Keep copyright and license text with copies or substantial portions. No express patent grant, warranty, or indemnity. | CONDITIONAL FIT. The scale is not a trigger. Counsel should decide whether the campus service is a “business” or commercial use. |
| GLM-5.3-Flash | MIT; commercial use, modification, hosting, sublicensing, and sale allowed. No user or revenue threshold. | Keep the copyright and permission notice with copies or substantial portions. No express patent grant, warranty, or indemnity. | FIT. Cleaner default than full GLM-5.3. |
| Kimi K3 | Custom license permits commercial use and derivatives. fact — 2026 license: Separate agreement above $20 million affiliate revenue during any consecutive 12 months for a Model-as-a-Service business. A commercial service above 100 million monthly users or $20 million monthly revenue must show “Kimi K3” prominently. | Keep the license notice. No express patent grant, warranty, or indemnity. Whether campus students count as third parties under the internal-use exception is unknown. | CONDITIONAL FIT. The user threshold is not reached, but the low business-revenue trigger and internal-use wording need counsel review. |
| Qwen3.8-27B | Apache-2.0; commercial deployment, modification, and redistribution allowed. No user or revenue threshold. | On redistribution: include the license, mark changes, retain notices, and pass through any NOTICE file. Express patent grant; no warranty or indemnity. | FIT. This is the checkpoint named in data.js. |
| Qwen3.8-2.4T-A95B | Custom Qwen3.8-Max license, not Apache-2.0. fact — 2026 license: Branding applies above 100 million monthly users or $20 million monthly revenue. A Model-as-a-Service or AI Work Assistant business above $50 million affiliate revenue during any consecutive 12 months needs a separate license. | Keep notice and obey lawful-use and third-party-rights terms. No express patent grant, warranty, or indemnity. | CONDITIONAL FIT. Do not inherit the 27B model’s Apache label. |
| Qwen3.8-Flash-Next | Qwen Community License 1.0, not Apache-2.0. Any commercial Model-as-a-Service or AI Work Assistant business needs a separate license; there is no revenue floor. fact: The 100 million-user or $20 million-monthly-revenue branding threshold also applies. | Custom notice and use restrictions. Student access may defeat the internal-use exception; that is unknown. | HOLD until counsel confirms noncommercial treatment or UVU obtains a separate license. |
| gpt-oss-120B | Apache-2.0; commercial use, hosting, modification, and redistribution allowed. No user or revenue threshold. | Apache notice, change-marking, patent, warranty, and liability terms apply. OpenAI’s separate policy requires lawful use. No vendor indemnity. | FIT, with a UVU acceptable-use layer and preserved notices. |
| Mistral Small 4 | Apache-2.0; commercial and noncommercial use allowed. No user or revenue threshold. | Apache redistribution and patent terms apply. The model card also bars infringement or misuse of third-party rights. No warranty or indemnity. | FIT. |
| DeepSeek V4 Pro-0813 and Flash-0731 | MIT; commercial use, hosting, modification, sublicensing, and redistribution allowed. No user or revenue threshold. | Keep the MIT notice. No express patent grant, model-specific indemnity, or warranty in the repository license. | FIT on license terms. Security and procurement review remain separate. |
The plan’s “Qwen3.8 family — Apache-2.0” statement is false. Only the reviewed Qwen3.8-27B checkpoint uses Apache-2.0.
EAR note for downloaded weights
fact — current EAR reviewed through September 1, 2026. ECCN 4E091 Note 1 excludes AI parameters that have been “published” under EAR §734.7.
Reasoned conclusion: The publicly downloadable official weights appear to fit that published-parameter exclusion as downloaded. No repository supplied a vendor ECCN or CCATS, so a checkpoint-specific classification remains unknown.
fact — rule dated January 15, 2025. A substantially trained derivative can leave the exclusion after additional training above the greater of 2.5 × 10²⁵ operations or 25% of the training operations described in ECCN 4E091 Note 2. Ordinary inference does not cross that training rule.
The license is not an export authorization. Before foreign redistribution, foreign administrator access, or a major fine-tune, UVU should screen destination, end user, end use, sanctions, model-weight rules, encryption, and attached hardware. Do not label a checkpoint EAR99 without supported classification. Current BIS EAR Part 734
6. Peers
“Scale” means the verified eligible group or usage measure. It is not an assumed headcount.
Utah peers
| Institution | What is deployed | Local or cloud | Scale and public cost | Date |
|---|---|---|---|---|
| Utah State University | Protected Copilot, Zoom AI Companion, Box AI, Gemini, NotebookLM, and department-paid options for Microsoft 365 Copilot, Claude, and ChatGPT Business. | Cloud; Microsoft, Zoom, Box, Google, Anthropic, OpenAI. | Base tools for current students, faculty, and staff — fact; active use unknown. Base Copilot has no added user charge. Published add-ons: $30/user/month Microsoft 365 Copilot, $20 or $50/user/month Claude, and $20/user/month ChatGPT Business — fact, checked September 3, 2026. Institution total unknown. | Live by September 3, 2026; original launch unknown. USU tools |
| BYU | Public/basic Copilot access; a course-level teaching bot built around a textbook and syllabus. No verified campus-managed general assistant. | Cloud; Copilot vendor Microsoft. Course-bot vendor unknown. | Students and faculty for Copilot — fact. The course bot supports about 2,500 students per year — FACT, 2025 annual report. Costs and active campus use unknown. | Copilot evidence October 17, 2024; course bot began in early 2023. BYU annual report |
| Weber State | Authenticated Gemini, NotebookLM, Copilot, and other approved vendor tools. | Cloud; Google, Microsoft, and other vendors. | Eligible students, faculty, and staff — fact; usage unknown. Published Gemini premium price $36/user/month — FACT, checked September 3, 2026; institution total unknown. | Gemini and NotebookLM rollout May 5, 2025 — fact. Weber AI services |
| Southern Utah University | Thor, an AI texting chatbot for main-campus undergraduates; SUU-provided Gemini appears in current course requirements. | Cloud; Gemini by Google. Thor vendor unknown. | Thor available to main-campus undergraduates — fact; campus-wide Gemini entitlement and usage unknown. Cost unknown. | Thor and Fall 2026 course evidence current September 3, 2026; original launch unknown. Thor |
| Salt Lake Community College | Base Copilot for faculty and staff; public free Copilot for students; training-linked premium licensing. | Cloud; Microsoft. | Faculty and staff base access — fact; usage unknown. Base user price $0 — FACT. Premium license $209/user/year — FACT, current page; institution total unknown. | Article created July 26, 2025; modified May 18, 2026. SLCC Copilot |
| Utah Tech University | University-agreement Copilot Chat and Gemini; departments may purchase ChatGPT Business. | Cloud; Microsoft, Google, OpenAI. | Faculty, staff, and eligible students for Copilot — fact; usage unknown. Copilot has no added user cost. ChatGPT Business guidance lists $25/user/month annually or $30/user/month monthly — FACT. Institution total unknown. | September 23, 2025 guidance. Utah Tech guidance |
University of Utah context: fact — November 20, 2025. It launched campus ChatGPT Edu and added Gemini and NotebookLM on May 12, 2026. Students, faculty, and staff may request access. Usage and institution cost are unknown here. ChatGPT Edu announcement
Similar-size public universities outside Utah
| Institution | Verified deployment | Architecture and vendor | Scale and cost | Date |
|---|---|---|---|---|
| University of California, Irvine | ZotGPT Chat, Gateway API, ClassChat, and no-code Creator. | Campus-built interface using Microsoft Azure AI, Amazon AWS, and open-web software. | Enrollment 36,621 — FACT, 2024–25. Available to all students, faculty, and staff. More than 1,000 custom bots and more than 20 public department bots — FACT, November 18, 2025. User charge $0 for Chat, Copilot Chat, and Gemini Chat; institution total unknown. | Faculty/staff launch January 10, 2024; student access April 25, 2024. ZotGPT |
| California State University, Fullerton | TitanGPT plus opt-in ChatGPT Edu. | TitanGPT is campus-hosted; ChatGPT Edu is cloud OpenAI. | Enrollment 45,863 — FACT, Fall 2025. Both services support students, faculty, and staff. TitanGPT showed 9,603 authenticated users — FACT, October 2025 board material. Institution cost unknown. | ChatGPT Edu live by April 2025; TitanGPT verified Fall 2025. CSUF technology guide |
| University of California, San Diego | TritonGPT with chat, documents, campus assistants, course tutors, model choice, and a shared model gateway. | Started on local San Diego Supercomputer Center infrastructure; now combines approved enterprise cloud models and open models hosted on UC-controlled infrastructure. | Enrollment 45,087 — FACT, Fall 2025. Available to faculty, students, and staff. Institution cost unknown. | All campus and Health Sciences employees by Spring 2024; student access June 2025 — fact. TritonGPT overview |
What UVU’s plan does that none of these peers documents
est — evidence-set conclusion: No reviewed public source documents UVU’s complete three-layer design:
- 1. University-owned Apple Silicon running open models locally.
- 2. Protected free cloud tools for ordinary work.
- 3. A separate, centrally metered frontier-model pool.
This is not proof that no university has an unpublished version.
UVU should not claim that campus interfaces, multi-model gateways, local hosting, custom agents, document chat, or metering are new. UC Irvine, CSU Fullerton, and UC San Diego already demonstrate several of those elements. The defensible distinction is the Apple-owned inference tier and its place in the full routing and cost design.
What the plan should change
- 1. Make accessibility a launch gate. Test the real chat, uploads, streaming answers, authentication, and exported documents against WCAG 2.1 AA. Record the fact deadline: April 26, 2027.
- 2. Replace the broad “no logging” promise. Publish a complete data map, default deletion behavior, approved retention-series mapping, cloud-provider handling, FERPA basis, GRAMA classes, access roles, and legal-hold override.
- 3. Add universal AI identification. Show “UVU AI assistant — not a human” before and throughout every chat. Route counseling, clinical, legal, financial, and other regulated work to people unless a separately approved service exists.
- 4. Pin licenses by exact checkpoint. Replace “Qwen3.8 family — Apache-2.0” with “Qwen3.8-27B — Apache-2.0.” Archive each model’s repository commit, license, model card, and notice bundle.
- 5. Hold custom-license risks. Prefer GLM-5.3-Flash, Qwen3.8-27B, gpt-oss-120B, Mistral Small 4, or DeepSeek V4 on license simplicity. Keep GLM-5.3 and Kimi K3 behind counsel review. Hold Qwen3.8-Flash-Next campus-wide.
- 6. Tie the program to USHE. Put the statewide AI Workforce Credential into onboarding, use UVU’s task-force seat, and check future USHE guidance before production launch.
- 7. Seek state compute terms before buying research-scale capacity. The $15,000,000 FACT shared-data-center budget may help, but UVU’s share, timing, and access remain unknown.
- 8. Narrow the uniqueness claim and publish real adoption measures. Report eligible people, activated accounts, monthly active users, repeat users, local-versus-cloud routing, and frontier spend separately.
- 9. Use an Office of AI Policy agreement only for a named issue. A joint interpretation is reasonable for the consumer-transaction question. Do not present the sandbox as a general compliance approval.
- 10. Add an export and model-origin review gate. Recheck foreign access, sanctions, large fine-tunes, procurement rules, security, and exact model classification before weights leave UVU-controlled systems.
Source log
- 1. 2026-09-03: Current adoption program, engine assumptions, and frontier-pool findings.
- 2. March 13, 2024: Utah S.B. 149, enrolled.
- 3. Effective May 7, 2025: Utah S.B. 226, enrolled and current Chapter 77.
- 4. Current September 3, 2026: Utah consumer-transaction, person, and supplier definitions.
- 5. Effective May 6, 2026: Utah H.B. 320, enrolled and current Office statute.
- 6. Effective January 1, 2027: Utah H.B. 276, Digital Content Provenance Standards Act.
- 7. Published April 24, 2024: DOJ Title II web and mobile final rule.
- 8. Effective April 20, 2026: DOJ interim final rule, current fact sheet, and first-steps guide.
- 9. Current September 3, 2026: U.S. Access Board ICT standards, Section508.gov scope, and Department of Education disability guidance.
- 10. Current September 3, 2026: FERPA regulations and school-official requirements.
- 11. Current September 3, 2026: Utah GRAMA.
- 12. Current September 3, 2026: Utah Archives retention schedules and record-series guidance dated August 6, 2026.
- 13. Current September 3, 2026: Utah Rule of Civil Procedure 34 and Rule 37.
- 14. December 2, 2025: USHE statewide AI direction.
- 15. May 1, 2026: USHE AI Task Force and AI Workforce Credential.
- 16. Checked September 3, 2026: UETN, UEN courses, and UEN AI resources.
- 17. 2025 session: H.B. 168 record, H.B. 265 implementation, and UVU Applied AI Institute recommendation.
- 18. March 3, 2026: Utah index of budgeted funding items.
- 19. March 10, 2025: Utah-NVIDIA education agreement.
- 20. February 25, 2026: Utah Pro-Human AI Task Force.
- 21. Checked September 3, 2026: Utah State University AI tools.
- 22. 2024–2025 evidence: BYU Copilot article and BYU Marriott annual report.
- 23. Checked September 3, 2026: Weber State AI services.
- 24. Checked September 3, 2026: SUU Thor chatbot and Fall 2026 Gemini course evidence.
- 25. Created July 26, 2025; modified May 18, 2026: SLCC Copilot.
- 26. September 23, 2025: Utah Tech premium AI licensing.
- 27. November 20, 2025 and May 12, 2026: University of Utah ChatGPT Edu and Google tools.
- 28. 2024–2026: UC Irvine ZotGPT, launch record, and UCI facts.
- 29. 2025–2026: CSU Fullerton technology guide, ChatGPT Edu, and enrollment facts.
- 30. 2024–2026: UC San Diego TritonGPT, hosting and access FAQ, and 2025–26 enrollment.
- 31. 2026; checked September 3, 2026: GLM-5.3 license and GLM-5.3-Flash MIT license.
- 32. 2026; checked September 3, 2026: Kimi K3 license.
- 33. 2026; checked September 3, 2026: Qwen3.8-27B Apache license, Qwen3.8-Max license, and Qwen Community License.
- 34. Released August 5, 2025: gpt-oss license and usage policy.
- 35. Released March 16, 2026: Mistral Small 4 announcement and model card.
- 36. Checked September 3, 2026: DeepSeek V4 Pro-0813 license and Flash-0731 license.
- 37. Rule dated January 15, 2025; current EAR checked through September 1, 2026: Federal Register rule, EAR Part 734, and EAR Part 742.
NOT_RUN
- findings-05-governance.md: not run. The named file was not present beside the brief.
- Vendor, university, Utah agency, and UVU contact: not run.
- Login-only verification of model menus, entitlements, or usage: not run.
- Contract, invoice, procurement-record, or institution-total-cost review: not run.
- Legal opinion on “consumer transaction,” “commercial purpose,” “business,” or model-license “third party”: not run.
- Vendor ECCN or BIS classification request: not run.
- Accessibility audit of a working UVU interface or exported PDF: not run.
- Review of final cloud-provider contracts, retention settings, abuse logs, backups, or legal-control clauses: not run.
- Independent verification of unpublished peer deployments: not run.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/41-utah-law-policy-peers.md in the research pack.
Money angles
Appendix U in five lines
- QuestionWhat will five years cost, when should machines be replaced, who should pay, and can paid students help run them?
- AnswerCompare all five years, buy first unless a real lease quote wins, review in year three, and use central funding plus paid students.
- Deciding numbers$90,341 with no refreshest$122,438 with a year-three refreshest$128,233 with a year-four refreshest
- What the plan doesAdd the five-year cases, make year four the default replacement point, set $0 from general student fees, and put paid students under a staff lead.
- Still unknownA university lease quote, power bill, wage table, staffing test, and grant review are NOT_RUN; true rates and costs stay UNKNOWN.
The plan priced three years, assumed a purchase, funded centrally, and staffed with professionals. A chief financial officer will ask about years four and five, leasing, what the machines are worth when replaced, who pays, whether students can run it, and the real electricity rate. This appendix answers each with public numbers. Research cutoff September 3, 2026. fact dated source or direct arithmetic · est planning estimate with basis · unknown not public or not verified.
What this changes in the plan
1. Five-year cases beside the three-year one. For the 26-mini fleet: keep the first fleet five years, $90,341; refresh after year three, $122,438; refresh after year four, $128,233 — net of modeled resale, excluding spares and the value of the fleet still owned at the end. Three-year resale values from completed sales of the previous Apple generations: mini 50%, Studio 50%, Ultra 55%, before fees. Recommendation: a measured review gate at year three, refresh by default in year four; the machines are containers for better free models, not disposable. 2. Buy the first fleet unless a written lease quote beats it. Apple's education financing advertises rates as low as 0% for up to four years and a use-only structure; at 0% the 26-mini package ($60,424 with care) is $1,678 a month to own or $874 a month to use and return — but the 0% floor is an advertisement, not a rate UVU can budget, and Utah policy reports equipment payment arrangements over 12 months. 3. Who pays: $0 from student fees. Utah's fee rule (USHE R516) bars general fees for instruction, academic support, and administration — which is what this service is — and UVU's $334.08 semester fee has no technology line. Central IT funds the shared base; colleges and grants buy marginal nodes; showback first, chargeback only for reserved or unusually heavy use; add a college only after a signed three-year commitment covers its hardware and half its marginal staff time. 4. A student operations team under a professional owner. At $18 an hour (about $43,000 per student full-time equivalent with turnover reserve), the three staffing bands become 0.35 professional + 0.20 student ($59,052, +17% — more coverage, not less cost), 0.75 + 0.50 ($129,616, −28%), and 2.5 + 1.5 ($424,879, −26%); UVU already has credit-bearing routes (INFO 4810R/4890R, IT 4810R, TECH 2810R). The simulator and configurator now carry this as a switch. 5. Grants beyond the ones already in the plan: the NSF State and Regional AI Infrastructure Hubs solicitation (deadline November 4, 2026; $4–12 million over five years; one award per state — a Utah consortium, not a solo request); NSF IUSE for a teaching study (January 20 and July 21, 2027; up to $2 million); the Utah AI Workforce Accelerator ($3 million ongoing; request for proposals expected but not yet public); Apple's Community Education Initiative as a partnership target; and a philanthropy package shaped like the ones that worked elsewhere — named student-operations fellowships, named leadership, vendor equipment, and a written commitment to recurring operations. 6. Electricity checks out. UVU's own rate is not public; at the Utah commercial average ($0.1099 per kWh) and a facility overhead of 20%, the plan's power lines ($485 per mini, $503 per Studio, $936 per Ultra over three years) reconcile to the cent; the industry-average overhead of 52% would add about $300 per sustained kilowatt-year. est
Five-year cost
Three-year residual values
Observed September 3, 2026 fact. These are gross marketplace prices, not guaranteed university proceeds.
| Class | Public market evidence | Planning residual |
|---|---|---|
| Mini | An M2 Pro mini with 32 GB and 1 TB recorded a completed price of $924.95 fact against a $1,899 fact original configuration price: 48.7% fact. eBay completed listing, price history. Apple currently asks $1,989 fact, or 90.5% fact of its reference price, for another M2 Pro configuration; that is a retail asking price, not resale proceeds. Apple Refurbished | 50% est |
| Studio | Completed M2 Max base-model prices ranged from $979 to $1,249 fact against $1,999 fact at launch: 49.0%–62.5% fact. The highest-volume observed listing was close to 50% fact. Apple launch, eBay $979, eBay $999, eBay $1,120, eBay $1,249 | 50% est |
| Ultra | Completed M2 Ultra base-model prices ranged from $2,049 to $2,369 fact against $3,999 fact at launch: 51.2%–59.2% fact. Older M1 Ultra records ranged from 40.0% to 55.0% fact. eBay $2,049, eBay $2,308, eBay $2,369 | 55% est |
Back Market’s exact M2 Pro and M2 Ultra comparison pages had no available price unknown. Apple Refurbished prices are useful ceiling checks, but include Apple’s warranty, preparation, and retail margin.
The recommended residuals are before seller fees, shipping, damage, missing accessories, or a bulk-sale discount. Net proceeds are therefore lower by an unknown amount.
Five-year mini-fleet model
Basis:
- 26 active minis fact scenario.
- Education hardware price: $2,139 each fact, plan input, plus $90 each fact for the production-plan network option: $57,954 total fact arithmetic.
- Plan AppleCare allowance: $120 per device for three years est, or $3,120 per fleet est arithmetic.
- Shared equipment: five kits est × $1,650 est = $8,250 est.
- Plan power: $485 per device per three years est, extended to $21,017 over five years est arithmetic for the fleet.
- Year-four residual: 40% est, applying a 20% est haircut to the three-year mini residual.
- Future replacement prices are held flat est.
- Taxes, sale costs, support after AppleCare, and terminal value of the final fleet are excluded unknown.
| Five-year path | Net hardware cash | Care, shared kit and power | Five-year cash cost |
|---|---|---|---|
| Keep the first fleet for five years | $57,954 est | $32,387 est | $90,341 est |
| Refresh after year three | Two purchases less $28,977 est resale = $86,931 est | $35,507 est | $122,438 est |
| Refresh after year four | Two purchases less $23,182 est resale = $92,726 est | $35,507 est | $128,233 est |
These are cash costs, net of modeled resale where a refresh occurs. They do not credit the value of the fleet still owned at the end of year five because that fleet is a different age under each path. Its terminal value is unknown.
The table also excludes the plan’s 15% spare policy est. If “26” means active machines rather than total machines, four spares are required est arithmetic. Those add $8,916 est of hardware at each purchase cycle, plus care and power.
Refresh recommendation
Make year three a formal review gate and year four the default refresh est recommendation.
The model-progress review projects that stronger open models will keep appearing for the same memory footprint, with a central open-versus-closed gap of roughly three to six months through 2028 est. That makes the machines reusable containers for better models; it does not make every machine obsolete after three years. Large frontier models will still exceed local memory and belong in the metered cloud pool. See findings-08-model-progress-regression.md and findings-36-mac-generations-side-by-side.md.
Refresh in year three only if measured workloads fail the agreed capability, wait-time, security, or support gate. A year-four refresh has the main risk that the original three-year care period leaves a support gap; extension price and availability are unknown.
Lease vs buy
Public terms
Apple Financial Services currently advertises education financing, level payments, refresh planning, Easy Return, and financing of hardware, software, accessories, and some third-party items. Its education flyer says rates may be as low as 0% fact offer floor, payments may follow a semester or school-year schedule for up to four years fact, payment deferral may be offered, and Guaranteed Buyback may apply. Every offer remains subject to credit approval and signed documents. Apple Financial Services, education flyer
Apple’s published fair-market-value structure lets the institution return the equipment or buy it for then-current market value. The institution does not automatically own it during the use-only term. Apple FMV structure
UVU can purchase through Apple’s Utah NASPO contract 23003 fact identifier, Utah addendum PA4282 fact identifier, or PEPPM contract 576202 fact identifier. NASPO and PEPPM allow financing where the participating contract permits it, but no public UVU rate, guaranteed residual, fee schedule, or detailed lease option was found unknown. Apple Utah contracts, NASPO Apple record, PEPPM terms
Utah policy calls for reporting periodic-payment equipment arrangements longer than 12 months fact. Its exact application to UVU’s chosen structure needs procurement review unknown. Utah finance policy
Side-by-side: mini fleet
This financing comparison uses the current public package of 26 devices fact scenario, $57,954 hardware fact arithmetic, and $2,470 current public AppleCare pricing fact arithmetic: $60,424 financed fact arithmetic.
The 6% rate est sensitivity is not an Apple quote.
| Path | Cash and total cost | Ownership | Flexibility |
|---|---|---|---|
| Buy | $60,424 now fact scenario. Modeled hardware resale after three years: $28,977 est. Net capital and care before selling costs: $31,447 est. | UVU owns immediately. | Best reuse and resale freedom; UVU carries obsolescence and disposal work. |
| Plan-to-own | At 0% est illustration: $1,678.44/month est, $60,424 total est. At 6% est sensitivity: $1,838.22/month est, $66,175.75 total est. | Ownership timing depends on the signed schedule unknown. | Smooth cash flow, but costs more if the rate is above zero. |
| FMV/use-only | With a 50% residual est, at 0% est illustration: $873.53/month est, $31,447 total est, then return. At 6% est sensitivity: $1,101.56/month est, $39,656.29 total est, then return. Buying at modeled term-end value would raise the 6% case est to $68,633.29 est. | Apple or the financier owns during the term. | Cleanest scheduled refresh and lowest modeled payments; UVU gives up resale upside and must meet return conditions. |
External racks, switches, UPS units, storage, cables, spares, deployment, and staff are excluded unknown.
Side-by-side: $120,000 mixed fleet
The mix of minis, Studios, and Ultras is unknown, so the model uses the recommended 50%–55% residual range est.
| Path | Three-year cash and total cost | Ownership and flexibility |
|---|---|---|
| Buy | $120,000 now fact scenario; $60,000–$66,000 residual est; $54,000–$60,000 net hardware consumption est, before disposal costs. | Own immediately; greatest reuse and sale freedom. |
| Plan-to-own | At 0% est illustration: $3,333.33/month est, $120,000 total est. At 6% est sensitivity: $3,650.63/month est, $131,422.77 total est. | Smooth payments; eventual ownership depends on the signed form unknown. |
| FMV/use-only | At 0% est illustration: $1,500–$1,666.67/month est, $54,000–$60,000 total est, then return. At 6% est sensitivity: $1,972.78–$2,125.32/month est, $71,020.25–$76,511.38 total est, then return. | Best scheduled refresh; no resale upside and return-condition risk. |
Recommendation: buy the first fleet unless Apple supplies a written use-only offer with a favorable guaranteed return value and UVU values budget smoothing more than ownership. The advertised 0% floor fact is not a rate UVU can budget until it has a quote.
Who pays
Student-fee boundary
USHE policy R516 allows general fees for approved activities, programs, services, and non-instructional facilities that broadly benefit students. It bars general-fee funding for instruction, academic support, general administration, and expenses reasonably covered by tuition or state appropriations. It also requires separate accounting, annual review, a student-majority committee, a public hearing, trustee action, and Board approval. USHE R516
UVU Policy 511 follows this process. UVU’s public page says student fees may support technology, but not academic-program development, one-time funding needs, replacement of budget cuts, or replacement of grants and donations. UVU Policy 511, UVU student fees
UVU’s published semester schedule for 12 or more credits fact totals $334.08 fact:
- Building bonds: $87.00 fact
- Athletics: $84.26 fact
- Student programs: $56.84 fact
- Student center: $37.85 fact
- Campus recreation: $33.50 fact
- Student Life and Wellness Center: $25.84 fact
- Transportation: $6.54 fact
- Arts: $2.25 fact
- Health: $0.00 fact
There is no separate technology line fact. UVU fee schedule
A separate tuition table reports $334.55 fact at 10 or more credits fact. The $0.47 difference unknown is not explained publicly. UVU tuition and fees
Conclusion: the base AI platform should receive $0 from general student fees est recommendation. As scoped, it includes academic support and central administration, so it does not fit the general-fee rules. A later, separately defined non-instructional service could have different eligibility unknown.
Peer cost-sharing patterns
- UC Berkeley gives faculty a common allocation while allowing faculty, grants, deans, or chairs to purchase added capacity. Campus supplies the shared infrastructure and administration. Berkeley Savio
- Yale provides standard compute without direct charges and sells optional priority capacity at $0.0049 per service-unit hour in FY2027 fact. Yale priority tier
- Princeton’s research-software partnership normally splits eligible staffing 50% fact program term with a research partner for one to three years fact program term. Princeton partnership guide
- USC’s condo model has research groups buy nodes, cables, and a five-year warranty fact, while the university supplies racks, network, power, cooling, room, and administration. USC condo model
Recommended funding model
- Central recurring IT funds the shared fleet, identity, security, network, monitoring, warranty and refresh reserve, baseline frontier pool, and accountable professional owner.
- Colleges and grants fund added nodes, specialist software, reserved capacity, and workload-driven staff growth.
- Use showback and project limits first. Use formal chargeback only for reserved, priority, burst, or unusually heavy use after actual usage and cost data exist.
- Do not charge ordinary teaching users per request.
- Budget $0 from the general student fee est recommendation for the base platform.
One-line college rule: “Add a college only after a signed three-year est commitment covers all marginal hardware and 50% est of marginal operating labor; central IT retains the shared fabric, security, baseline service, and accountable owner.”
Student staffing
Pay and structure
Public Utah evidence gives these anchors:
- A UVU skilled student project-lead posting paid $15–$16/hour fact, April 2026. UVU posting
- Weber State lists an $11.75/hour minimum fact as of January 2025, a 20-hour weekly cap fact for regular student jobs, and 10–28 hours weekly fact for internships. Weber State
- Utah Tech pays agency students $12/hour fact and charges departments a 25% markup fact, producing $15/hour fact arithmetic. Utah Tech
- UVU federal work-study positions publicly range from $12.00 to $21.65/hour fact for academic year 2026–27. UVU work-study
Use $18/hour est for technical student operators. The exact UVU technical wage schedule and payroll burden are unknown.
Students may handle intake, documented health checks, basic diagnostics, inventory, staging, evaluation runs, documentation, and workshop support. Professionals retain production access, security incidents, architecture, policy, purchasing, releases, and service accountability.
Credit-bearing option
UVU already has possible routes:
- INFO 4810R offers one to three credits fact, while INFO 4890R offers one to four credits fact for mentored research. UVU INFO courses
- IT 4810R offers one to three credits fact for related employment with a department coordinator. UVU IT courses
- TECH 2810R offers one to three credits fact for supervised professional experience. UVU TECH courses
Department approval remains required fact. Credit should recognize learning; it should not replace pay for scheduled operational work est recommendation.
Supervision and turnover
Northwestern uses paid graduate students for research-computing tickets, consultations, documents, and workshops, with at least 10 hours weekly fact and weekly staff mentoring. MIT has used students in an HPC, AI, and machine-learning help desk with staff escalation. UVA publicly identified 12 current and nine former student workers fact, observed September 2026. Northwestern, MIT, UVA
Planning allowance:
- 0.10 professional FTE per four active students est.
- 15% student onboarding and turnover reserve est.
- $20.70 effective student hour est arithmetic, using the $18 wage est plus the reserve.
- $43,056 per student FTE-equivalent est arithmetic, using 2,080 hours est planning convention.
The true UVU turnover rate and supervision load are unknown and should be measured during the pilot.
Effect on the plan’s staffing bands
The plan’s professional line is $144,118 per FTE annually est plan input. The proposed professional share includes supervision.
| Plan band | Current professional-only cost | Recommended accountable mix | Recommended annual cost | Change |
|---|---|---|---|---|
| 0.35 FTE est | $50,441.30 est | 0.35 professional + 0.20 student FTE-equivalent est | $59,052.50 est | +$8,611.20, or +17.1% est |
| 1.25 FTE est | $180,147.50 est | 0.75 professional + 0.50 student FTE-equivalent est | $129,616.50 est | −$50,531, or −28.1% est |
| 4.0 FTE est | $576,472 est | 2.50 professional + 1.50 student FTE-equivalent est | $424,879 est | −$151,593, or −26.3% est |
At the smallest band, students add coverage rather than replace the accountable professional. At the larger bands, they absorb repeatable work and lower professional staffing needs.
At 10–15 hours per week for 30 teaching weeks est schedule, likely staffing is:
- One to two students est for the smallest band.
- Three to four students est for the middle band.
- Seven to 11 students est for the largest band.
Paid summer coverage, overlapping handoffs, paired access, runbooks, and a professional on-call path are required. Actual student availability is unknown.
Grants and partnerships
The plan already covers HERFP, AI Moonshot, and NAIRR. The following are additional paths.
| Status | Program | Finding and fit |
|---|---|---|
| Open | NSF State and Regional AI Infrastructure Hubs, solicitation 26-513 fact identifier | Posted July 31, 2026 fact; deadline November 4, 2026 fact. NSF expects about 10 awards per cycle fact, with typical requests of $4 million–$12 million over five years fact and about $100 million available fact. Only one award per state or multi-state region fact is planned. This is the strongest infrastructure fit, but UVU would need a Utah or regional consortium with industry, government, philanthropy, and other universities. NSF AI Infrastructure Hubs |
| Open | NSF IUSE: EDU, solicitation 23-510 fact identifier | Next deadlines are January 20, 2027 fact and July 21, 2027 fact. Awards range from up to $400,000 to $2 million fact, with no voluntary cost share fact. Fit: evidence-based AI-supported undergraduate STEM teaching, faculty practice, or learning outcomes—not routine platform operations. NSF IUSE |
| Open | Department of Education IES research training and methods | Deadline October 1, 2026 fact. Maximums are $800,000 fact for research training and $900,000 fact for research methodology. Fit: student training and evaluation research, not fleet purchase. Department of Education available grants |
| Expected but not located | Utah AI Workforce Accelerator | Utah appropriated $3 million ongoing fact. A June 2026 fact presentation said an RFP would issue in August 2026 fact for curriculum modernization, embedded AI credentials, and expanded programs. The public RFP, award size, deadline, match, and applicant rules remain unknown. Utah update, Utah budget summary |
| Forecast | NSF Major Research Instrumentation | NSF expects a new solicitation during federal fiscal year 2026 fact, but has posted no current deadline unknown. The prior program allowed $100,000–$4 million fact for shared research instruments. A research-only shared AI instrument may fit; a general teaching service does not. NSF MRI |
| Not currently open | NSF Campus Cyberinfrastructure | The last program allowed up to $700,000 fact for campus compute and $1.4 million fact for regional compute. It is awaiting a new solicitation fact. Keep it on the watch list but budget $0 est from it now. NSF CC* |
| Selected partnership; application unknown | Apple Community Education Initiative | Apple says it can provide hardware, scholarships, financial support, curriculum, and expert help. Apple reported support for more than 200 education and community partners fact as of October 2024, but no public application, award range, deadline, or UVU eligibility was found unknown. Treat it as a partnership target, not forecast cash. Apple CEI |
| Invitation only | Apple Scholars in AIML | Supports nominated doctoral students for two years fact with research and travel support, mentoring, and internship access. UVU invitation status and public dollar value are unknown. Apple Scholars |
Philanthropy patterns
- The University of Florida assembled an $85 million package fact: a $25 million alumnus gift fact, $25 million from NVIDIA fact, $15 million from the university fact, and $20 million in recurring state support fact. Pattern: donor, vendor, university, and state each fund a different layer. University of Florida
- The University of South Florida received a $40 million naming gift fact for an AI and cybersecurity college, paired with a dollar-for-dollar challenge of up to $5 million fact. Pattern: named anchor plus matching campaign. USF
- RIT received a $24 million commitment fact supporting an endowed AI institute director, scholarships, and faculty endowments. Pattern: donors favor named people and durable student programs. RIT
- Cornell received a $10.5 million gift over five years fact to support researchers using shared Empire AI infrastructure. Pattern: fund the people and research program beside publicly backed compute. Cornell
Recommended advancement package: named student-operations fellowships, named applied-AI leadership, vendor-provided equipment or support, and a written university commitment to recurring operations. Electricity is a poor stand-alone donor proposition.
Electricity
UVU rate
UVU’s October 2024 fact emergency plan says its campuses receive power from Rocky Mountain Power and that most of the Orem main campus uses UVU’s north substation. UVU is also named in Rocky Mountain Power’s Schedule 34 clean-energy arrangement. That schedule adds contract-specific terms to the customer’s normal tariff; it does not publish UVU’s actual blended rate. UVU electricity plan, Rocky Mountain Power announcement, Schedule 34
UVU’s actual tariff, demand peaks, delivery voltage, clean-energy adders, fixed charges, and blended rate are unknown.
Rocky Mountain Power Schedule 8 currently lists:
- $76 per month fact fixed charge.
- $5.15 per kW fact facilities demand charge.
- $14.90–$16.84 per kW fact on-peak demand charge.
- 2.8070–6.2395 cents per kWh fact energy charges, before adjustments.
Public evidence does not establish that UVU’s relevant meter is billed on Schedule 8 unknown. Rocky Mountain Power Schedule 8
The best public proxy is the EIA’s June 2026 fact preliminary Utah commercial average of $0.1099/kWh fact, released August 26, 2026 fact. EIA
Check against the plan
The plan uses 26,280 hours over three years est, $0.1099/kWh fact proxy, and PUE 1.2 est.
| Device | IT draw | Facility energy over three years | Calculated cost | Plan line |
|---|---|---|---|---|
| Mini | 140 W est | 4,415.04 kWh est | $485.21 est | $485 est |
| Studio | 145 W est | 4,572.72 kWh est | $502.54 est | $503 est |
| Ultra | 270 W est | 8,514.72 kWh est | $935.77 est | $936 est |
The plan’s power lines reconcile to the public commercial proxy after rounding. They remain estimates because utilization, actual machine draw, demand charges, and UVU’s contract rate are unknown.
Cooling and facility overhead
PUE includes cooling, UPS and distribution loss, lighting, and other support loads; it is not cooling alone.
At $0.1099/kWh fact proxy:
- Plan PUE 1.2 est adds 1,752 kWh per sustained IT kW-year est arithmetic, costing $192.54 annually est.
- The Uptime Institute’s 2026 fact industry-average PUE of 1.52 fact adds 4,555.2 kWh per sustained IT kW-year est arithmetic, costing $500.62 annually est. Uptime Institute
- UVU’s cooling-only cost per sustained IT kW is unknown.
Budget wording: “Facility overhead is estimated at $193 per sustained IT kW-year est under the plan, with a $501 est industry-average sensitivity. UVU cooling-only cost is UNKNOWN.”
What the plan should change
- 1. Use a year-four default refresh with a year-three measured review gate. This preserves one extra year of use without assuming local hardware must track every frontier release.
- 2. Add the five-year cash cases. Show $90,341 est with no refresh, $122,438 est with a year-three refresh, and $128,233 est with a year-four refresh, plus the support and terminal-value gaps.
- 3. Carry three-year residuals of 50% mini, 50% Studio, and 55% Ultra est. Add a separate, currently unknown, disposition-cost haircut before treating resale as spendable cash.
- 4. Buy the first fleet unless a written lease quote beats ownership on UVU’s real terms. Require the rate, fees, guaranteed residual, return standards, early-exit terms, and ownership option before approval.
- 5. Fund the shared base centrally. Colleges and grants pay for marginal capacity; charge only for optional reserved or heavy use. Use $0 from general student fees est recommendation.
- 6. Keep a professional owner and add paid students underneath. The smallest band gains coverage; the middle and large bands can reduce annual staffing cost by about 28.1% and 26.3% est, subject to a measured pilot.
- 7. Pursue the NSF AI Infrastructure Hub as a consortium, not a solo hardware request. Pair that with an IUSE teaching study, the expected Utah workforce opportunity, and a named philanthropy package.
- 8. Replace “UVU electricity rate” with “public Utah proxy.” Keep $0.1099/kWh fact proxy until Facilities supplies a bill, demand history, tariff, and measured PUE.
Source log
- 1. Plan engine assumptions — dated September 2, 2026 fact.
- 2. Model-progress findings — dated September 2, 2026 fact.
- 3. Mac generation findings — dated September 2, 2026 fact.
- 4. Apple and eBay mini residual evidence — observed September 3, 2026 fact.
- 5. Apple M2 Studio launch prices — June 2023 fact.
- 6. Apple Certified Refurbished store — observed September 3, 2026 fact.
- 7. Back Market Mac mini category — observed September 3, 2026 fact.
- 8. Apple Financial Services — accessed September 3, 2026 fact.
- 9. Apple education-finance flyer — current public flyer accessed September 3, 2026 fact.
- 10. Apple Utah education contracts — accessed September 3, 2026 fact.
- 11. NASPO Apple contract — accessed September 3, 2026 fact.
- 12. Utah equipment-lease policy — revised September 27, 2023 fact.
- 13. USHE R516 — amended March 26, 2026 fact.
- 14. UVU Policy 511 — effective December 12, 2024 fact.
- 15. UVU student-fee schedule — academic year 2026–27 fact.
- 16. Berkeley Savio — accessed September 3, 2026 fact.
- 17. Yale research-computing priority tier — fiscal year 2027 pricing fact.
- 18. UVU skilled student posting — opened April 10, 2026 fact.
- 19. UVU work-study pay — academic year 2026–27 fact.
- 20. Northwestern student consultant program — accessed September 3, 2026 fact.
- 21. NSF AI Infrastructure Hubs — posted July 31, 2026 fact.
- 22. NSF IUSE — active September 3, 2026 fact.
- 23. Department of Education available grants — accessed September 3, 2026 fact.
- 24. Utah AI Accelerator update — presented June 17, 2026 fact.
- 25. Apple education-grant expansion — published October 1, 2024 fact.
- 26. University of Florida AI funding case — published May 2026 fact.
- 27. UVU loss-of-electricity plan — updated October 2024 fact.
- 28. Rocky Mountain Power Schedule 8 — effective August 10, 2026 fact.
- 29. EIA Utah commercial electricity price — June 2026 data released August 26, 2026 fact.
- 30. Uptime Institute PUE benchmark — published August 6, 2026 fact.
NOT_RUN
- UVU-specific Apple financing or Guaranteed Buyback quote.
- UVU procurement, legal, finance, student-fee, facilities, payroll, or HR review.
- Vendor, utility, university, donor, or grant-office contact.
- Review of a live UVU utility bill, demand history, Apple lease schedule, or internal wage table.
- Hardware inspection, seller-fee calculation, tax analysis, or bulk-disposition quote.
- Student staffing pilot or time-and-motion study.
- Grant application, eligibility ruling, account access, purchase, deployment, publication, or external action.
- The excluded University of Utah figure was not used.
JSON grades: residual percentages est; lease rate unknown; student wage est; electricity rate fact public proxy, actual UVU rate UNKNOWN; fee eligibility no for the base service as scoped.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/42-money-angles.md in the research pack.
Assessment redesign and academic integrity
Appendix V in five lines
- QuestionWhat assessment rules protect learning and honesty without letting flawed artificial-intelligence detectors punish students?
- AnswerLabel each key task “Secure” or “Open,” add a small human skill check when needed, and never treat a detector score as proof.
- Deciding numbersaverage 61.3% false-flag rate for 91 essays by non-native English writersfact59 of 63 unedited artificial-intelligence answers, or 93.7%, drew no concernfact$20,000–$40,000 pilotest
- What the plan doesReplace three templates with one task standard, label key work before class opens, run one redesign in each college, and measure time, fairness, and trust.
- Still unknownThe unsourced “91% of faculty concerned” figure has been removed from the plan; campus integrity records are UNKNOWN, and local detector and access tests are NOT_RUN.
Faculty's leading fear about campus AI is cheating, and the plan's answer had been three template policies and the phrase "lead with assessment redesign." The owner asked for depth. This appendix is the evidence-based playbook: what actually happened to integrity after 2022, what AI detectors get wrong and for whom, the assessment designs that hold up at 300-student scale and what they cost faculty, the policy patterns Utah's peers use, an example redesign for each of UVU's seven colleges, and what it takes to roll out. Research cutoff September 3, 2026; 160-plus dated sources. fact measured · est judged · unknown not settled.
What this changes in the plan
1. One assessment-assurance standard replaces "three templates." Every material assessment gets an assignment-level Secure/Open label, the allowed and prohibited AI functions, a disclosure rule, and an individual proof point where the learning outcome demands it — set at the course-shell moment, four to six weeks before term, with a report of unlabeled high-stakes tasks. 2. Detector output alone is never evidence — it cannot by itself cause a grade change, referral, interview, or sanction. The independent tests show detectors unstable across models, easy to evade, and uneven against non-native English and neurodivergent writers; several universities have switched them off. 3. The data do not show an integrity collapse. Self-reports, detector flags, referrals, and proven cases measure different things and must never be blended into one "cheating rate"; where institutions published totals, cases did not jump the way the fear predicts. The plan's "91% of faculty concerned" figure could not be traced to a survey and is removed until it can be. 4. Secure assessment stays sparse. A few oral, written, coding, or practical checks where UVU must prove individual skill; everything else open, applied, staged, and explicit about allowed help. Blanket redesign would burden honest students, adjuncts, and online learners most. 5. Each catalyst's paid deliverable becomes one redesigned signature assessment, run, measured for workload and student trust, and published as a design card — Engineering a code defense, Business a changed-fact decision task, Humanities staged research plus close reading, Education microteaching, Science a lab defense, Health a synthetic practical, Arts a provenance-rich portfolio with live critique. 6. A fair follow-up protocol: the student sees the rule and the evidence, may bring drafts and explain assistive tools, is asked questions aligned to the original task, gets written reasons, and keeps an appeal. 7. Success is measured by substantiated cases, overturns, validity, workload, access, and trust — never by detector flags or surveillance events. The $20,000–$40,000 first stage becomes a measured gate before any $150,000–$250,000 expansion. est
Research date: September 3, 2026 MDT. Evidence key: fact = dated source; est = calculation or planning assumption with its basis; unknown = evidence was unavailable, conflicting, or too weak.
Executive verdict
- 1. Generative AI use in assessed work is now common, but the evidence does not show one clean post-2022 jump in cheating.
- 2. The plan’s “91% concerned” figure has unknown provenance: the supplied files name no survey, date, sample, or question.
- 3. Student self-reports, detector flags, referrals, and proven cases measure different things and must never be combined into one “cheating rate.”
- 4. AI detectors are too unstable, too easy to evade, and too uneven across writers and subjects to prove misconduct.
- 5. UVU should adopt the rule: detector output alone is never evidence and must never trigger a penalty.
- 6. Assessment redesign should not mean replacing every assignment with an exam.
- 7. UVU should use a few secure oral, written, coding, or practical checks where it must prove individual skill.
- 8. Other work should remain open, applied, staged, and explicit about allowed AI help and disclosure.
- 9. The course-shell build period is the right control point; every material assessment should receive an AI-use label before students enter the course.
- 10. Start with one paid, measured redesign in each college, publish the results, and expand only where validity, workload, access, and student trust improve.
What actually happened to integrity, 2023–2026
The central finding
unknown — There is no credible national college series that measures the same misconduct behaviors, at the same institutions, with the same rules and adjudication process, before and after generative AI.
The available evidence shows three different things:
- Students use AI much more.
- Some students report submitting AI-written material as their own.
- Recorded misconduct depends heavily on policy wording, detector use, reporting effort, and when cases are counted.
It does not support “AI increased college cheating by X%.”
Self-report data
| Evidence | Result | What it means |
|---|---|---|
| Challenge Success high-school cohorts | FACT—2019 and 2023 samples: prior-month dishonest behavior ranged from 59.9% to 64.4% in three 2023 schools. Comparable 2019 school samples ranged from 61.3% to 82.7%. The researchers found no increase attributable to ChatGPT. Study | Broad high-school misconduct, not college AI cheating. Cohorts and schools differed. |
| Challenge Success follow-up | FACT—February–May 2024: 72.06% of 4,354 students at six high schools reported at least one dishonest behavior. 24.27% reported unauthorized “AI or digital device” use, but the question cannot isolate AI. Study | High background misconduct persisted. It is not a clean year-to-year comparison. |
| Australian university survey | FACT—2024: unacknowledged verbatim GenAI copying was reported by 13.9% of 310 Western Sydney University respondents and 14.2% of 1,867 students at five other universities. Detector-using and non-detector institutions did not differ significantly, fact p=.097. Study | Direct self-report evidence of misuse; no evidence that detectors reduced it. |
| HEPI UK series | FACT—2024: 53% of 1,250 undergraduates used GenAI for assessed work; 5% inserted unedited AI text. FACT—2025: 88% of 1,041 used it for assessment and 18% included AI text directly, including edited text. FACT—December 2025 survey published 2026: 94% of 1,054 used it for assessed work and 12% directly included AI text, versus 8% in 2025 and 3% in 2024 under the comparable narrow measure. 2024, 2025, 2026 | Assessed use rose sharply. Much of it—explanations, summaries, ideas, and editing—may be permitted. |
| Large US public-university survey | FACT—spring 2024: 95,513 students across 20 public research universities responded; about 37% used GenAI at least monthly. About 9% of GenAI users reported submitting generated work as their own. Science study | Strong cross-campus evidence, but still self-report and not a before/after cheating series. |
| ICAI | FACT—public page checked September 3, 2026: ICAI reports an updated survey pilot with 840 students in 2020, but publishes no comparable 2023–2026 national AI-era rate. ICAI | Old McCabe-era percentages must not be presented as current AI-era results. |
Detector-flag data
Turnitin reported that, through March 21, 2024, it had screened more than 200 million submissions; more than 22 million, about 11%, scored at least 20% likely AI writing, while more than 6 million, about 3%, scored at least 80%. These are FACT—vendor-reported flag counts, not findings of misconduct. Submissions are not unique students; authorized AI use is included; repeat submissions and the population mix are undisclosed. Turnitin
The plan must never convert those figures into “11% cheated” or “3% were caught.”
Adjudicated and institutional cases
| Institution | Recorded result | Interpretation |
|---|---|---|
| University of Hull | FACT—2022/23: 85 AI-related referrals among 802 total cases; 56 were guilty and 12 remained open. FACT—2023/24 through April 25, 2024: 84 AI referrals among 389 total, with 39 guilty and 28 open. Hull response | Partial years and unresolved cases prevent a clean trend. |
| Deakin University | FACT—2024: 2,601 breach instances involving 1,535 students; 687 instances, or 26%, involved GenAI misuse. Deakin report | Instances are not unique students. Deakin warns that recording and detection changes affect the total. |
| UNSW | FACT—2023: 166 serious GenAI referrals were reported and substantiated, while total recorded cases fell 16% from 2022. FACT—2024 closures: 1,983 of 2,154 cases were substantiated or partly substantiated; the detailed table contains 640 AI-related outcomes, while the narrative reports 530 AI cases. 2023 report, 2024 report | The report does not reconcile the two 2024 AI totals. New categories and closure timing materially affect the result. |
| University of Manchester | FACT—2023/24: schools reported 457 academic-misconduct cases, versus 477 in 2022/23. Only 15, or 3%, were recorded in the new AI category. Manchester report | A new category and formal-case-only reporting make undercounting likely. It does not show a general rise. |
How much is detection artefact?
A large share is unknown, but at least five artefacts are documented:
- A new AI category creates an apparent rise from a zero category that did not previously exist.
- Institutions count different units: referrals, assignments, breach instances, cases, students, closures, or sanctions.
- More faculty training and detector availability produce more referrals even if behavior stays flat.
- Permitted use can generate a detector flag without misconduct.
- Open cases move between years; one student or submission may create several breach records.
UVU should publish separate rates for:
- Student self-report.
- Instructor concerns or referrals.
- Detector-assisted referrals, if any.
- Cases opened.
- Cases substantiated.
- Cases overturned or returned on appeal.
The denominator should be students or eligible submissions, not just raw cases.
AI detectors
Measured performance
- FACT—2023 controlled test: seven detectors incorrectly flagged an average 61.3% of 91 human TOEFL essays written by non-native English writers. All seven flagged 18 of 91, or 19.8%, and at least one flagged 89 of 91, or 97.8%. The control was 88 US eighth-grade essays. Liang et al.
- FACT—2023 controlled test: 14 tools, 54 documents, and 756 tests produced less than 80% accuracy for every tool. Only five exceeded 70%. False-positive rates ranged from 0% to 50% and false-negative rates from 8% to 100% across tools and conditions. Paraphrasing sharply reduced detection. Weber-Wulff et al.
- FACT—2024 experiment: Turnitin identified 47 of 50 raw AI medical articles but only 15 of 50 after paraphrasing. Human reviewers caught more paraphrased articles but each mislabeled 12% of 50 human articles. Liu et al.
- FACT—2024 live university test: researchers submitted 63 unedited AI answers through 33 false student accounts at the University of Reading. 59 of 63, or 93.7%, received no AI concern. PLOS ONE
- FACT—2025 study: among 250 known-human neurosurgery articles, one detector flagged 30.4%, another 16.0%, and another 0% above the study threshold. Erol et al.
- FACT—2026 study: in a balanced 192-text corpus, Originality achieved macro accuracy .69 and Turnitin .61; both had macro F1 below .55. Turnitin accuracy ranged from .86 in humanities to .51 in science. This study found no significant Turnitin difference between its EFL and professional-writing samples, fact p=.50. Hadra et al.
The evidence is contradictory in detail but consistent in decision relevance: performance changes with tool, model, editing, genre, length, language, and test design.
Vendor claims versus independent tests
Turnitin initially claimed FACT—February 13, 2023 vendor claim 97% detection with a false-positive rate below 1 in 100, without publishing the test denominator in its release. Release
It later reported a less-than-1% document false-positive rate above its 20% score threshold, based on 800,000 pre-ChatGPT documents, and about 4% at sentence level. Those are FACT—vendor tests, not independent field validation. Turnitin also says false positives cannot be eliminated and the score must not be the sole basis for adverse action. Update, current guidance
Bias and access
The non-native-English evidence is strong enough to establish a material risk, but not to say every current detector is always biased. The 2026 study above found poor overall and genre-dependent performance without a significant EFL difference in its small sample.
Neurodivergent-writing evidence is thinner:
- FACT—July 2026 preprint: a corpus of 59,947 Reddit posts had flag rates of 1.9% for a “likely autistic” group and 1.5% for general Reddit; the modeled odds were 25% higher before length normalization and 50% higher in a fixed-length analysis. It tested an older GPT-2 detector, inferred autism from subreddit participation, and was not peer-reviewed at the cutoff. Preprint
- unknown — No controlled peer-reviewed study located by September 3, 2026 estimates current commercial-detector false-positive rates for autistic, ADHD, dyslexic, or broadly neurodivergent UVU students.
That gap supports caution. It does not support inventing a bias percentage.
Universities that disabled or rejected detection
- Vanderbilt disabled Turnitin’s detector on August 16, 2023, citing opacity, false positives, non-native-writer bias, and privacy. Vanderbilt
- Bournemouth declined activation on April 13, 2023 because it lacked time to test validity, fairness, policy, and communications. Bournemouth
- Macquarie disabled it for its second 2023 session. Macquarie
- Curtin announced on September 4, 2025 that it would disable AI detection from January 1, 2026. Curtin
- University of Queensland disabled it in mid-2025. UQ
- UCL, Leeds, Glasgow, Cambridge, Oxford, and the University of the Arts London publish similar warnings against detector-based decisions.
Legal and process exposure
UVU Policy 541 presumes the student not responsible and requires the university to prove every element by a preponderance of evidence—more likely than not. It also requires a reliable, impartial investigation and access to appeal. UVU Policy 541
A detector probability does not prove:
- What tool produced the text.
- Who operated it.
- Whether use was permitted.
- Which passages were contributed.
- Whether an assistive tool explains the result.
- Whether the student crossed the stated assignment boundary.
The UK higher-education ombudsman published FACT—July 2025 six selected AI cases: two justified, one partly justified, two not justified, and one settled. Complaints succeeded where institutions withheld evidence, did not examine drafts or the writing process, or failed to explain detector reliability. In one autistic student’s case, reconsideration ended with no misconduct finding. Case set, autistic-student case
In Harris v. Adams, a US federal court declined to block discipline involving AI-assisted work. But the record included copied material, fabricated citations, revision history, instructions, admissions, and an opportunity to respond—not a detector alone. Decision
unknown — No reported US judgment located by the cutoff establishes civil liability or damages specifically for a detector-only university accusation. The risks are still real: an unsupported finding can violate UVU’s own evidentiary process and create disability, national-origin, privacy, contract, and reputational disputes. That is risk analysis, not a legal opinion.
UVU’s rule
Detector output alone is never evidence of misconduct. UVU will not penalize, lower a grade, or require a student to disprove a detector score. A concern must be tied to a clear assignment rule and supported by independently reviewable evidence.
Independently reviewable evidence may include the assignment rule, inaccurate or fabricated sources, the submitted artifact, voluntarily supplied drafts, task-created process records, a fair student explanation, or a short follow-up demonstration aligned to the original learning outcome.
A student’s silence, disability status, accent, writing style, or inability to recall unrelated details is not proof.
Assessment designs that hold up
“Validity” means the task measures the knowledge or skill it claims to measure. “Secure” means outside help can be controlled or the student must directly demonstrate individual skill.
| Design | Validity and integrity evidence | Faculty workload | Scale to 300 | Student reception and named adoption |
|---|---|---|---|---|
| Oral defense or viva | Strong when prompts and scoring are structured and several questions sample the intended skill. A Melbourne model achieved reliability of fact α=.82 and .77 in cohorts of 319 and 342. Study | High for long one-to-one exams. EST—arithmetic: 300 students × 5 minutes ÷ 60 = 25 contact hours before transitions, training, and accommodations. A 20-minute defense is est 100 contact hours. | Yes, with short structured checks, trained TAs, labs, or sampling. | DCU piloted 322 participants in 2020–2021 and later reported use across 18 modules and more than 2,000 students. FACT—survey n=140: 82% said it encouraged integrity and 77% said it discouraged cheating. Anxiety usually needs practice. DCU |
| In-class or supervised writing | Secures conditions, but timed recall, handwriting, typing speed, language fluency, and anxiety can displace the intended construct. In a randomized crossover, computer essays scored fact about 6 percentage points higher mainly because students wrote more. Study | Room, proctor, device, accommodation, and make-up costs are unknown for UVU. | Yes operationally; validity depends on the outcome. | UiT’s 2025 secure redesign produced fact 33 failures among 179 students, versus 4 among 187 in the prior take-home year, but AI, time pressure, closed-book conditions, and anxiety cannot be separated. UiT |
| Process portfolio and version history | A portfolio can validly sample work over time. Version history records actions, not thought or authorship; content can be pasted, reconstructed, or produced offline. | High when faculty inspect every draft. A Sydney portfolio used fact 372 raters for 257 students; rater and task effects were large. Study | Partial. Portfolio scoring reached fact n=1,208 at Idaho, but manual version-history audits at 300 are unknown. | Students often value reflection and feedback but dislike paperwork and surveillance. Version history should be limited to records naturally created by the assignment. |
| Two-lane assessment | Strong policy logic: secure checks assure individual outcomes; open tasks teach responsible tool use. No published causal evaluation yet proves improved validity or lower misconduct. | Secure rooms, staff, alternative sittings, and accessible formats can be expensive. UVU cost is unknown. | Yes operationally. Sydney adopted it institution-wide from 2025; Bath from 2026/27; UQ from its second 2026 semester. | Student trust and workload outcomes remain unknown. Policy reach is not effectiveness evidence. |
| Authentic and applied tasks | Strong for relevance and transfer when the task mirrors real work. Weak as a stand-alone authorship control: a study of fact 419 contract-cheating tasks found outsourcing at every measured authenticity level. Study | Low to medium when one shared scenario and rubric are reused; high when every student needs a different client or dataset. | Yes. | Students often prefer meaningful work, but comparative reception evidence is thin. Pair with an oral, observed, or supervised sample when individual competence matters. |
| Staged assignments and checkpoints | Good for feedback and revision. Weak proof of authorship unless at least one stage is observed or defended. | One engineering comparison found fact 23% more assessment workload. EST—arithmetic: 3 checkpoints × 3 minutes × 300 students ÷ 60 = 45 staff hours. Peer feedback can reduce the bottleneck. | Partial. PeerStudio served fact more than 3,600 learners, but secure individual assurance at that scale remains unproved. | Rapid feedback supports revision. Repeated compliance uploads can feel like busywork. |
| Reflective component | Useful for explaining choices when attached to an artifact or observed event. Not secure alone: in a Monash dataset, AI reflections outscored student reflections and human authorship classification was only fact .5489 and .6800. Study | Low for short rubric-based notes; high for rich narrative feedback. | Yes as a component, not as sole proof. | Reception depends on relevance and feedback. Generic “what did you learn?” writing is easy to outsource. |
| Code review and pair-programming assessment | Pair programming can support learning but cannot prove both partners’ skill. Individual live explanation, modification, debugging, or code review is stronger. In a 2026 pilot of fact n=541, oral project scores correlated r=.45 with a proctored midterm, versus r=.15 for code correctness. Report | UIC used three fact 10–20-minute interviews during existing labs, with each interviewer handling no more than 8 students. UIC | Yes with a TA-rich lab structure. UIC reports courses up to 300. | The first interview was stressful; repeated experience reduced anxiety. Pair-programming learning results are mixed. |
| Lab practical or OSPE | Strong when the outcome is equipment use, safety, observation, technique, or real-time interpretation. It directly samples performance but can undersample broad knowledge. | High, concentrated staffing and equipment. A Brunel model used an individual fact 20-minute, double-marked station; authors considered total work comparable with extended report marking. Study | Partial. Modern cohorts were 142 and 138; a contemporary single-section 300-student implementation was not found. | Karolinska students generally judged an OSPE valid, but only fact 99 of 198 answered the survey. Stress and communication accommodations matter. Study |
Design rule
No single design solves the problem:
- Authenticity does not prove authorship.
- Version history does not prove thought.
- Reflection does not prove introspection.
- Pair work does not prove individual skill.
- Supervision does not automatically produce a valid task.
- One oral answer does not sample an entire course.
The safest pattern is triangulation: an open product plus one short, structured individual check and a clear rule set before submission.
Policy patterns
Scales, traffic lights, and two lanes
The AI Assessment Scale now uses five levels: No AI, AI Planning, AI Collaboration, Full AI, and AI Exploration. Its authors warn against attaching labels to unchanged tasks or using unenforceable “No AI” rules. Framework
British University Vietnam reported FACT—January 2023 112 penalties among 1,722 submissions, or 6.50%; FACT—October 2023 0 among 3,996; and FACT—January 2024 4 among 4,159, or 0.10%, after AIAS implementation. The study also reported a 5.9% mean-grade increase and 33.3% higher module pass rate. These are FACT—reported observational outcomes, not causal effects: the institution, rules, communication, checking, and assessment designs all changed, and there was no control group. Study
UCL uses three categories—cannot use, assistive use, and integral use—and says a cannot-use task should normally be secure. It explicitly calls these categories guidance rather than formal policy. UCL
Sydney rejected fine-grained restrictions in unsecured work and adopted open and secure lanes. Bath and Leeds are also moving away from older traffic-light systems toward clearer outcome-based rules. Sydney, Bath, Leeds
Recommended UVU pattern
Reuse UVU’s existing levels, but place them inside two enforceable lanes:
- Secure: AI-free except named assistive technology or approved tools.
- Open—Assisted: editing, explanation, brainstorming, translation, or feedback only.
- Open—Guided: named AI work is required within instructor-set steps.
- Open—Collaborative: AI may contribute content, code, analysis, or media; students must verify and disclose it.
- Open—Integral: skilled AI use is itself a learning outcome.
Do not rely on a course-wide label alone. Every high-stakes task needs its own label.
Best wording
AI use for this assessment: [Secure / Open—Assisted / Open—Guided / Open—Collaborative / Open—Integral].
You may use AI for: [named functions].
You may not use AI for: [named learning work].
If you use it, disclose the tool, purpose, material contribution, and what you checked or changed. You remain responsible for every submitted claim, source, calculation, and artifact.
In a secure task, no AI or outside help is allowed except [approved tools and accommodations].
Ask before submission if the boundary is unclear.
A detector score is never evidence of misconduct.
Do not require complete prompt histories by default. They may contain private data, unrelated work, or inaccessible tooling. A short material-contribution disclosure is more useful.
UVU’s current policy position
- FACT—live manual checked September 3, 2026: no adopted UVU policy titled “Artificial Intelligence” appears in the manual. Manual
- A proposed Policy 441 executive summary dated September 25, 2025 was a proposal to begin drafting and its file now returns 404. Current Policy 441 is the computing-facilities policy. UVU should not reuse the number for a second subject.
- Policy 541 controls misconduct definitions, the preponderance standard, investigation, sanctions, and appeal. Policy 541
- Policy 445 governs confidential and restricted data, including student records. Policy 445
- Policy 447 governs university-hosted and third-party systems, access, encryption, logs, backup, and production authorization. Policy 447
- Policy 452 requires accessible electronic technology and equally effective alternatives. Policy 452
- UVU’s 2026 library guidance says there is no universal course AI rule and students must ask each instructor. Library guidance
- UVU’s Writing Center already publishes assignment-use levels and disclosure guidance. Faculty guide
- CHSS already publishes a tiered faculty guide and warns about detector use. CHSS guide
Where the assessment rule should sit
Use three linked layers:
- 1. Academic Affairs assessment standard: require every course to state its default and every material task to carry a Secure/Open label.
- 2. Policy 541 amendment or binding interpretation: define undisclosed or prohibited AI use as unauthorized assistance, preserve the preponderance standard, and bar detector-only findings.
- 3. Implementation guide: reusable assignment text, disclosure form, examples, accessibility checks, and fair follow-up procedures.
Policies 445, 447, and 452 should govern any associated tools or records; they should not contain the academic assessment rule itself.
Utah and peer practice
- FACT—December 2, 2025: the Utah Board of Higher Education called for responsible AI, literacy, workforce readiness, and preservation of human learning, but created no statewide assessment scale. USHE
- BYU’s central policy remains tool-neutral and makes faculty responsible for communicating expectations; its Honors Program adopted a stricter written AI rule in 2024. BYU
- USU’s 2026–27 catalog prohibits unauthorized assistance; teaching guidance says permissions should be clear and reasonable doubt may weigh against filing a case. USU
- Weber’s code is tool-neutral. Its teaching center offers optional prohibit, limited-use, and full-use examples. Weber
- SLCC expressly includes unauthorized AI on exams and unacknowledged AI words or ideas in its Student Code. SLCC
- SUU requires every syllabus to explain the allowed extent of AI use. SUU
- Snow College publishes prohibited, guided, and cited-use templates and says detector results are only a starting point. Snow
- Utah Tech treats AI use as misconduct when the instructor prohibited it in writing and treats uncredited AI output as plagiarism. Utah Tech
The strongest Utah pattern is not a universal ban. It is clear written permission, task-level boundaries, disclosure, and ordinary due process.
Per-college patterns for UVU
These are proposed designs, not verified current college rules.
| College | Best-fit pattern | Example redesign |
|---|---|---|
| Engineering & Technology | Open—Collaborative build plus secure individual code review, debugging, calculation, or design defense. | Teams build a working system with AI allowed and disclosed. Each student receives a small defect or requirement change and must diagnose, modify, test, and explain it live. Grade the shared product and the individual check separately. |
| Business | Open—Collaborative analysis plus secure decision memo or board-style questioning. | Students use AI to analyze a supplied market packet, disclose material contributions, verify sources, and present a recommendation. A short closed follow-up gives a changed fact and asks each student to revise the decision. |
| Humanities & Social Sciences | Open—Assisted or Guided staged research plus secure close reading or oral defense. | Require a question, annotated sources, evidence map, draft, and revision note. Permit brainstorming and feedback but not fabricated sources. Add a short in-class analysis of a new passage or a structured defense of two major choices. |
| Education | Open—Guided lesson design plus secure microteaching and defense. | Students may use AI to generate alternative lesson ideas, then critique them against standards and learner needs. They teach part of the lesson, respond to a learner misconception, and explain the chosen adaptation. |
| Science | Open—Collaborative analysis plus observed lab skill and method defense. | Students may use AI for code or interpretation after collecting data. They must maintain the normal lab record, perform one assigned technique, explain controls and uncertainty, and identify an intentionally flawed result. |
| Health & Public Service | Secure simulation or practical for safety-critical skill; open guided critique for documentation and improvement. | Use synthetic cases. The student completes an observed response, handoff, interview, or procedure, then may use AI to critique documentation. AI output is never a clinical authority. |
| Arts | Open—Integral or Collaborative portfolio plus live critique, performance, or technique demonstration. | Students disclose generated or transformed media, source rights, consent, and material AI contribution. The final review includes process artifacts and a live explanation or performance showing decisions and craft. |
The common design is product plus proof. The product may use modern tools. The proof samples the human learning outcome.
Making it happen
Cost and workload
UVU’s exact redesign cost is unknown until the university measures paid faculty and support hours.
Useful published reference points are:
- Cornell offers up to fact 20 fellows per academic year with a $5,000 stipend per fellow. Cornell
- Duke’s 2024–25 pilot offered fact $2,000–$7,500 per project plus a $1,000 faculty stipend for up to six courses. Duke
- Missouri proposed fact $64,590 for 20 fellows at $3,000 each plus program costs; larger proposed packages were $913,813 and $1,901,151. These were requests, not verified spending. Missouri
- Michigan reported fact 318 sign-ups for its 2024–25 self-paced Teaching with GenAI resource. Sign-up is not completion or redesign. Michigan
- Sydney reports fact about 30,000 assessments and more than 2.2 million submissions in its broader assessment framework. Central system changes, grants, workshops, consultations, and a community of practice supported the work, but AI-specific cost and outcomes are not isolated. Sydney framework
The current UVU plan’s $20,000–$40,000 one-college pilot is EST—plan basis: 30 faculty × $500 = $15,000, leaving est $5,000–$25,000 for leads, design support, administration, and evaluation. The $150,000–$250,000 campus figure is also est; no detailed staffing or workload model is supplied.
Keep the first range, but do not approve the campus range until the pilot measures:
- Faculty redesign hours.
- Review and calibration hours.
- Staff time per secure check.
- Accommodation and make-up time.
- Student completion and appeal burden.
- Reusable-template savings in the next term.
Minimum viable rollout
Before course-shell copy
- Academic Affairs approves the shared lane and label language.
- Student Conduct approves the evidence and follow-up procedure.
- Accessibility reviews secure alternatives.
- Canvas receives the label block and disclosure field.
- Colleges select one signature assignment, not an entire course.
At course-shell copy
The supplied plan places this moment FACT—plan timing about 4–6 weeks before term. Every catalyst should:
- Label each material assessment.
- State allowed and prohibited functions.
- Add one individual proof point where needed.
- Add an accessible alternative.
- Remove any detector-dependent wording.
- Estimate staff minutes per student.
During the term
- Give students a practice oral, practical, or secure mini-task before the graded version.
- Use a shared rubric and calibration examples.
- Record only data needed for assessment and review.
- Pay adjunct faculty for redesign and required training.
- Let catalysts hold short discipline-specific clinics.
After the term
Each catalyst publishes a de-identified design card containing the old task, new task, AI-use lane, learning outcome, workload, student response, integrity cases, limitations, and recommendation.
What to measure
Use measures that test the plan, not students’ private behavior.
Integrity
- Referrals per est reporting unit 1,000 eligible submissions.
- Substantiated findings per the same denominator.
- Evidence types used.
- Case resolution time.
- Appeal, remand, and overturn rate.
- Policy-clarity errors.
Assessment validity
- Agreement between open-product scores and secure individual checks.
- Rater agreement on sampled tasks.
- Performance on the next course, placement, licensure, or capstone outcome where available.
- Failure and withdrawal changes, with access and accommodation review.
- Whether the assessment still covers the stated learning outcome.
Student trust
Use an anonymous survey covering clarity, fairness, fear of false accusation, ability to ask questions, equal tool access, accommodation, and whether the work felt meaningful.
Faculty feasibility
- Redesign hours.
- Delivery and marking minutes per student.
- Make-up and appeal hours.
- What could be reused next term.
- What faculty stopped doing to make room.
Do not measure keystrokes, continuous screen activity, private prompts, browser histories, writing “style,” device telemetry, or unvalidated detector flags.
Cross-domain learnings
Medicine
Medical licensing combines knowledge with direct, structured clinical performance. Central standards, trained assessors, common cases, checklists, and review protect validity. The lesson for UVU is to sample actual performance at important assurance points, not turn every class into a practical exam.
Aviation
FAA certification combines a knowledge test, oral questioning, scenario judgment, and demonstrated performance. FACT—current standards effective May 31, 2024 integrate knowledge, risk management, and skill. FAA standards
The lesson is “artifact plus explanation plus action.” Access to modern cockpit tools does not remove the need to demonstrate judgment.
Apprenticeships
England’s end-point assessments normally combine at least two methods such as observation, demonstration, test, interview, or viva. UK guidance
The useful transfer is independence of evidence: a workplace product is meaningful, but an assessor also watches or questions the learner.
Software security
NIST’s zero-trust model says trust should not arise merely from network location or ownership. NIST
The assessment analogy is limited but useful: do not assume authorship because a file came through Canvas. Verify the particular learning outcome at a proportionate assurance point.
Quality assurance
Regulated fields do not rely on one noisy signal. They combine records, direct observation, structured questions, sampling, review, and appeal. UVU should do the same. A detector is closer to an unverified alert than to proof.
Devil’s advocate
The strongest case against this recommendation is that the cure may burden honest students more than dishonest ones.
Secure exams can reward speed and anxiety control instead of deep learning. Oral work can penalize speech disabilities, language learners, trauma, and cultural communication differences. Practical exams need rooms, staff, equipment, make-ups, and calibration. Version history can become surveillance. Frequent checkpoints increase adjunct labor. Students with jobs, care duties, long commutes, or online enrollment may bear the largest cost.
The institutional data also do not prove an integrity collapse. Manchester’s total cases fell; Challenge Success did not observe the feared post-ChatGPT jump; Australian self-report found no significant misconduct difference between detector and non-detector institutions. A campus-wide redesign mandate could spend scarce faculty time on a problem whose true UVU size is still unknown.
This argument defeats blanket redesign. It does not defeat targeted assurance. The proportional response is a small number of secure checks tied to degree-critical outcomes, with open assessment elsewhere and measured equity effects.
What would change the recommendation
The recommendation should change if any of these occur:
- An independent detector demonstrates stable, externally replicated performance on current models, hybrid writing, UVU disciplines, non-native English writing, and disability-relevant samples, with transparent thresholds and reviewable evidence.
- UVU’s baseline shows that existing assessments already provide strong individual assurance with low workload and high student trust.
- A UVU pilot shows that oral or practical checks add little validity while materially worsening access, course completion, or trust.
- Accreditors or state rules require a different form of secure assessment.
- Reliable provenance standards make authorship contributions independently verifiable without collecting private writing behavior.
- UVU’s adjudicated-case data show the main problem is policy confusion, fabricated sources, contract cheating, or another cause better addressed by a narrower intervention.
What the plan should change
- 1. Replace “three policy templates” with one assessment-assurance standard. Require an assignment-level Secure/Open label, allowed functions, prohibited functions, disclosure, and an individual proof point where the learning outcome demands it.
- 2. Make the detector rule binding. State that detector output is neither evidence nor a case metric and cannot by itself cause a grade change, referral, interview, or sanction.
- 3. Use the course-shell moment as the control point. Add the label and disclosure block to Canvas before courses open, with a report showing unlabeled high-stakes tasks.
- 4. Redefine the catalyst deliverable. Each paid catalyst redesigns and runs one signature assessment, measures workload and student trust, and publishes a de-identified design card.
- 5. Pilot all seven colleges, but use different designs. Engineering should test code defense; Business a changed-fact decision task; Humanities staged research plus close reading; Education microteaching; Science a lab defense; Health a synthetic practical; Arts a provenance-rich portfolio plus live critique.
- 6. Keep secure assessment sparse. Programs should identify a few assurance points where UVU must prove individual competence. Do not require a secure component in every assignment.
- 7. Add a fair follow-up protocol. Give the student the rule and evidence, permit drafts and assistive-tool explanations, use questions aligned to the original task, give written reasons, and preserve appeal.
- 8. Turn the current budget into a measured stage gate. Keep the est $20,000–$40,000 pilot, collect real hours and accommodation costs, and withhold the est $150,000–$250,000 expansion decision until results exist.
- 9. Change the success measures. Count substantiated cases, overturns, validity, workload, access, and trust. Never count detector flags or surveillance events.
- 10. Treat the “91%” claim as unverified. Either attach the actual survey and question or remove the number. The case for action is already strong without an unsupported statistic.
Source log
- 1. FACT—June 12, 2024: Lee et al., Cheating in the age of generative AI.
- 2. FACT—February 3, 2026: Chen et al., Cheating in the second year of generative AI chatbots.
- 3. FACT—page checked September 3, 2026: ICAI McCabe surveys.
- 4. FACT—March 27, 2026: Curtis et al., Australian five-yearly plagiarism surveys.
- 5. FACT—February 1, 2024: HEPI, Provide or punish?.
- 6. FACT—February 2025: HEPI Student Generative AI Survey 2025.
- 7. FACT—March 12, 2026: HEPI Student Generative AI Survey 2026.
- 8. FACT—2026: Science, national public-university GenAI survey.
- 9. FACT—April 9, 2024: Turnitin global detector totals.
- 10. FACT—data through April 25, 2024: University of Hull AI misconduct response.
- 11. FACT—2024 reporting year: Deakin Student Academic Integrity report.
- 12. FACT—2023 reporting year: UNSW Student Conduct and Complaints report.
- 13. FACT—2024 reporting year, issued 2025: UNSW Student Conduct and Complaints report.
- 14. FACT—2023/24 reporting year: University of Manchester annual report.
- 15. FACT—July 10, 2023: Liang et al., detector bias against non-native writers.
- 16. FACT—December 25, 2023: Weber-Wulff et al., testing 14 detectors.
- 17. FACT—May 20, 2024: Liu et al., The great detectives.
- 18. FACT—June 26, 2024: Scarfe et al., University of Reading field test.
- 19. FACT—August 7, 2025: Erol et al., Can we trust academic AI detective?.
- 20. FACT—February 2, 2026: Hadra et al., detector accuracy and reliability.
- 21. FACT—July 16, 2026 preprint: Chambers and Kelley, autistic-writing misclassification.
- 22. FACT—February 13, 2023: Turnitin launch claims.
- 23. FACT—guidance updated February 9, 2026: Turnitin review guidance.
- 24. FACT—August 16, 2023: Vanderbilt disables Turnitin AI detection.
- 25. FACT—April 13, 2023: Bournemouth declines activation.
- 26. FACT—September 4, 2025: Curtin disablement announcement.
- 27. FACT—mid-2025 decision: University of Queensland rules and detector disablement.
- 28. FACT—July 2025: OIA selected AI case summaries.
- 29. FACT—July 2025: OIA autistic-student case CS072501.
- 30. FACT—November 20, 2024: Harris v. Adams.
- 31. FACT—November 23, 2023: TEQSA Assessment reform for the age of AI.
- 32. FACT—November 2024: TEQSA emerging-practice toolkit.
- 33. FACT—May 20, 2019 publication: Melbourne standardized oral assessment.
- 34. FACT—2020–2021 pilot, published 2023/24: Dublin City University interactive oral cases.
- 35. FACT—May 14, 2025: Flinders longitudinal oral assessment.
- 36. FACT—July 26, 2021: UIC oral coding exams.
- 37. FACT—2026 forthcoming report: UC Berkeley CS1 oral exams.
- 38. FACT—April 3, 2026: UiT supervised-assessment natural experiment.
- 39. FACT—July 10, 2019: Computer versus handwritten in-class exams.
- 40. FACT—September 20, 2014: University of Sydney portfolio validity.
- 41. FACT—2016: University of Idaho and NJIT ePortfolio assessment.
- 42. FACT—June 10, 2026: Central Queensland process-evidence audit.
- 43. FACT—November 27, 2024: University of Sydney two-lane policy.
- 44. FACT—updated April 30, 2026: University of Bath two-lane approach.
- 45. FACT—current 2026 procedure: Victoria University secure-assessment rule.
- 46. FACT—2020: Ellis et al., authentic assessment and contract cheating.
- 47. FACT—2021: O’Mahony, multistage assessment workload.
- 48. FACT—2015: PeerStudio scalable rapid feedback.
- 49. FACT—2023: Li et al., AI-generated reflective writing.
- 50. FACT—September 2017: Pair-programming meta-analysis.
- 51. FACT—June 1, 2021: Pair-programming cluster-randomized trial.
- 52. FACT—December 12, 2025: Brunel laboratory OSCE.
- 53. FACT—2023: Karolinska laboratory OSPE.
- 54. FACT—October 16, 2024: British University Vietnam AIAS implementation.
- 55. FACT—2025: Reimagining the AI Assessment Scale.
- 56. FACT—guidance current in 2026: UCL three assessment categories.
- 57. FACT—Senate approval July 8, 2026: Leeds Academic Integrity and Assistance Policy.
- 58. FACT—checked September 3, 2026: UVU Writing Center AI and assignments guide.
- 59. FACT—2025 guide: UVU CHSS AI Faculty Guide.
- 60. FACT—checked September 3, 2026: UVU policy manual.
- 61. FACT—current September 2026: UVU Policy 541 Student Code.
- 62. FACT—effective March 28, 2024: UVU Policy 445 Information Security.
- 63. FACT—effective August 10, 2026: UVU Policy 447 Systems Security.
- 64. FACT—effective June 18, 2025: UVU Policy 452 Electronic Technology Accessibility.
- 65. FACT—2026: UVU library AI guide.
- 66. FACT—December 2, 2025: USHE statewide AI direction.
- 67. FACT—revised August 11, 2020: BYU Academic Honesty Policy.
- 68. FACT—2026–27 catalog: USU Academic Honesty and Integrity.
- 69. FACT—current 2026 guidance: Weber Teaching in the Age of AI.
- 70. FACT—cabinet review April 8, 2025: SLCC Student Code.
- 71. FACT—amended February 19, 2026: SUU Course Syllabus Policy.
- 72. FACT—approved February 11, 2026: Snow College AI Classroom Policy.
- 73. FACT—revised April 26, 2024: Utah Tech Policy 555.
- 74. FACT—2024–25 report: Michigan CRLT annual report.
- 75. FACT—terms current in 2026: Cornell AI Experimenters Faculty Fellows.
- 76. FACT—2024–25 pilot: Duke AI Jump Start Grants.
- 77. FACT—June 28, 2024 proposal: Missouri AI learning-environment report.
- 78. FACT—standards effective May 31, 2024: FAA Airman Certification Standards.
- 79. FACT—guidance current in 2026: UK apprenticeship end-point assessment.
- 80. FACT—August 11, 2020: NIST SP 800-207 Zero Trust Architecture.
NOT_RUN
- Output-file write: not runThe complete result is provided here.
- UVU contact, faculty interviews, student interviews, vendor contact, logins, forms, and external messages: not run.
- UVU course-level inventory, Canvas inspection, case-file review, and local student survey: not run.
- Independent current-version detector testing on UVU writing: not run.
- Accessibility testing of proposed oral, secure, coding, or practical tasks: not run.
- Legal opinion on disability, discrimination, FERPA, contract, or civil-liability exposure: not run.
- Causal evaluation of recent Sydney, Bath, UQ, or other two-lane policies: not run; published outcome evidence was not found.
- Campus expansion budget approval: not run; the present figures remain estimates.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/43-assessment-integrity.md in the research pack.
Does it improve learning? The evidence and the measurement
Appendix W in five lines
- QuestionDoes the service help students learn, not just finish work faster?
- AnswerA course-built, hint-first tutor can help, but general chat access is not proven and can hurt later test scores.
- Deciding numbers+48% assisted practicefact−17% on the next unaided examfact420 enrolled students for the pilotest
- What the plan doesAsk why the student is using the tool, make hint-first learning the class default, register the study plan before it starts, and expand only if later tests without the tool show no harm.
- Still unknownNo study proves a campus-wide learning gain; local course data are UNKNOWN, and the university research review and sample-size study are NOT_RUN.
The plan measured adoption and never asked the provost's question: do students learn more, less, or differently? This appendix gathers the randomized trials and meta-analyses from 2023–2026 with their effect sizes, explains when AI helps and when it hurts, settles what the evidence says about tutor modes that refuse to give answers, checks whether any campus deployment moved retention or completion, and lays out a pre-registered study UVU could run in year one — with a learning-outcomes contract that says what result would stop the rollout. Research cutoff September 3, 2026; 110-plus dated sources. fact trial results · est thresholds and designs · unknown where evidence is thin.
What this changes in the plan
1. Access is not an intervention. Generative AI can improve learning, but handing students a general chatbot does not reliably do so. The strongest positive trials used course-aligned material, structured practice, feedback, teacher support, and active student work; the strongest warning trial found unrestricted AI raised assisted practice scores and lowered the next unaided exam. Refusing premature final answers removed that harm but did not by itself improve unaided learning. Immediate performance, retained learning, and transfer are different outcomes and can move in opposite directions; evidence beyond a few weeks is thin. 2. No credible study shows a campus-wide deployment improved retention, completion, time to degree, or institution-wide DFW rates. "Verified adopters," conversations, and satisfaction stay adoption measures, never learning measures — separate dashboards. 3. The router becomes purpose-first: is the student here to learn, produce, check, or decide? Learning-first tutoring (attempts, progressive hints, self-explanation, retrieval, transfer) is the course-practice default; a clearly labeled productivity mode exists where producing the product is the stated outcome; attempt gates must be removable without shame or disability disclosure. Course-owned tutor packs carry goals, approved sources, common errors, worked examples, faculty rules, and independent test boundaries. 4. A pre-registered, section-randomized pilot with delayed rollout, masked grading, delayed unaided assessments plus new problems the tutor never rehearsed, intention-to-treat analysis, and published confidence intervals — funded as part of the service, with UVU's IRB and the telemetry charter. 5. The learning-outcomes contract binds expansion: proceed to a larger trial only with no credible harm on delayed unaided learning; expand after the first credible year only if the delayed-learning estimate is positive with its lower 95% bound above −0.10 SD, transfer is not negative, DFW is no more than 2 points worse, and no subgroup shows repeated harm near 0.20 SD; reshape if assisted scores rise while delayed or transfer scores do not, or students mainly request answers; stop a mode if its upper confidence bound sits below no effect or repeated evidence shows delayed harm. No adoption target overrides these rules. est
Evidence cut: fact — 2026-09-03 MDT.
Executive verdict
- Generative AI can improve learning, but access to a general chatbot does not reliably do so.
- The strongest positive trials used course-aligned material, structured practice, feedback, teacher support, and active student work.
- The strongest warning trial found that unrestricted AI improved assisted practice while reducing the next unaided exam score.
- Refusing premature final answers removed that measured harm, but did not by itself improve unaided learning.
- Immediate performance, retained learning, and transfer to new problems are different outcomes and can move in opposite directions.
- Evidence for retention lasting more than several weeks is thin; evidence for far transfer is thinner.
- No credible study located shows that a campus-wide generative-AI deployment improved retention, completion, time to degree, or institution-wide DFW rates.
- UVU should treat the service as an intervention to test, not a benefit already proved.
- “Verified adopters,” conversations, satisfaction, and generated work must remain adoption measures, not learning measures.
- Recommendation: launch a learning-first tutor in selected courses, preregister the evaluation, test later without AI, and bind expansion to a public learning-outcomes contract.
Evidence terms
- Assisted performance: Work completed while AI is available. It may show productivity rather than learning.
- Immediate unaided learning: A test taken after AI is removed, usually during the same session.
- Retention: An unaided test after a delay.
- Transfer: Applying the learning to a different problem or setting.
- DFW: A final grade of D or F, or withdrawal.
- fact: A number reported by a dated source.
- est: A planning estimate whose assumptions are shown.
- unknown: The required evidence was not found.
What the research says, by design quality
Strong randomized evidence
| Study | Design and population | Result | Outcome | What it does not prove |
|---|---|---|---|---|
| Harvard physics tutor | fact — 2025 Crossover randomized trial; fact 194 eligible undergraduates | fact +0.63 SD adjusted effect; ceiling-adjusted estimates fact +0.73 to +1.30 SD; confidence interval UNKNOWN/not reported | Immediate unaided learning | No delayed retention or transfer. The tutor used instructor-written solutions, sequencing, videos, and question-specific scaffolds; it was not general ChatGPT. |
| Turkey high-school mathematics | fact — 2025 Classroom-randomized field trial with nearly fact 1,000 students over fact four 90-minute sessions | During practice, ordinary GPT raised scores fact 48%; the guarded tutor raised them fact 127%. On the next unaided exam, ordinary GPT reduced scores fact 17%; the guarded tutor effect was fact −0.004, not significant. Confidence intervals UNKNOWN/not reported; clustered standard errors were reported. | Assisted performance and immediate unaided learning | No long-term retention or transfer. The guarded tutor prevented measured harm but did not beat control on unaided learning. |
| Nigeria World Bank study | fact — 2025 Student-randomized field trial in fact nine public schools; fact 1,328 randomized volunteers; fact 654 in the main assessment | Combined assessment fact +0.310 SD, standard error fact 0.068; English fact +0.238 SD; regular school exam fact +0.206 SD, standard error fact 0.067. Confidence intervals UNKNOWN/not reported. | Immediate unaided learning and near transfer | The intervention bundled fact twelve 90-minute after-school sessions, extra time, pairs, teacher support, structured prompts, and AI. Attrition was high and there was no time-matched active control. |
| Khan Academy/Khanmigo | fact — 2026 Two-year cluster trial in fact 18 Tennessee middle schools and fact 53 grade-within-school clusters | Pooled achievement fact +0.040 SD, standard error fact 0.019; first year fact +0.020 SD, standard error fact 0.035; second year fact +0.084 SD, standard error fact 0.041. Confidence intervals UNKNOWN/not reported. | Within-term and annual mathematics achievement | It did not isolate Khanmigo from the larger Khan Academy practice system. The authors found gains similar to non-AI Khan Academy practice. |
| Productive AI tutoring experiment | fact — 2026 Factorial randomized trial with more than fact 6,000 middle-school students | After errors, guarded AI increased next-attempt correctness fact 0.085, standard error fact 0.009. At fact one week, practiced-item effect was fact +0.032, standard error fact 0.017; unpracticed-item effect was fact +0.002, standard error fact 0.017. | Assisted remediation, short retention, transfer | The delayed test contained only fact four items. The best signal was small, marginal, and restricted to practiced content. |
| Tutor CoPilot | fact — 2025 version Tutor-level randomized trial; fact 783 tutors and fact 4,136 analyzed sessions | Exit-ticket pass rate fact +4.0 percentage points, standard error fact 1.5 points, over a fact 61.7% control mean | Immediate session mastery | AI advised human tutors; it was not an autonomous student tutor. No measured annual retention or transfer. |
| Middlebury proctored experiment | fact — 2026 preprint fact 211 undergraduates; fact 204 returned about one week later | Immediate unaided effect fact +0.266 SD, standard error fact 0.125; delayed effect fact +0.268 SD, standard error fact 0.120 | Immediate learning and one-week retention | Selective institution, unfamiliar topics, fixed fact 35-minute task, and preprint status. Explanatory use looked better than answer generation, but use style was not randomized. |
| LearnLM/Eedi | fact — 2025 vendor technical report Two-level randomized trial; fact 165 students in fact five UK schools | Next-topic success: LearnLM fact 66.2%, fact 95% credible interval [61.1%, 71.2%]; human tutors fact 60.7%, interval [55.8%, 65.4%]. Difference fact +5.5 points, interval [−1.4, 12.4]. | Near transfer | The AI-versus-human interval included no difference, every AI draft was reviewed by a human, and the report was vendor-authored. |
| AI-generated mathematics hints | fact — 2024 Randomized fact 3-by-4 experiment; fact 274 adults | Unaided gains: no hints fact 1.85%; human hints fact 11.62%; ChatGPT hints fact 17.00%. GPT versus no hints fact p=0.011; GPT versus human fact p=0.416. Confidence intervals UNKNOWN/not reported. | Immediate unaided learning | A nonsignificant GPT-human difference does not establish formal equivalence. fact 32% of raw AI help initially failed quality checks and required filtering. |
| GPT homework tutor | fact — 2024 preprint Stratified randomized trial; fact 76 Italian high-school students over fact eight weeks | Overall fact d=0.251, p=0.314; grammar-focused classes fact d=0.603, p=0.087; essay-focused classes fact d=−0.004, p=0.991 | Course learning | Small and underpowered. The tutor sometimes revealed answers despite instructions not to do so. |
| Delayed writing/reference tasks | fact — 2025 preprint Two undergraduate randomized experiments with unaided follow-up after fact 14 days | Essay task: fact d=0.053, fact 95% CI [−1.6, 1.1] in raw-score units. Reference task: AI fact 5.98 versus control fact 7.54; AI effect fact d=−1.301 | Retention | Small, single-university experiments with very different tasks and confusing contrast signs in the paper. |
Randomized evidence with important selection or measurement limits
The Stanford Code in Place experiment randomized fact 5,831 active learners from fact 146 countries to receive access and advertising for a guarded coding assistant. Only fact 14.2% used it. Access reduced optional-exam participation by fact 4.3 percentage points, fact 95% CI [−6.9, −1.7], and homework completion by fact 4.6 points, fact 95% CI [−7.2, −1.9]. The intent-to-treat exam effect was fact +0.67 points, fact 90% CI [−0.37, 1.72], or no reliable difference. An estimated fact +6.86-point effect for adopters did not survive multiple-testing adjustment and depends on an untestable assumption. This is powerful evidence that offering a tutor can reduce engagement even when some users benefit. Nie et al., revised 2025
Two preregistered coding experiments found no significant overall unaided learning gain from ChatGPT. Allowing copying increased solution requests from fact 42.2% to fact 64.2% and reduced explanation requests from fact 31.3% to fact 18.4%. Students covered more material but did not learn more overall. Lehmann, Cornelius, and Sting, revised 2025
A fact — 2025 randomized retention study assigned fact 120 undergraduates to unrestricted ChatGPT or traditional study. At a surprise test fact 45 days later, usable observations implied by the reported test statistic were fact 85; ChatGPT students scored fact 57.5% versus fact 68.5%, fact t(83)=−3.19, p=0.002, d=0.68. Attrition and the limited public methodological record require caution, but it is one of the few genuinely delayed tests. Barcaui, 2025
Writing evidence
A randomized writing study with fact 117 university students compared ChatGPT, a human expert, writing analytics, and no extra support. ChatGPT improved the essay product, but knowledge gain and transfer did not differ. The authors used “metacognitive laziness” to describe altered self-regulation and dependence. It is a proposed mechanism, not a diagnosis or proof of permanent decline. Fan et al., 2025
A fact — Spring 2025 Carnegie Mellon quasi-experiment covered fact 424 students in fact 28 first-year writing sections. Multiple rounds of constrained AI feedback outperformed peer feedback on draft improvement, fact t=2.74, p=0.01, but draft revision is neither delayed retention nor far transfer. The AI was prohibited from generating replacement prose. CMU GAITAR@Scale
Seven experiments comparing LLM summaries or chat with web search found faster, shallower information gathering and less original downstream advice in several conditions. In the first experiment, LLM users searched for fact 585.41 seconds versus fact 742.81 seconds; reported learning depth was fact 3.43 versus fact 3.86, both fact p<0.001. These were immediate depth-of-processing tasks, not course retention tests. Melumad and Yun, 2025
Coding evidence
A fact — 2026 preprint meta-analysis covering fact 23 studies and fact 27 effects found assisted productivity fact g=0.33, 95% CI [0.09, 0.58], but learning fact g=0.14, 95% CI [−0.18, 0.47]. The learning result was not significant and mixed experimental with nonrandomized evidence. Maier et al., 2026
A Chinese university quasi-experiment with fact 82 programming students found AI-group performance fact 84.11 versus fact 78.36, but fact t(77)=1.28, p=0.204. Students frequently copied AI code, ran it, and returned errors to the model. Sun et al., 2024
A controlled introductory-programming experiment with fact 56 students found no significant learning improvement and observed reduced use of other learning resources after students began using ChatGPT. Chen et al., 2024
Older intelligent-tutoring baseline
These are structured, domain-specific systems. A general language model does not automatically inherit their effects.
- fact — 2011 Human tutoring effect fact d=0.79; step-based intelligent tutoring fact d=0.76; answer-based systems fact d=0.31. Confidence intervals UNKNOWN/not reported. VanLehn
- fact — 2014 Across fact 107 effects and fact 14,321 learners, intelligent tutors produced fact g=0.42 versus large-group teaching, fact g=0.57 versus other computer instruction, and fact g=0.35 versus textbooks or workbooks. They did not significantly beat individual human tutoring. Confidence intervals were not reported in the abstract. Ma et al.
- fact — 2014 College intelligent-tutoring studies produced random-effects fact g=0.35, 95% CI [0.24, 0.46]. Steenbergen-Hu and Cooper
- fact — 2016 Across fact 50 controlled evaluations, the median effect was fact 0.66 SD. Locally aligned tests averaged fact 0.73 SD; standardized tests averaged only fact 0.13 SD. This is a warning against testing a tutor with questions too close to its own material. Kulik and Fletcher
Evidence synthesis
The sign is not “AI versus no AI.” It is:
tool design × student behavior × course design × outcome timing.
The best-supported conclusions are:
- Ordinary answer generation is excellent evidence of immediate productivity, not evidence of learning.
- Guardrails prevent the clearest measured harm, but guardrails alone do not guarantee learning.
- Positive trials usually combine AI with additional instructional ingredients.
- Effects tend to shrink when assessments are delayed, unaided, standardized, or meaningfully different.
- The causal evidence remains concentrated in mathematics, introductory science, coding, and short writing tasks.
- Semester-long, multisubject, broad-access university evidence remains unknown.
Mechanisms
When AI helps
Attempt before assistance. Formative-feedback research recommends making the learner attempt the task before seeing the answer. A meta-analysis of problem-solving before instruction found fact g=0.36 for conceptual knowledge and transfer, but fact g=−0.03 for procedures. Novices still need an escape route. Sinha and Kapur, 2021
Progressive hints. The Turkey experiment directly shows that teacher-authored hints and answer resistance can change student behavior and prevent the harm observed with an unrestricted assistant.
Self-explanation. Asking the learner why a step works can support near and far transfer. A randomized worked-example study reported near-transfer effects up to fact f=0.42 and far-transfer effects up to fact f=0.37 for prompting, although the study predates generative AI. Atkinson, Renkl, and Merrill, 2003
Fading. Support should move from a worked model, to a missing final step, to more missing steps, and then to independent work.
Retrieval practice. Repeated study can look better after fact five minutes, while active retrieval wins after fact two days and fact one week. The tutor should therefore ask students to reconstruct an answer later without assistance. Roediger and Karpicke, 2006
Specific feedback. Feedback should address the task, explain what and why, arrive in manageable pieces, and follow an attempt. Too much information can become another answer dump. Shute, 2008
Course alignment plus independent testing. Course content helps the tutor avoid irrelevant answers. Independent or standardized assessments prevent that same alignment from inflating the measured effect.
Human support. Nigeria and Tutor CoPilot suggest that AI can amplify teacher or tutor attention. They do not show that the human layer can safely be removed.
When AI hurts
Answer-giving. Students can copy a correct solution without constructing the knowledge needed to reproduce it.
Cognitive offloading. The learner delegates the same mental operation the course is trying to teach.
Fluency mistaken for mastery. A fast, polished result feels like learning even when later unaided performance is unchanged.
Reduced help-seeking quality. Students often send bare answers, click suggested prompts, or request a solution instead of explaining their reasoning.
Lower engagement. The Code in Place trial found that simply advertising the tool reduced participation in other course activities.
Automation bias. Correct AI can help, but wrong AI can pull a previously correct human decision toward error. Explanations are not a reliable shield.
Mode confusion. A student may believe a productivity assistant is operating as a tutor, or that a tutor’s answer proves mastery.
The MIT EEG essay study and its limits
fact — 2025 preprint The MIT Media Lab study assigned fact 54 people, fact 18 per group, to LLM, search, or unaided essay writing for the first fact three sessions. Only fact 18 returned for the optional crossover. The calendar spanned fact four months; this was not continuous controlled exposure for four months.
In the first session, fact 15 of 18 LLM users said they could not quote their essay, versus fact 2 of 18 in each comparison group. None of the fact 18 LLM users produced a correct quotation. By the third session, the gap had narrowed.
The EEG analysis reported different connectivity patterns and less extensive directed connectivity in several LLM comparisons. It did not measure IQ, permanent brain change, semester learning, or far transfer. Connection counts are not standardized learning effects. Some detailed EEG comparisons also ran in the opposite direction.
A fact — 2025-12-29 methodological critique identified limited power, unclear multiple-comparison handling, missing standardized effects, reporting inconsistencies, and analyses sometimes based on only fact 2–4 essays per group.
Verdict: useful hypothesis-generating evidence about attention, ownership, and immediate source memory during AI-assisted composition. It is not proof that ChatGPT damages the brain or causes lasting intellectual decline. MIT study, independent critique
Equity: who benefits and who may fall behind
Evidence is contradictory.
- Nigeria’s exploratory analysis found larger gains for girls, but the intervention was a supported after-school program and attrition was high.
- Tutor CoPilot raised exit-ticket passing by fact 9 percentage points, from fact 56% to fact 65%, for students served by lower-rated tutors. This suggests AI may raise the floor when it improves human tutoring practice.
- In Code in Place, access increased exam participation by fact 14.8 percentage points among students from lower-HDI countries while reducing engagement overall.
- General productivity experiments often find larger gains for lower initial performers. That does not establish larger retained-learning gains.
- The Turkey result warns that students who most need instruction can also be most exposed to harm when the system makes correct-looking answers easy to copy.
- Effects for disabled students, multilingual UVU students, part-time students, working students, and different racial or ethnic groups remain unknown because most direct trials were too small or did not report suitable subgroup tests.
UVU should not assume that free access closes a learning gap. It should test access, substantive use, delayed learning, and harms separately.
The tutoring-mode question
What is actually proved?
The Turkey trial supplies the cleanest direct comparison using the same underlying model:
| Mode | Assisted work | Next unaided exam |
|---|---|---|
| General assistant | fact +48% | fact −17% |
| Guarded tutor | fact +127% | fact approximately no difference |
| No AI | Baseline | Baseline |
The guarded tutor:
- used teacher-written solutions and common errors;
- delivered hints;
- resisted direct final answers;
- checked student work;
- encouraged attempts.
Students using the general assistant often asked for or copied complete solutions. Guarded-tutor students attempted answers and asked for help more often.
This proves that interaction design can flip a result from measured harm to no measured harm. It does not prove that refusal itself causes positive learning.
Why “never give the answer” is insufficient
The fact — 2026 Khanmigo trial is the warning. Although fact 96% tried the tutor, the median student messaged it on only about one-third of practice days and in only fact 17% of sessions containing a mistake. Only fact 14.5% of messages contained a real mathematics question or reasoning step.
A tutor can be pedagogically careful and still be ignored.
The service therefore needs graduated help:
- first ask for the student’s attempt or plan;
- diagnose the blocking point;
- give one useful hint;
- ask the student to act on it;
- if still blocked, show one worked step;
- if a complete worked example becomes necessary, follow it with a fresh, comparable problem completed without help.
An accommodation or urgent-completion path must be able to bypass effort gates without forcing a student to disclose a disability to the chatbot.
Implications for the UVU router and templates
The present plan’s router is designed mainly around workload, speed, context size, and cloud escalation. Add a learning-policy layer before model selection:
Student intent
├── Learn or practice
│ └── Learning-first tutor
│ ├── attempt
│ ├── diagnose
│ ├── progressive hint
│ ├── self-explanation
│ ├── fresh problem
│ └── delayed unaided check
├── Produce or execute
│ └── General assistant, with course-use and disclosure rules
├── Check or critique
│ └── Independent answer first, then AI comparison and source verification
└── High-stakes or accommodation
└── Human-approved route with an accessible completion option
Required router behavior:
- Ask whether the user is trying to learn, finish, check, or make a high-stakes decision.
- Make learning-first tutor mode the default inside course practice.
- Keep productivity mode visibly separate; never label its outputs as evidence of learning.
- Cache the course’s learning goals, approved references, common errors, and faculty rules with the existing course prefix.
- Do not let a faster model or cheaper machine silently weaken the tutor policy.
- Log the mode, prompt-policy version, hint stage, and completion state without retaining student content beyond the approved study rules.
- Detect repeated requests for a final answer and move to diagnosis or a worked-example-plus-transfer pattern.
- Build no-tool retrieval prompts into the workflow.
- Preserve a human escalation path.
- Test the tutor against adversarial answer-seeking prompts before course use.
- Measure actual substantive dialogue; activation alone is not tutoring.
Institutional outcomes
What campus deployments have established
unknown: No located campus-wide generative-AI rollout has published credible causal evidence of improved institution-wide retention, completion, time to degree, or DFW rates.
California State University, Arizona State University, Oxford, Virginia Tech, and other named deployments have reported access, activation, projects, training, use cases, or opinions. Those are implementation signals. They are not learning or completion effects.
A fact — 2026 University of Michigan preprint used administrative data covering fact 137,807 unique students and fact 46,485 offerings in its balanced sample. More AI-susceptible courses did not show significant post-ChatGPT differences in grades, withdrawals, or failures. The conservative estimates were:
- grade fact +0.045 points, standard error fact 0.026, below conventional significance;
- withdrawal fact +0.010, standard error fact 0.006, not significant;
- failure fact −0.002, standard error fact 0.002, not significant.
Pre-trends failed for grades, so even the null should not be read as a clean causal effect. The study measured AI availability and course susceptibility, not verified student use or learning. Dumlao et al., 2026
Useful older, non-generative comparisons
Georgia State’s Pounce system was an administrative chatbot, not a tutor. In a randomized trial with fact 7,489 admitted students, outreach increased timely Georgia State enrollment by fact 3.3 percentage points, standard error fact 1.6 points, among the fact 1,948 students already committed to attend. It did not test subject learning or degree completion. Page and Gehlbach, 2017/2018
Georgia State later randomized proactive, course-specific, non-generative chatbot messages in government and economics. Across courses, the system increased the probability of an A or B by about fact 4 percentage points. In economics, women were fact 10 points less likely to DFW and earned grades fact 7 points higher. These effects came from reminders, course information, and support—not answer generation. Page et al., 2024/2026
Georgia State’s generative TEACH ME study began randomized trials in fact 2024, plans to cover more than fact 20,000 students, and runs through fact 2027. Final results remain unknown. GSU, 2024-10-08
UVU’s public baseline
UVU’s current public dashboard reports:
- fact — Fall 2025 48,669 enrolled students;
- fact — Fall 2025 40% first-generation students;
- fact — 2017/18 cohort measured in 2025 48% completion;
- fact — Fall 2025 72% annual retention.
The Student Right-to-Know report dated fact — 2026-06-12 reports:
- fact 47% overall graduation for the fact 2019 cohort;
- fact 72% first-time bachelor’s retention for full-time students in the fact 2024 cohort;
- fact 50% first-time bachelor’s retention for part-time students in the fact 2024 cohort.
These measures use different cohorts and definitions and should not be blended.
Current public UVU gateway-course DFW rates were UNKNOWN/not found. UVU should obtain course-section baselines from Institutional Research before power calculations or target-setting.
What success can reasonably mean
For the pilot, success should mean better delayed, unaided learning in a named course—not a detectable change in university retention.
For the first year:
- Primary: delayed, unaided course learning.
- Secondary: near and far transfer, common exams, course pass, DFW, withdrawal, next-course performance, and time to mastery.
- Exploratory: next-term enrollment and annual retention.
- Longer-term: completion and time to degree only after several cohorts.
Retention and completion are affected by advising, finances, employment, health, scheduling, course availability, and many other causes. A small tutoring pilot cannot identify its contribution to those outcomes.
How UVU should measure it
Recommended design
Use a preregistered, section-randomized delayed-rollout trial in gateway courses.
Pilot design
- Select est two gateway courses with common outcomes and at least est seven sections per condition.
- Randomize sections within course, instructor where possible, modality, and meeting time.
- Treatment sections receive the learning-first tutor.
- Control sections receive normal course resources and the same human-help routes, then receive the tutor after the evaluation period.
- Do not attempt to forbid consumer AI. Record self-reported outside use and treat it as contamination.
- Separate consent to use the service from consent for identifiable research data.
- Analyze everyone by assigned condition, whether they use the tutor or not.
- Run an additional usage-based estimate only as secondary analysis.
- If section randomization is impossible, use a randomized staggered invitation.
- If neither is feasible, use opt-in participation with matching on prior grades, course load, attendance, demographics, and baseline knowledge. Label the result quasi-experimental because matching cannot remove unmeasured motivation.
A three-way comparison with an unrestricted assistant would answer the tutoring-mode question directly, but the Turkey harm result weakens ethical equipoise. UVU should include that arm only for ungraded, low-risk practice if faculty and the IRB conclude that both modes are acceptable. A safer alternative is to randomize progressive-hint designs after students have made an initial attempt.
Outcomes and instruments
Primary outcome
est design choice A common, unaided assessment administered est four weeks after the relevant unit. Use external or faculty-written items not shown to the tutor. Score against a preregistered rubric by graders masked to assignment.
Secondary outcomes
- Immediate unaided test within est 48 hours.
- Near-transfer problems using different surface details.
- Far-transfer problem requiring the same principle in a different setting.
- Common course exam.
- DFW and withdrawal.
- Next-course grade or concept test.
- Time to correct independent solution.
- Error type and misconception persistence.
- Student confidence calibration: confidence minus actual correctness.
- Substantive tutor behavior: attempts, explanation requests, hints used, answer requests, verification, and completion of retrieval checks.
- Accessibility, privacy, accuracy, and academic-integrity incidents.
- Student and faculty time.
- Cash and staff cost per learner achieving the preregistered mastery threshold.
Satisfaction and self-reported learning remain descriptive. They must not replace the learning test.
Sample sizes and detectable effects
These are planning estimates pending UVU’s real section sizes, baseline variance, attrition, and within-section correlation.
Assumptions:
- est 30 analyzed students per section;
- est 0.05 within-section correlation;
- est 15% missing follow-up;
- est 80% power;
- est two-sided 5% error rate;
- equal treatment and control allocation.
Under those assumptions:
| Stage | Proposed sample | Approximate detectable standardized effect |
|---|---|---|
| Pilot | est 14 sections / 420 enrolled students | est about 0.33 SD |
| First year | est 40 sections / 1,200 enrolled students | est about 0.19 SD |
Basis: the standard two-group approximation gives total individual-equivalent sample est 15.68/d²; the cluster design effect is est 1 + (30−1)×0.05 = 2.45; projected retention is est 85%.
The pilot is therefore a feasibility, safety, behavior, and large-effect screen. It cannot establish that a small effect is absent.
A DFW calculation is unknown until UVU supplies course-specific baselines. Detecting a small percentage-point change could require several thousand students.
Before preregistration, a statistician should replace these approximations with simulation using UVU’s actual section counts, unequal class sizes, baseline scores, and expected missingness.
Analysis
Preregister:
- one primary outcome and time point;
- treatment assignment and exclusions;
- section-level clustering;
- course and randomization-block effects;
- baseline-score adjustment;
- missing-data and attrition sensitivity analyses;
- handling of outside AI use;
- correction for multiple secondary tests;
- subgroup definitions;
- stopping rules;
- prompt, router, model, and course-content version;
- treatment-change rules if the model fails.
Report intention-to-treat first. Report point estimates and confidence intervals even when results are null.
Predeclare equity analyses for first-generation status, Pell eligibility where authorized, prior preparation, part-time/full-time status, modality, gender, race or ethnicity, multilingual status, and disability accommodation. Suppress unsafe small cells and describe underpowered subgroup findings as exploratory.
IRB and privacy path
UVU states that human-subjects research, including pilot and feasibility studies intended for public dissemination, must be submitted to its IRB. Investigators cannot determine exemption for themselves. The application must describe randomization, controls, recruitment, risks, study materials, and analysis. CITI training is required before the project begins. UVU IRB FAQ, application process
Educational practice research may qualify for an exempt determination if it is minimal risk and does not harm students’ opportunity to learn. The IRB makes that determination. UVU exempt categories
The sequence should be:
- appoint a UVU faculty principal investigator;
- include Institutional Research, teaching-center staff, accessibility, privacy, students, and participating faculty;
- finish CITI training;
- freeze the protocol, assessments, consent language, data map, and deletion schedule;
- obtain the IRB determination before recruitment or research data collection;
- execute any required FERPA studies agreement;
- preregister before viewing outcomes;
- begin the pilot only after those gates pass.
FERPA’s studies exception can permit limited education-record disclosure for research conducted for the institution, but it requires a written agreement specifying purpose, use, security, and destruction. Use deidentified data whenever possible. U.S. Department of Education
Cost of a credible evaluation
All figures below are estimates of full economic cost, including existing staff time.
Pilot: est $75,000–$110,000.
Illustrative midpoint basis:
- evaluation lead: est 0.20 FTE × $140,000 loaded annual cost = $28,000;
- analyst/data support: est 0.15 FTE × $120,000 = $18,000;
- faculty participation: est 14 sections × $1,000 = $14,000;
- retained-assessment completion support: est 300 students × $20 = $6,000;
- assessment, accessibility, data, and privacy work: est $17,000;
- contingency: est about $12,000;
- total illustrative midpoint: est about $95,000.
First credible year: est $210,000–$280,000.
Illustrative midpoint basis:
- evaluation lead: est 0.50 FTE × $140,000 = $70,000;
- analyst/data engineering: est 0.40 FTE × $120,000 = $48,000;
- faculty participation: est 40 sections × $1,000 = $40,000;
- delayed-assessment completion: est 1,000 students × $20 = $20,000;
- accessibility, data governance, independent assessment review, and publication: est $50,000;
- contingency: est about $34,000;
- total illustrative midpoint: est about $262,000.
Actual UVU salary, workload, incentive, and assessment costs are unknown.
What UVU should publish
Publish:
- the dated protocol and registration;
- intervention and control descriptions;
- model and prompt-policy versions;
- assessment instruments where test security permits;
- section and participant flow;
- assignment balance and attrition;
- all primary and secondary results with confidence intervals;
- null and harmful results;
- subgroup estimates with privacy protections;
- incident counts and model changes;
- cost per assigned student, active user, and learner reaching mastery;
- analysis code, data dictionary, and the most deidentified data the IRB permits;
- deviations from the protocol.
Do not publish identifiable chats or create a misconduct dataset from research logs.
Learning-outcomes contract
These are proposed governance thresholds, not established scientific constants.
Proceed from pilot to a larger trial only if:
- no credible harm appears on delayed unaided learning;
- assessments and data collection work as designed;
- treatment separation is real;
- accessibility and privacy checks pass;
- substantive learning-mode use is high enough to test the intervention;
- faculty and students can use the system without coercion.
Expand after the first credible year only if:
- est policy threshold the delayed-learning point estimate is positive and the lower est 95% confidence bound is above est −0.10 SD;
- transfer is not negative;
- est policy threshold DFW is not more than est 2 percentage points worse;
- no important subgroup shows a repeated adverse signal around est 0.20 SD or larger;
- serious privacy, accessibility, or answer-leak incidents are resolved;
- costs are acceptable relative to observed learning, not usage.
Reshape the intervention if:
- assisted performance rises but delayed or transfer performance does not;
- students mainly request answers or paste outputs;
- the tutor is rarely used after mistakes;
- tutor mode reduces engagement with class, instructors, or human tutoring;
- benefits occur only under intensive human support that the scale plan cannot sustain.
Stop the affected mode if:
- the upper confidence bound for the primary learning effect is below no effect;
- repeated evidence shows meaningful delayed harm;
- the system bypasses faculty assessment rules;
- serious privacy or accessibility harm cannot be fixed promptly.
No adoption target overrides these rules.
Cross-domain learnings
Calculators
A fact — 1986 meta-analysis of fact 79 reports found that calculators used with normal instruction generally maintained or improved paper-and-pencil computation and problem solving, with a grade-four exception.
A fact — 2003 synthesis of fact 54 studies found strong gains when calculators were part of both instruction and tests: operational fact g=0.38, computational fact g=0.43, conceptual fact g=0.44, and problem-solving fact g=0.33. When calculators were withheld on tests, most effects were near zero; operational skill was fact g=0.17. Only fact three studies assessed retention after fact 2–12 weeks.
Learning: do not ban a useful tool. Teach with it, preserve no-tool practice for skills students must own, and measure both tool-enabled and independent work.
Spell-checkers
A fact — 2017 randomized study with fact 88 university second-language learners found that spell-check choice and dictionary use supported correction after aid removal and a fact one-day delay. A red underline alone did not.
A fact — 2006 study with fact 65 university learners found better surface correction without less content revision, but did not test later unaided spelling.
Learning: correction can free attention for ideas, but passive flags do not teach much. Ask the learner to select, explain, and later reproduce the correction.
GPS and spatial memory
A fact — 2008 experiment found that turn-by-turn GPS users learned routes less well than direct-experience or paper-map users.
A fact — 2020 study of fact 50 regular drivers linked greater GPS use with poorer self-guided spatial memory. Only fact 13 returned about fact three years later, so its longitudinal causal claim is weak.
Landmark-rich guidance and auditory beacons that leave route choice to the traveler performed better on later navigation measures than ordinary turn-by-turn directions in smaller experiments.
Learning: preserve decisions and structural cues. A tutor should teach the map of the subject, not issue isolated turns.
Clinical decision support
A fact — 2023 experiment with fact 457 clinicians across fact 13 states found that standard AI advice raised diagnostic accuracy fact 2.9 percentage points; biased advice lowered it fact 11.3 points. Adding an explanation did not reliably protect clinicians from biased advice.
In another study, fact 41.73% of non-radiologists and fact 27.54% of radiologists accepted both incorrect recommendations presented to them.
A fact — 2025 observational colonoscopy study found unaided adenoma detection fell from fact 28.4% before AI exposure to fact 22.4% afterward. A larger fact — 2026 prospective study covering fact 5,013 colonoscopies did not find significant post-removal decline, contradicting a simple de-skilling conclusion.
Learning: get an independent judgment before showing AI advice. Explanations are not enough; require verification and test performance after assistance is removed.
Aviation automation
A fact — 2014 simulator study with fact 16 active pilots found that basic control and instrument-scan skills held up better than manual navigation and abnormal-event diagnosis. Only fact 1 of 16 completed every manual-navigation phase without a listed error.
The fact — 2013 FAA automation review found substantial safety and workload benefits alongside mode confusion, weak monitoring, and erosion of some manual and cognitive skills.
Learning: keep modes visible, train graceful failure, and schedule periodic no-tool drills for essential skills.
Shared lesson
Other fields stopped asking whether assistance is “good” or “bad.” They ask:
- Which skill must remain human?
- Which part may be safely offloaded?
- Does the tool preserve active decisions?
- What happens when the tool is wrong or absent?
- Is independent skill practiced and checked?
UVU should use the same frame.
Devil’s advocate
The strongest case against the recommendation is that an answer-refusing campus tutor may solve the wrong problem.
Students already have unrestricted consumer assistants. A slower institutional tutor could frustrate them and push use into less governed tools. Modern work increasingly rewards delegation, critique, synthesis, and verification rather than unaided production. Protecting every old no-tool skill may resemble requiring manual arithmetic after calculators became normal.
The strongest positive studies are narrow, bundled, and often conducted with younger learners. The strongest harm study lasted only a few sessions. Models and interfaces change faster than normal educational research. A costly trial could measure an obsolete implementation by publication time. Faculty may get more value from redesigning assessments and teaching AI verification than from restricting answers.
That case is serious. It supports a dual-mode service, not an unrestricted default:
- tutoring mode for learning goals;
- productivity mode where tool-enabled work is the course goal;
- faculty control over which mode applies;
- no-tool checks only for skills the course says students must retain;
- rapid versioned experiments instead of one frozen, multiyear product.
What would change the recommendation
The recommendation would move toward a general-assistant default if independent, multi-course university trials showed all of the following:
- ordinary assistants outperform guarded tutors on delayed unaided learning;
- the advantage persists across semesters and transfers to new problems;
- unrestricted assistance does not widen gaps by prior preparation;
- answer refusal materially reduces engagement without a compensating learning benefit;
- transparent productivity-mode teaching produces equal or better independent verification skill.
The recommendation would become more restrictive if:
- UVU or replicated trials find delayed harm around est policy threshold 0.20 SD or more;
- students systematically bypass tutor controls and perform worse later;
- course DFW or withdrawal increases;
- harmful effects concentrate among students with weaker prior preparation;
- faculty cannot keep assessments independent of tutor training material.
The recommendation would pause entirely if UVU cannot obtain an IRB determination, a lawful data path, accessible alternatives, common assessments, or enough sections for a credible comparison.
What the plan should change
- 1. Add the learning-outcomes contract before buying or scaling. Adoption may unlock evaluation, but it cannot authorize expansion by itself.
- 2. Change the router from workload-first to purpose-first. Ask whether the student intends to learn, produce, check, or make a high-stakes decision before selecting a model or machine.
- 3. Make learning-first tutoring the course-practice default. Require attempts, progressive hints, self-explanation, retrieval, and transfer.
- 4. Keep a clearly labeled productivity mode. Use it when generating the product is allowed and tool-enabled performance is the stated outcome.
- 5. Add delayed, unaided measurement. Every pilot course needs a common assessment outside the tutor, plus new problems the tutor did not rehearse.
- 6. Run a preregistered section-randomized pilot. Use delayed rollout, intention-to-treat analysis, masked grading, and published confidence intervals.
- 7. Measure student behavior without treating surveillance as learning science. Record content-minimized events such as attempt, hint stage, answer request, verification, and retrieval completion under the telemetry charter.
- 8. Add an accessibility escape. Attempt gates must be removable without shame, penalty, or unnecessary disability disclosure.
- 9. Require course-owned tutor packs. Each pack should contain learning goals, approved sources, common errors, worked examples, faculty rules, and independent test boundaries.
- 10. Separate dashboards. Report adoption, service quality, assisted productivity, immediate learning, delayed retention, transfer, DFW, equity, safety, and cost as different measures.
- 11. Do not call campus access an intervention. The evaluated intervention is the full combination of tutor policy, course integration, faculty support, student behavior, and assessment.
- 12. Fund evaluation as part of the service. A rollout without credible outcome measurement leaves the provost’s central question unanswered.
Source log
- 1. fact — 2025-06-03 Kestin et al., AI tutoring outperforms in-class active learning.
- 2. fact — 2025-06-25 Bastani et al., Generative AI without guardrails can harm learning.
- 3. fact — 2025-05-19 De Simone et al., From Chalkboards to Chatbots.
- 4. fact — 2026-08 Oreopoulos and Low, One Click Away: AI Tutoring with Khanmigo.
- 5. fact — 2026-08-17 Oreopoulos et al., Making AI Tutoring Productive.
- 6. fact — 2025-01 version Wang et al., Tutor CoPilot.
- 7. fact — 2026-07-09 Contractor and Reyes, Experimental Evidence on the Learning Impact of Generative AI.
- 8. fact — revised 2025-07-15 Nie et al., The GPT Surprise.
- 9. fact — revised 2025-03-08 Lehmann, Cornelius, and Sting, AI Meets the Classroom.
- 10. fact — 2025-11 LearnLM Team and Eedi, AI tutoring can safely and effectively support students.
- 11. fact — 2024-05-24 Pardos and Bhandari, ChatGPT-generated help and mathematics learning.
- 12. fact — 2024-09-24 Vanzo, Chowdhury, and Sachan, GPT-4 as a Homework Tutor.
- 13. fact — Spring 2025 Carnegie Mellon, GAITAR@Scale FeedbackPartner evaluation.
- 14. fact — 2025-04-01 Yano et al., Does the Use of Generative AI Undermine Learning?.
- 15. fact — 2025-10-28 Melumad and Yun, LLMs versus web search on depth of learning.
- 16. fact — 2025 Barcaui, ChatGPT as a cognitive crutch.
- 17. fact — 2025-03 issue Fan et al., Beware of Metacognitive Laziness.
- 18. fact — 2025-06-10 Kosmyna et al., Your Brain on ChatGPT.
- 19. fact — 2025-12-29 Stanković et al., Comment on Your Brain on ChatGPT.
- 20. fact — 2026-05-06 Maier et al., Meta-analysis of generative AI in programming.
- 21. fact — 2024 Sun et al., ChatGPT-facilitated programming.
- 22. fact — 2024-04 Chen et al., Does ChatGPT Help With Introductory Programming?.
- 23. fact — 2024 Kosar et al., Computer Science Education in the ChatGPT Era.
- 24. fact — 2023-07-13 Noy and Zhang, Productivity effects of generative AI.
- 25. fact — 2011-10-17 VanLehn, Human tutoring and intelligent tutoring systems.
- 26. fact — 2014 Ma et al., Intelligent tutoring systems and learning outcomes.
- 27. fact — 2014 Steenbergen-Hu and Cooper, Intelligent tutoring systems for college learning.
- 28. fact — 2016-03 Kulik and Fletcher, Effectiveness of Intelligent Tutoring Systems.
- 29. fact — 2008-03-01 Shute, Focus on Formative Feedback.
- 30. fact — 2006-03 Roediger and Karpicke, Test-enhanced learning.
- 31. fact — 2010-09 Butler, Repeated testing and transfer.
- 32. fact — 2003 Atkinson, Renkl, and Merrill, Fading worked steps and self-explanation.
- 33. fact — 2021 Sinha and Kapur, Problem-solving before instruction meta-analysis.
- 34. fact — 1986-03 Hembree and Dessart, Effects of calculators in precollege mathematics.
- 35. fact — 2003 Ellington, Calculator effects on achievement and attitude.
- 36. fact — 2006 Ellington, Non-CAS graphing-calculator meta-analysis.
- 37. fact — 2017 Lin, Liu, and Paas, Spell checkers and incidental spelling learning.
- 38. fact — 2006-04-25 Figueredo and Varnhagen, Spell-checker revision study.
- 39. fact — 2022-04 McCarthy et al., Mechanical and strategy writing feedback.
- 40. fact — 2008-03 Ishikawa et al., Wayfinding with GPS, maps, and direct experience.
- 41. fact — 2020-04-14 Dahmani and Bohbot, GPS use and spatial memory.
- 42. fact — 2017-02-13 Gramann et al., Landmark-based navigation instructions.
- 43. fact — 2021-04-08 Clemenson et al., Navigation guidance preserving cognitive maps.
- 44. fact — 2023 Jabbour et al., AI advice, bias, and clinician decisions.
- 45. fact — 2021 Gaube et al., Physician reliance on correct and incorrect advice.
- 46. fact — 2024 Goh et al., Large language models and physician reasoning.
- 47. fact — 2025 Budzyń et al., Endoscopist de-skilling risk after AI exposure.
- 48. fact — 2026 Pedersen et al., Prospective endoscopy performance after AI removal.
- 49. fact — 2014 Casner, Geven, and Williams, Pilot manual skill under automation.
- 50. fact — 2013 FAA PARC/CAST, Operational Use of Flight Path Management Systems.
- 51. fact — 2017 De Boer and Hurts, Automation surprises in aviation.
- 52. fact — 2017/2018 Page and Gehlbach, Pounce randomized enrollment trial.
- 53. fact — 2024 working paper; 2026 publication Page et al., Course-specific academic chatbot trial.
- 54. fact — 2024-10-08 Georgia State, TEACH ME randomized trials launch.
- 55. fact — 2025-02-04 OpenAI and CSU, CSU campus-wide deployment announcement.
- 56. fact — 2025-10-02 Virginia Tech, ChatGPT Edu Pilot Findings Report.
- 57. fact — 2026-07-23 Dumlao et al., Generative AI availability, grades, and satisfaction.
- 58. fact — accessed 2026-09-03 UVU, current institutional data.
- 59. fact — 2026-06-12 UVU, Student Right-to-Know retention and graduation disclosure.
- 60. fact — accessed 2026-09-03 UVU, IRB FAQ.
- 61. fact — accessed 2026-09-03 UVU, IRB application process.
- 62. fact — accessed 2026-09-03 UVU, IRB exempt categories.
- 63. fact — accessed 2026-09-03 U.S. Department of Education, FERPA studies exception.
- 64. fact — 2012-09-04 CONSORT Group, cluster-randomized trial reporting guidance.
- 65. fact — accessed 2026-09-03 American Economic Association, RCT registry policy.
- 66. fact — 2026-03-04 OpenAI, study-mode learning experiment and measurement suite.
- 67. fact — 2026 Khan Academy, Khanmigo product experiments.
- 68. fact — 2026-05 technical report page Google, LearnLM evaluations.
NOT_RUN
- UVU contact: not run, as required.
- Vendor contact: not run, as required.
- Private UVU gateway-course DFW extraction: not run; public current rates were not found.
- Student-level UVU power simulation: not run; real section sizes, outcome variance, missingness, and within-section correlation were unavailable.
- IRB submission or determination: not run; only UVU’s public process was reviewed.
- Independent full-text risk-of-bias scoring for every source: not run.
- New confidence intervals from published standard errors: not run; intervals are given only where sources reported them.
- Meta-analysis combining generative-AI studies: not run; interventions and outcomes were too different for a defensible pooled estimate.
- Deployment, data collection, student tracking, messages, purchases, and public release: not run.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/44-learning-outcomes-evidence.md in the research pack.
Trust, sources, and error handling: the interface
Appendix X in five lines
- QuestionHow can students know whether an answer can be trusted, see what supports it, and get mistakes fixed?
- AnswerUse one trust rule that shows exact source passages and limits, says not found when needed, and tracks each mistake until it is fixed.
- Deciding numbers2,000 students served by Wilsonfact10-second first-word targetestat least 95% right-source rate for cited claimsest
- What the plan doesKeep one student front door, add exact-passage links, track repair cases, use hint-first tutoring, provide a real person handoff, and enforce pilot gates.
- Still unknownLive Wilson and LibreChat behavior, support routes, and the exact model-and-document error rate are UNKNOWN; signed-in tests and benchmarks are NOT_RUN.
Trust is won or lost at the moment a student sees a source, a doubt, a mistake, or a refusal — and the plan had chosen a front door and a router without designing that moment. This appendix measures how wrong each model class is by task, what grounding on documents fixes and what it does not, what the research says about over-trust and under-trust, which citation and uncertainty designs measurably help, how to handle errors and refusals, what conversation design does for tutoring, and how to test all of it inside the 14-week pilot. Research cutoff September 3, 2026; 110-plus dated sources. fact measured · est designs and thresholds · unknown untested at UVU.
What this changes in the plan
1. One campus trust contract across Wilson (UVU's existing course assistant), LibreChat, and every routed model: the same source states, uncertainty states, refusals, report route, and human route. The review's sharpest point: do not launch a second, competing student front door until UVU decides Wilson's role — the fleet can serve behind Wilson or inside approved course tools first. 2. A citation badge is not enough. Every sourceable claim opens the exact supporting passage; every answer shows whether it came from approved UVU material, an uploaded file, the open web, or model memory; an evidence-sufficiency gate lets the assistant say "not found" or "sources conflict" instead of filling the gap. Raw confidence percentages are not shown unless UVU validates them for that model, task, and course — the studies show confident tone and unverified numbers miscalibrate trust in both directions. 3. Errors become repair cases with a receipt, an owner, a status, a visible correction notice, and correction memory — not a thumbs-down counter. Verifiable tools (calculators, code tests) are used wherever the task permits. 4. Tutoring is student-first and hint-first, enforced outside the model prompt where sequence matters: the student attempts the work before advice; steps before final answers. 5. Identity and route stay visible — "UVU AI assistant — not a human" and the active route on screen throughout, which is also Utah's safe harbor. 6. Response time: keep the 10-second first-word target, add truthful status and restartable streaming. 7. Helpful refusals and a real, acknowledged human route before broad student use — and if a staffed route cannot be funded, remove the features that imply one; never fake it. 8. Test it in the pilot: named tasks, participants, calibration and error-recovery metrics, and launch gates; publish results by task and route, never one blended "accuracy" or "trust" score. est
Research date: fact — September 3, 2026, MDT.
Evidence labels: fact means a dated source reports it; fact-derived is arithmetic from reported data; est is a planning choice with its basis stated; unknown means the evidence does not support a number. Benchmark failure is not called hallucination unless the study defines it that way. Model names, statute numbers, weeks, and source-list numbers are identifiers rather than measurements.
Executive verdict
- 1. Keep LibreChat as an engine candidate, but do not make it a second, conflicting campus identity beside Wilson.
- 2. Give Wilson and LibreChat one shared trust contract: the same sources, uncertainty states, refusals, reports, and human routes.
- 3. A citation badge is not enough; every sourceable claim should open the exact supporting passage.
- 4. Show whether an answer came from approved UVU material, an uploaded file, the open web, or model memory.
- 5. If the available evidence does not answer the question, the assistant should say so instead of completing the gap.
- 6. Do not show a raw confidence percentage unless UVU validates it for that model, task, and course.
- 7. Make the student attempt the work before showing tutoring advice; use hints and steps before final answers.
- 8. Treat feedback as a repair case with a receipt, owner, status, and correction—not as a thumbs-down counter.
- 9. Keep “UVU AI assistant — not a human” visible throughout the interaction.
- 10. Do not launch broad student use until source fidelity, error recovery, crisis routing, and accessibility pass the pilot tests.
How wrong, how often
There is no defensible universal “hallucination rate.” The answer changes with the task, model settings, retrieval quality, tools, prompt, grading method, and whether abstention is allowed.
No public evaluation tests the plan’s current local, heavy, and frontier models on the same factual, citation, math, code, and uploaded-document tasks. That combined comparison is unknown and should be produced during the pilot.
Published model evidence
| Model class and dated source | Factual questions | Math | Code | Documents and citations |
|---|---|---|---|---|
| Current local class: Qwen3.8-27B, official card, fact — August 14, 2026 | GPQA accuracy fact 89.2%; raw miss fact-derived 10.8 percentage points. This is scientific reasoning, not a conversational hallucination test. | A comparable text-math rate is unknown. Visual MathVision accuracy was fact 90.0% without code execution and fact 94.6% with it. | SWE-bench Pro resolved fact 61.7%; unresolved fact-derived 38.3 points. LiveCodeBench scored fact 90.3%, under a different harness. | True factual-hallucination, citation, and uploaded-summary rates are unknown. Qwen model card |
| Older local comparator: Qwen3-32B thinking, technical report, fact — May 14, 2025 | GPQA accuracy fact 68.4% across fact 198 questions sampled 10 times, or fact-derived 1,980 generations. | MATH-500 accuracy fact 97.2%; raw miss fact-derived 2.8%. | LiveCodeBench accuracy fact 65.7%; raw miss fact-derived 34.3 points. | No matching citation or document-summary rate. Qwen3 report |
| Local-class caution: Gemma 3 27B IT, official card, fact — March 2025 | SimpleQA accuracy fact 10.0% over fact 4,326 questions. The remaining fact-derived 90.0 points include abstentions and errors; they are not all hallucinations. | MATH accuracy fact 89.0% over the fact 5,000-problem test set. | HumanEval pass rate fact 87.8% over fact 164 problems. | FACTS Grounding score fact 74.9; its deficit is not a literal count of bad responses. Gemma 3 model card |
| Heavy class: gpt-oss-120b, official card, fact — August 5, 2025 | On SimpleQA without browsing: accuracy fact 16.8%, hallucination fact 78.2%, and abstention/other fact-derived 5.0%, over fact 4,326 questions. PersonQA accuracy was fact 29.8% and hallucination fact 49.1%. | AIME 2025 with tools scored fact 97.9% over fact 30 problems; raw deficit fact-derived 2.1 points. | SWE-bench Verified resolved fact 62.4%, or fact 312 of 500 tasks; fact 188 of 500 remained unresolved. | Exact citation and current private-document rates are unknown. gpt-oss evaluation |
| Heavy alternative: Mistral Small 4, official card, fact — March 16, 2026 | GPQA accuracy fact 71.2% over the standard fact 198 prompts; raw miss fact-derived 28.8 points. | Auditable absolute results for the plan’s use are unknown. | Auditable absolute results in accessible text are unknown. | Hallucination, citation, and document-QA rates are unknown. Mistral Small 4 |
| Frontier comparator: GPT-5.6, official release, fact — July 9, 2026 | GPQA accuracy fact 94.6%; raw miss fact-derived 5.4 points. This is not a campus factuality rate. | FrontierMath scores were fact 89% on lower tiers and fact 83% on the hardest tier; deficits fact-derived 11 and 17 points. | SWE-bench Pro resolved fact 64.6%; unresolved fact-derived 35.4 points. | A current absolute uploaded-document or citation error rate is unknown. GPT-5.6 |
| Frontier browsing evidence: GPT-5 system card, fact — August 7, 2025 | On representative factual traffic with browsing, incorrect claims were fact 4.5% and responses containing a major factual error were fact 4.8%. Sample size was unknown. Without web on SimpleQA, accuracy was fact 55%, hallucination fact 40%, and abstention/other fact-derived 5%. | Not comparable here. | Not comparable here. | Browsing helps materially, but errors remain and the production grader was itself a model. GPT-5 system card |
Citations and uploaded documents
Grounding means giving the model retrieved evidence. It changes the failure pattern; it does not make the answer safe by itself.
- RAGTruth contains fact 17,790 retrieval-grounded responses. At least one annotated hallucination appeared in fact 1,724 of 5,934 QA responses, or FACT-derived 29.05%, and fact 1,686 of 5,658 summaries, or FACT-derived 29.80%. The models date mainly from 2023, and later work suggests subtle errors were undercounted. These are warning rates, not forecasts for UVU. RAGTruth, ACL 2024
- FACTS Grounding supplied the full source document for fact 1,719 prompts. Scores were fact 78.8 for GPT-4o, fact 83.6 for Gemini 2.0 Flash Experimental, and fact 79.4 for Claude 3.5 Sonnet. These scores combine judges and splits; their deficits are not literal error percentages. factS Grounding, January 6, 2025(https://arxiv.org/abs/2501.03200)
- In the ALCE citation benchmark, GPT-4 answering ELI5 questions with fact 20 retrieved passages achieved citation recall fact 48.5 and citation precision fact 53.4. About half of the tested long answers were not fully supported by their citations. On ASQA, citation recall was fact 73.0 and precision fact 76.5. ALCE, December 6, 2023
- A study of fact 636 generated citations found fabrication in fact 55% of GPT-3.5 citations and fact 18% of GPT-4 citations. Among citations that existed, substantive errors remained in fact 43% and fact 24%, respectively. These were closed-book 2023 systems, but they prove that bibliography-shaped text is not evidence. Walters and Wilder, September 7, 2023
- “Sufficient Context” found that irrelevant retrieval can reduce abstention and increase wrong answers. For one Gemma condition, incorrect answers rose from fact 10.2% without context to fact 66.1% with insufficient context. A sufficiency-aware answer gate improved correct-among-answered by fact 2–10 percentage points. ICLR 2025 paper
- A citation can support a sentence without having caused the model to produce it. In an adversarial test, a retrieval model attached apparently supporting citations to claims driven by uncited or random material in as many as fact 57% of one condition. Correctness is not Faithfulness, December 23, 2024
What remains wrong after grounding:
- The right passage was never retrieved.
- An old or unauthorized document outranks the current source.
- The passage is relevant but does not support the claim.
- The model merges facts from different people, dates, policies, or courses.
- A correct answer is surrounded by invented explanation.
- A summary omits an important exception or qualifier.
- A citation exists but points to the wrong section.
- Conflicting sources are silently collapsed into one answer.
- The model answers even though the evidence is incomplete.
- The automatic citation checker is also wrong. MiniCheck found only fact 75.3% balanced accuracy for GPT-4 across fact 10 grounding datasets. MiniCheck, November 12, 2024
Therefore, UVU should measure source retrieval, claim support, answer correctness, and abstention separately.
Trust calibration
Trust calibration means using the assistant when it is right and checking or rejecting it when it is wrong. High trust is not the goal. Appropriate trust is.
What the studies say
- Confident length can masquerade as quality. In controlled experiments with fact 60, 60, and 59 participants, longer explanations increased confidence without improving discrimination between right and wrong answers. Human discrimination was fact AUC 0.54 for long explanations and fact AUC 0.57 for uncertainty-only answers. Language tied to measured model uncertainty improved calibration, though it did not improve users’ underlying subject knowledge. Nature Machine Intelligence, January 21, 2025
- Citations can raise trust before they are checked. A randomized study reports fact 303 participants and fact 1,976 citation-bearing answers. Only fact 193 citations, or FACT-derived 9.77%, were opened or hovered over. fact 83 of 197 citation-condition participants, or FACT-derived 42.1%, checked at least one. Random citations still increased trust unless people inspected them. The paper has an internal count conflict: its printed condition totals sum to fact-derived 305, not the reported fact 303. AAAI, April 11, 2025
- Uncertainty wording can reduce blind agreement. In a preregistered medical-question study with fact 404 participants and an intentionally fact 50%-accurate assistant, agreement fell from fact 80.9% with unhedged answers to fact 74.8% with first-person uncertainty, while answer accuracy rose from fact 63.9% to 72.8%. It did not produce a significant increase in source checking. Kim et al., FAccT, June 5, 2024
- Confidence displays work only when the confidence is calibrated. A study with fact 72 participants found that a calibrated confidence display helped people distinguish when to follow an assistant, but did not improve combined human-AI accuracy. Zhang, Liao, and Bellamy, January 27, 2020
- Forcing a small pause reduces overreliance. In a study with fact 199 participants, asking for an independent judgment or delaying advice reduced wrong-AI overreliance from fact 0.64 to 0.48 on one decision measure. Participants liked the more demanding interfaces less. Buçinca et al., April 2021
- Visible errors usually hurt more than correct answers help. Algorithm-aversion experiments found that people abandoned an algorithm after observing its mistakes even when it still outperformed a human forecaster. Across the reported studies, fact 610 of 741 participants, or FACT-derived 82.3%, saw the algorithm outperform the human. Dietvorst, Simmons, and Massey, 2015
- But the first error does not always cause permanent distrust. A legal-advice experiment with fact 208 participants and fact 14 cases found an immediate trust drop after a planted error, followed by relatively rapid behavioral recovery. This contradicts a universal “one error destroys trust” claim. Kahr et al., April 5, 2024
- Honest advance warning can soften later failure. A randomized chatbot experiment with fact 558 users found that candid low-performance messaging reduced discontinuance after a response failure; inflated claims did not. Weiler, Matt, and Hess, December 22, 2021
Interface conclusion
Use evidence-based uncertainty, not personality-based uncertainty.
Good:
- “The current course policy says …”
- “I found this in the syllabus and assignment page.”
- “The sources conflict: the syllabus says X, while the newer announcement says Y.”
- “I could not find this in the materials I searched.”
- “This is an inference, not a statement in the source.”
- “I can offer a general explanation, but it is not verified against UVU material.”
Avoid:
- “I’m definitely right.”
- “I’m 92% confident” when the number has not been validated.
- A long rationale that adds no evidence.
- Green checkmarks based only on model self-confidence.
- Citations that open only a home page or document title.
Showing sources and uncertainty
Recommended answer anatomy
For institutional, course, policy, research, and uploaded-file answers:
- 1. Give the short answer.
- 2. Highlight each checkable claim.
- 3. Attach an inline citation to the exact passage.
- 4. On selection, show the passage, title, owner, effective date, page or timestamp, and access boundary.
- 5. Provide “Show what I used,” listing all retrieved sources—not just those the model cited.
- 6. Mark conflicts, missing evidence, and expired sources.
- 7. Provide “Report this answer” beside the evidence, not buried in settings.
The source panel should distinguish:
- Approved UVU source
- Course source
- Your uploaded file
- Open-web source
- General model knowledge — not independently verified
A document title alone is not enough. The user should be able to inspect the supporting sentence without searching the entire file.
What campus systems document today
| System | Published behavior | Important gap |
|---|---|---|
| UVU Wilson | fact — current page checked September 3, 2026: uses UVU resources on the website and available course material in selected Canvas courses. It cannot access grades, private files, or tests, and students acknowledge that it is AI and may hallucinate. A Qualtrics feedback link exists. Wilson AI | Public documentation does not show inline claim citations, uncertainty states, a feedback receipt, correction status, response time, or crisis flow. Actual authenticated behavior is unknown. |
| Wilson course deployment | fact — February 11, 2026 interview: UVU’s CIO reported use in fact 44 courses serving 2,000 students, after a fact fall 2023 biology pilot. Wilson can link to the relevant point in a lecture recording and is intended to coach rather than give answers. EdTech interview | This is a named executive interview, not an independent outcome evaluation. Source-checking behavior is unknown. |
| TritonGPT | fact — August 13, 2024 release: “See Context” exposes source documents, webpages, relevance scores, and direct interaction with selected sources. UC San Diego release | Current live behavior and student source-opening rates are unknown. A relevance score is not proof that a passage supports a claim. |
| TritonGPT course tutors | fact — 2025–26 published results: instructors can include or remove Canvas sources and select Socratic or directive behavior. Among fact 68 survey responses, fact 81% said it helped explain concepts, fact 86% found it easy, and fact 67% wanted it in future courses. Instructional program | Response rate and course mix are unknown. The survey does not measure citation checking or factual accuracy. |
| ZotGPT ClassChat | fact — page updated May 20, 2026: faculty can upload curriculum, set instructions and guardrails, restrict access, and view activity. UCI ClassChat | Public evidence for per-claim citations, uncertainty labels, crisis language, or closed-loop corrections is unknown. |
| TitanGPT | fact — page published August 18, 2026: available to students and employees; users are told they remain responsible for mistakes and must follow instructor policy. CSUF TitanGPT | Public source, uncertainty, correction, and crisis details are unknown. |
| Khanmigo | fact — documentation updated through August 2026: uses moderation, adult notification, limitations messaging, feedback categories, and optional severe-content administrator alerts. Safety, feedback | It does not publish a correction service level or exact student-facing crisis script. These are vendor descriptions, not independent outcome evidence. |
Measured checking behavior by students inside Wilson, ZotGPT, TritonGPT, TitanGPT, or Khanmigo is unknown. The best direct citation study above was not campus-specific. A separate undergraduate study with fact 66 students found a mean of only fact 1.76 correct classifications out of 4 after a short reference-verification lesson. Franzoni Velázquez et al., 2024
This supports two actions: make checking much easier, and teach lateral reading—opening another source or site to verify ownership and claims. In a college intervention with fact 230 students, the share that both used lateral reading and correctly judged at least one source rose from fact 7.0% at baseline to fact 61.0% in the intervention group at post-test. Brodsky et al., 2021
Error handling and feedback loops
“Report this response”
LibreChat can display thumbs-up and thumbs-down controls, but its published behavior is rating capture, not case management. LibreChat interface documentation, checked September 3, 2026
The pilot needs this flow:
- 1. User selects Report this response.
- 2. User chooses: wrong answer, wrong/outdated source, unsupported claim, harmful or biased, bad refusal, privacy concern, tutoring problem, or other.
- 3. The interface previews what will be sent and lets the user omit their prompt where policy permits.
- 4. The system records the response, model and router versions, retrieved-source identifiers, and course/source version.
- 5. The user receives an immediate case number and status.
- 6. A named owner reviews it.
- 7. An approved correction updates the authoritative source or correction layer.
- 8. The affected answer shows “Corrected” with date and explanation.
- 9. The reporter receives closure when contact is permitted.
Planning service levels:
- Receipt: est within 2 seconds. Basis: the plan’s existing upload-acknowledgment target.
- Human ownership: est within 1 business day. Basis: a campus support target, not published evidence.
- Ordinary correction or reasoned rejection: est within 5 business days. Basis: a proposed pilot operating target.
- Urgent privacy, safety, or cross-user leakage: follow the incident process, not the ordinary queue.
Correction memory
Do not store corrections as the model’s personal “memory.”
LibreChat memory is a user-controlled key/value store, not a governed institutional correction record. LibreChat memory documentation
Use a separate correction record containing:
- Claim or source affected.
- Approved replacement.
- Source owner and reviewer.
- Effective and expiry dates.
- Course and audience scope.
- Reason for the change.
- Retrieval re-index status.
- Tests rerun.
- Earlier answer identifiers that need a correction notice.
A student correction suggestion is evidence for review, not automatic truth.
Escalation to a person
The interface should name the destination before sending anything:
- Course meaning or grading policy → instructor or named course staff.
- Tutoring help → academic tutoring.
- Account or service failure → service desk.
- Accessibility barrier → accessibility support.
- Privacy or data concern → privacy/security queue.
- Immediate safety concern → emergency and crisis services.
Never say “I alerted someone” until an acknowledged route confirms receipt.
Response-time perception
The program’s current service target is est p95 first token within 10 seconds for tutoring/chat, with an est p95 queue budget within 2 seconds. Service-engineering findings
A 2026 experiment with fact 240 participants, fact 3 tasks per participant, and first-token delays of fact 2, 9, or 20 seconds found that the fact 2-second condition could be perceived as less thoughtful or useful than fact 9- or 20-second conditions, although long waits also caused frustration. Response latency study, 2026
Keep the est 10-second operational target, but do not fake thinking. Show real status:
- “Searching approved course sources.”
- “Checking whether the passages support the answer.”
- “Using the calculator.”
- “The local service is busy; you can wait or use the approved cloud route.”
- “The answer stopped. Retry from the last completed step.”
Streaming improves perceived responsiveness and recovery. It is not evidence of correctness.
Refusals
A useful refusal has three parts:
- 1. State the boundary.
- 2. Give a short reason.
- 3. Offer the closest safe action.
Examples:
- Academic work: “I can’t complete this graded answer for you. Show me your first step, and I’ll give one hint.”
- Grades: “I can’t assign or predict your grade. I can help compare your draft with the published rubric, one criterion at a time.”
- Missing evidence: “I couldn’t find that in the approved course materials. I can show what I searched or help you ask the instructor.”
- Regulated advice: “I can give general information, but I can’t provide personal legal, medical, mental-health, or financial advice. Here is the appropriate UVU service.”
In a refusal study with fact 480 participants and fact 3,840 comparisons, partial compliance—safe general help without dangerous detail—reduced negative reactions by more than fact 50% compared with a flat refusal. “Let Them Down Easy,” November 2025
Crisis message
Use the crisis pattern already developed in the duty-of-care findings:
I’m sorry you’re dealing with this. I’m an AI tutor, not a crisis service. If you or someone else may be in immediate danger, call 911 or UVU Police now. For suicide or emotional crisis, call or text 988. You do not need to repeat the details here.
This is an est design based on UVU, NIMH, and federal crisis guidance—not a claim that a bot can assess risk. Duty-of-care findings, UVU crisis services, NIMH guidance revised 2024
The assistant should stop ordinary tutoring, avoid long disclaimers, use the smallest necessary question, and never promise monitoring, confidentiality, dispatch, or response time unless those are real. Publicly validated false-positive and false-negative rates for named campus crisis bots remain unknown.
AI identity and Utah law
Utah’s current disclosure chapter took effect fact May 7, 2025. It requires disclosure on a clear request in covered consumer transactions and advance disclosure for regulated occupations’ high-risk AI interactions. Its safe harbor is broader: clearly identify the system at the outset and throughout as generative AI, not human, or an AI assistant. Whether a no-charge UVU assistant is a covered consumer transaction is unknown. Utah Code Title 13, Chapter 77
Use:
UVU AI assistant — not a human
Place it above the first prompt and keep it visible in the header, shared chats, errors, mobile views, and routed model changes.
Disclosure’s trust effect is mixed:
- A randomized field study involving fact 11,000 truck drivers found initial voice-bot disclosure reduced response probability by about fact 11%. Management Science, August 26, 2024
- In an online experiment with fact 194 participants, only fact 24% in the disclosed condition correctly recalled the chatbot introduction. AI & Society, January 5, 2024
These settings do not predict campus behavior. They do show that one small opening label is insufficient. Identity disclosure is a transparency and legal control, not a substitute for evidence or quality.
Conversation design for tutoring
Recommended default
The assistant should:
- 1. Ask what the student is trying to learn.
- 2. Ask for the student’s current work or first attempt.
- 3. Identify one misconception or missing step.
- 4. Give one hint.
- 5. Wait for a response.
- 6. Increase help gradually: prompt → hint → partial step → analogous worked example → full explanation when course policy allows.
- 7. Ask the student to summarize or apply the idea without assistance.
- 8. Offer a direct-explanation mode for accessibility, review, or explicit instructor-approved use.
“Socratic” should mean guided questioning, not endless questions. Students must be able to request a direct explanation where that is pedagogically and academically permitted.
Evidence
- A preregistered field experiment with nearly fact 1,000 high-school mathematics students found that unrestricted GPT access improved assisted-practice grades by fact 48% but reduced later unaided grades by fact 17%. A safeguarded tutor improved assisted practice by fact 127% without the later loss. Its guardrails used teacher material, hints, and limits on giving answers. Bastani et al., June 25, 2025
- A Harvard physics trial with fact 194 eligible students found better immediate post-test results from an expert-built AI tutor than from an active-learning class. The system prompt alone was not reliable enough to maintain the lesson sequence; the researchers added a custom structure that moved students through each problem part and supplied expert solutions. Scientific Reports, June 3, 2025
- Tutor CoPilot’s preregistered field trial involved fact 900 tutors, fact 1,800 students, and more than fact 550,000 messages. Access increased lesson-topic mastery by fact 4 percentage points, with a fact 9-point gain for students of lower-rated tutors. It increased guiding questions and reduced answer-giving, but no significant end-of-year math-test gain was reported. Tutor CoPilot, revised January 26, 2025
These results support structured tutoring, not a generic “be Socratic” prompt.
“Explain your reasoning” prompt families
Ask the student to expose their reasoning:
- “What have you tried?”
- “Which rule do you think applies?”
- “Where did your result stop matching the example?”
- “Explain this step in your own words.”
- “What evidence supports that claim?”
- “How would you check the answer without the assistant?”
Ask the model for:
- A concise rationale.
- The evidence and tools used.
- A checkable intermediate result.
- The next instructional move.
- A different worked example.
Do not present the model’s hidden chain of thought as reliable proof. Visible reasoning can rationalize a biased answer; one NeurIPS study found accuracy reductions as large as fact 36% across 13 tasks while explanations failed to mention the biasing prompt feature. Turpin et al., NeurIPS 2023
What LibreChat and the router can enforce
LibreChat can configure:
- Persistent agent names and descriptions.
- Welcome text and terms.
- Course-specific system instructions.
- Role permissions.
- File search and file citations.
- Model and endpoint availability.
- Feedback controls.
- Uploaded text.
- Streaming behavior.
LibreChat model specifications, agent citation configuration, upload behavior
The router should enforce outside the model:
- Data classification and allowed destination.
- Course and user access to each source.
- Local, free-cloud, or metered-frontier routing.
- Calculator, code runner, or other verified tool use.
- Evidence-required mode.
- No-sufficient-source abstention.
- Safety classification and crisis path.
- Pinned model, prompt, retrieval, and policy versions.
- Grade and private-file exclusion.
A custom service or interface layer is needed for:
- Claim-to-passage highlighting.
- “Show what I used.”
- Source freshness, conflict, and authorization state.
- Structured reports and correction status.
- Human acknowledgement and escalation.
- Governed correction memory.
- Calibrated uncertainty state.
- Consistent identity disclosure across all routes.
- Sequential tutoring state that a prompt alone cannot reliably maintain.
Evaluation
Participants
- Students: est 48 unique participants. Basis: est 4 rounds or strata of 12, covering first-generation and commuter students, writing/research work, quantitative/code work, and accessibility or English-language needs. Participants may overlap characteristics.
- Instructors: est 12 unique participants. Basis: course-policy, source-control, tutoring, and correction-workflow review.
- Support and safety staff: est 6 unique participants. Basis: service, accessibility, privacy, and synthetic crisis tabletop roles.
- Total: est 66 unique participants.
This is a formative usability sample, not a powered learning-effect trial. Faulkner found that samples of fact 5 users could uncover anywhere from fact 55% to 99% of observed usability problems, which supports repeated diverse rounds rather than one tiny test. Faulkner, 2003
A formal outcome claim needs a separate power analysis after UVU obtains baseline variance; required enrollment is unknown.
Pilot tasks
Each participant receives only role-appropriate tasks:
- Find a course rule and inspect the exact passage.
- Ask a question the approved sources do not answer.
- Resolve two sources that conflict by date.
- Summarize an uploaded document containing a planted unsupported claim.
- Judge an answer with a valid citation and one with a wrong citation.
- Solve a math problem using the calculator and verify the result.
- Repair code and run a provided test.
- Complete a tutoring problem without receiving the answer first.
- Report an incorrect answer, follow its status, and inspect the correction.
- Encounter an academic-integrity or regulated-advice refusal.
- Recover from delay, interrupted streaming, or route failure.
- Staff only: run synthetic crisis and privacy scenarios.
Metrics
Measure behavior, not only opinion:
- Task success against an instructor-approved answer key.
- Time on task.
- Correct acceptance of correct advice.
- Correct rejection of wrong advice.
- Overreliance and underreliance.
- Confidence-to-correctness calibration and Brier score.
- Source opens, exact-passage views, and successful source verification.
- Claim-level citation precision and citation coverage.
- Correct abstention when evidence is absent.
- Trust before an error, immediately after it, and after repair.
- Error-recovery success.
- Report receipt, ownership, correction, and closure times.
- First-token and complete-answer latency by route.
- Abandonment and retry behavior.
- System Usability Scale, a fact 10-item standardized questionnaire, plus open comments. Brooke, 1996
- Keyboard, screen-reader, zoom, contrast, focus, and streaming-announcement task success.
Proposed launch gates
These are planning requirements, not published norms:
- Cross-user or unauthorized-source disclosure: est 0 observed cases.
- Claim-level citation precision: est at least 95% on the pilot gold set.
- Citation coverage for sourceable claims: est at least 90%.
- Abstention on unsupported institutional questions: est at least 90%.
- Core-task success: est at least 80% in every tested student group.
- Wrong-answer overreliance: est no more than 20% on planted-error tasks.
- Synthetic imminent-risk message and route: est 100% correct, with est 0 unacknowledged claimed handoffs.
- Accessibility blockers: est 0 unresolved launch-blocking defects.
- SUS planning threshold: est at least 70, used as a comparison signal rather than proof of safety.
Basis: these are risk-based pilot gates selected to force correction before scale. UVU should revise them after the first baseline round rather than lowering them to excuse a failing interface.
Schedule inside the plan’s 14-week sequence
| Weeks | Work and proof |
|---|---|
| est Weeks 1–2 | Approve task corpus, source versions, issue categories, consent language, and research/IRB determination. Establish baseline answers and planted errors. |
| est Week 3 | Expert review of citations, uncertainty states, refusals, keyboard flow, screen-reader announcements, and source authorization. |
| est Week 4 | Moderated round with est 12 students. Measure first error, source checking, abstention, and repair discovery. |
| est Week 5 | Fix only observed high-severity problems; freeze the next test build and source set. |
| est Week 6 | Moderated round with est 12 new students and 6 instructors. Test course policy, tutoring sequence, and instructor source controls. |
| est Week 7 | Synthetic safety and incident tabletop with est 6 support staff. Prove acknowledgement and records. |
| est Weeks 8–10 | Limited field pilot with est 24 new students and 6 additional instructors. Collect aggregate behavior, not private prompt surveillance. |
| est Week 11 | Correct sources and interface defects; freeze model, router, retrieval, and prompt versions. |
| est Week 12 | Blinded challenge set covering all three model routes, missing evidence, conflicting sources, math, code, summaries, and citations. |
| est Week 13 | Retest with an est 12-person subset from the field pilot; no additional unique student count. Verify repair and accessibility regressions. |
| est Week 14 | Publish the evidence packet and decide: stop, extend pilot, or scale. No automatic scale-up. |
Cross-domain learnings
- Clinical decision support: the FDA’s fact January 2026 guidance emphasizes that a professional must be able to independently review the basis of a recommendation. UVU should apply the same idea: show inputs, sources, limitations, and known unknowns rather than asking users to trust a score. FDA guidance
- Aviation: flight systems show mode, intended action, and limits because hidden mode changes create “automation surprise.” UVU should always show which assistant, route, evidence mode, and tools are active. FAA automation report
- Safety engineering: NIST separates transparency from accuracy. A transparent system can still be wrong, but opacity prevents review, accountability, and repair. NIST AI RMF, January 26, 2023
- Information literacy: lateral reading works better than teaching users to judge a page by appearance. Source links should make checking fast, but the pilot must also teach students to leave the answer and inspect independent evidence.
- Tutoring: good tutors control the help sequence. The strongest education studies added teacher-authored material, question order, hints, and deliberate withholding. They did not rely on a personality prompt alone.
Devil’s advocate
The strongest case against this recommendation is that it creates too much interface and operational work around a system that may still be unreliable.
The evidence is fragmented. Many error studies use older models or laboratory questions. Many trust studies use medicine, forecasting, legal vignettes, or crowdsourced participants rather than UVU students. Citation panels can increase trust without checking. Uncertainty warnings can lower useful reliance. Cognitive forcing adds friction and is often disliked. A custom source, correction, and escalation layer can become a larger project than the local-model service.
UVU also already has Wilson. Adding LibreChat as a visible second front door could split support, source governance, analytics, and student expectations. The simpler safe path may be to use the new model fleet behind Wilson or within narrowly approved course tools, instead of launching another general student chat interface.
That objection is strong. It changes the implementation path, but not the need for source proof, abstention, repair, and human routes.
What would change the recommendation
- A live UVU evaluation shows that Wilson already provides exact claim-to-passage citations, source conflicts, correction status, and acknowledged escalation.
- The selected LibreChat release supplies the same controls without a maintained custom fork.
- A shared UVU benchmark shows one route is reliable enough to simplify the uncertainty design.
- Student testing shows the source panel materially harms task completion without improving verification, and a simpler passage preview performs better.
- UVU counsel gives a narrower written interpretation of the identity requirement.
- Accessibility testing shows the proposed inline evidence interaction is not usable and identifies a better equivalent.
- Pilot data show that structured tutoring harms learning or produces unacceptable frustration in UVU courses.
- A staffed feedback or crisis route cannot be funded. In that case, remove the claims and features that imply such support; do not fake the route.
What the plan should change
Ranked:
- 1. Use one campus trust contract across Wilson, LibreChat, and every routed model.
- 2. Do not create a competing student front door until UVU decides Wilson’s role.
- 3. Require exact-passage citations for institutional, course, research, and uploaded-document claims.
- 4. Add an evidence-sufficiency gate that can return “not found” or “sources conflict.”
- 5. Build structured reporting, governed corrections, and visible correction notices.
- 6. Keep the AI identity and active route visible throughout every interaction.
- 7. Enforce student-first, hint-first tutoring outside the model prompt where sequence matters.
- 8. Use calculators, code tests, and other verifiable tools for tasks that permit them.
- 9. Add helpful refusals and a real, acknowledged human route before broad student use.
- 10. Preserve the existing EST 10-second chat target, but add truthful status and restartable streaming.
- 11. Run the proposed trust and usability work inside the 14-week pilot before scaling hardware or access.
- 12. Publish results by task and route; never publish one blended “accuracy” or “trust” score.
Source log
Source-log ordinals are reference labels, not measurements.
- 1. fact — accessed September 3, 2026: Current UVU program.
- 2. fact — accessed September 3, 2026: Service-engineering findings.
- 3. fact — accessed September 3, 2026: Duty-of-care findings.
- 4. fact — accessed September 3, 2026: Utah-law and peer findings.
- 5. fact — August 14, 2026: Qwen3.8-27B model card.
- 6. fact — May 14, 2025: Qwen3 technical report.
- 7. fact — March 2025: Gemma 3 model card.
- 8. fact — August 5, 2025: gpt-oss-120b evaluation.
- 9. fact — March 16, 2026: Mistral Small 4 model card.
- 10. fact — July 9, 2026: GPT-5.6 release.
- 11. fact — August 6, 2026: GPT-5.6 safety update.
- 12. fact — August 7, 2025: GPT-5 system card.
- 13. fact — October 30, 2024: SimpleQA.
- 14. fact — August 2024: RAGTruth.
- 15. fact — January 6, 2025: factS Grounding(https://arxiv.org/abs/2501.03200).
- 16. fact — December 6, 2023: ALCE.
- 17. fact — September 7, 2023: Fabrication and errors in ChatGPT citations.
- 18. fact — November 12, 2024: MiniCheck.
- 19. fact — November 9, 2024: Sufficient Context.
- 20. fact — December 19, 2024: LongBench v2.
- 21. fact — June 16, 2024: Groundedness in retrieval-augmented long-form generation.
- 22. fact — July 2025: GaRAGe.
- 23. fact — December 23, 2024: Correctness is not Faithfulness.
- 24. fact — April 11, 2025: Citations and Trust.
- 25. fact — January 21, 2025: What large language models know and what people think they know.
- 26. fact — June 5, 2024: First-person uncertainty and overreliance.
- 27. fact — January 27, 2020: Confidence and explanation study.
- 28. fact — April 2021: Cognitive forcing functions.
- 29. fact — 2015: Algorithm aversion.
- 30. fact — April 5, 2024: One-error trust dynamics.
- 31. fact — December 22, 2021: Inoculation messages against chatbot failure.
- 32. fact — 2024: Student evaluation of generated references.
- 33. fact — 2021: College lateral-reading intervention.
- 34. fact — checked September 3, 2026: Wilson AI.
- 35. fact — February 11, 2026: UVU Wilson deployment interview.
- 36. fact — January 28, 2025; updated February 3, 2026: UCI ClassChat deployment.
- 37. fact — updated May 20, 2026: ZotGPT ClassChat.
- 38. fact — checked September 3, 2026: TritonGPT overview.
- 39. fact — August 13, 2024: TritonGPT “See Context” release.
- 40. fact — 2025–26: TritonGPT instructional program.
- 41. fact — effective June 3, 2025: TritonGPT terms.
- 42. fact — August 18, 2026: TitanGPT.
- 43. fact — updated July 1, 2026: Khanmigo safety features.
- 44. fact — updated August 20, 2025: Khanmigo flagged conversations.
- 45. fact — updated October 28, 2025: Khanmigo feedback.
- 46. fact — June 25, 2025: Generative AI without guardrails can harm learning.
- 47. fact — June 3, 2025: Structured AI physics tutor RCT.
- 48. fact — revised January 26, 2025: Tutor CoPilot.
- 49. fact — NeurIPS 2023: Unfaithful chain-of-thought explanations.
- 50. fact — November 2025: Let Them Down Easy.
- 51. fact — 2026: Response latency experiment.
- 52. fact — checked September 3, 2026: LibreChat agent citation configuration.
- 53. fact — checked September 3, 2026: LibreChat interface and feedback configuration.
- 54. fact — checked September 3, 2026: LibreChat uploaded-text behavior.
- 55. fact — checked September 3, 2026: LibreChat user memory.
- 56. fact — effective May 7, 2025: Utah Code Title 13, Chapter 77.
- 57. fact — August 26, 2024: Voice-chatbot identity disclosure field experiment.
- 58. fact — January 5, 2024: Disclosed versus undisclosed chatbots.
- 59. fact — January 26, 2023: NIST AI Risk Management Framework.
- 60. fact — January 2026: FDA clinical decision-support guidance.
- 61. fact — checked September 3, 2026: FAA automation human-factors report.
- 62. fact — revised 2024: NIMH crisis guidance.
- 63. fact — checked September 3, 2026: UVU crisis services.
- 64. fact — 2003: Faulkner usability sample study.
- 65. fact — 1996: System Usability Scale.
NOT_RUN
- not run: Saving this report to the requested output file.
- not run: Authenticated testing of Wilson, LibreChat, ZotGPT, TritonGPT, TitanGPT, or Khanmigo.
- not run: Benchmarking the exact quantized models, prompts, router, retrieval index, and Apple Silicon machines proposed by the plan.
- not run: Inspecting current campus feedback queues, crisis staffing, source indexes, logs, or correction records.
- not run: Contacting UVU, any university, or any vendor.
- not run: A statistical power calculation for learning outcomes; baseline UVU variance and minimum meaningful effect are unknown.
- not run: Legal advice or a UVU counsel determination about Utah’s “consumer transaction” boundary.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/45-trust-sources-interface.md in the research pack.
Accessibility and language
Appendix Y in five lines
- QuestionWhat must work so disabled students, Spanish speakers, and phone users can use the service?
- AnswerLaunch text chat first, keep LibreChat only if the exact build passes, and test uploads, voice, language, and phones one by one.
- Deciding numbersApril 26, 2027 deadlinefact18–24 paid studentsest300 paired English-and-Spanish tasksest
- What the plan doesFreeze and test the exact build, pay disabled and bilingual testers, test English and Spanish, check weak networks, and provide accessible Hypertext Markup Language pages first.
- Still unknownLibreChat’s fit with access rules and current language, phone, and weak-network needs are UNKNOWN; live screen-reader, speech, Spanish, and network tests are NOT_RUN.
The plan made accessibility a launch gate with a federal deadline and said little else. A chat that streams answers, accepts uploads, reads documents, and may speak is hard to make accessible, and UVU serves many first-generation and Spanish-speaking students. This appendix names the WCAG 2.1 AA criteria that bite on a streaming chat and how to audit them with real assistive technology, tests the chosen front door's actual accessibility evidence, covers on-device voice, neurodivergent design, Spanish-language quality of the candidate models against UVU's demographics, and phones on weak Wi-Fi — with a prioritized backlog, testers, cost, and the launch-gate checklist. Research cutoff September 3, 2026; 130-plus dated sources. fact standards, dates, published measurements · est designs and priorities · unknown what UVU has not published.
What this changes in the plan
1. LibreChat moves from "chosen front door" to "conditional candidate pending exact-build proof." No public evidence shows any current release is WCAG 2.1 AA conformant; the pilot must audit the exact pinned build with screen readers (VoiceOver, NVDA, JAWS), switch access, and magnification on complete conversations, and keep an alternative front door ready. 2. Accessible text chat launches first; uploads, math, diagrams, PDF export, and voice each pass their own gate later. Streaming is treated as a visual effect: announce short states and complete messages, never every token; focus stays under the student's control. 3. Language controls are separate: interface language, content language, and reply language, each chosen by the student; a paired UVU English/Spanish benchmark for every exact deployed model and quantization — only gpt-oss-120B has a maker-published Spanish score, and it does not prove tutoring quality, so Spanish parity is a launch test, not an assumption. 4. Reading and voice support as reversible choices — read-aloud, plain-language, short-answer, and step-by-step modes — never diagnosis-based presets; on-device speech is feasible but accent, Spanish, latency, and error rates are launch tests. 5. Paid disabled and bilingual student testers, with an IRB determination before recruitment; a funded accessibility owner; a release-specific conformance report, a known-issues statement, and a barrier-response route. 6. Phones and weak Wi-Fi: captive-portal, weak-network, interrupted-upload, and data-use tests, because UVU's phone-only share and preferred-language mix are not published. 7. Generated documents ship as accessible HTML first; PDF is an additional audited artifact. The federal deadline for UVU is April 26, 2027; existing equal-access duties apply now. est
Research cutoff: fact — 2026-09-03 MDT Scope: Read-only review of the current plan, named findings, public standards, official product records, UVU pages, and published research. Method: every date, contradiction and evidence grade is kept explicit, and anything that could not be proved is marked unknown rather than smoothed over.
Executive verdict
- Proceed with a conditional pilot, but launch accessible text chat before voice, complex uploads, diagrams, or PDF export.
- fact — deadline April 26, 2027: UVU’s web and mobile services generally must conform to WCAG 2.1 Level A and AA; existing ADA effective-communication duties apply now.
- UNKNOWN: Public evidence does not prove that any current LibreChat release is WCAG 2.1 AA conformant.
- Treat streaming as a visual effect: announce short states and complete messages, not every generated token.
- Keep focus under the student’s control and make every action work by keyboard, switch, touch, and screen reader.
- On-device voice is feasible, but accent, Spanish, latency, and hallucination performance remain launch-test questions.
- Offer read-aloud, plain-language, short-answer, and step-by-step modes as reversible choices; do not create diagnosis-based presets.
- FACT: Only gpt-oss-120B has a located maker-published Spanish knowledge score; it does not prove Spanish tutoring quality.
- UNKNOWN: UVU’s current phone-only share, preferred-language mix, poor-Wi-Fi rate, and exact accessible-chat demand are not published.
- The plan needs a funded accessibility owner, paid disabled and bilingual testers, a pinned front-door build, and evidence for each release.
WCAG labels, release numbers, dates, source-log ordinals, and arithmetic expressions identify standards or artifacts. Measured quantities and planning quantities are graded fact, est, or unknown.
WCAG 2.1 AA for a streaming chat
The legal scope is the service, not its landing page. It includes authentication, chat history, model selection, streaming, uploads, errors, generated content, mobile layouts, and downloads. Contractor or open-source status does not transfer UVU’s responsibility. Newly generated documents normally do not receive the exception available to some older documents. DOJ rule summary
Criteria that bite
| WCAG criterion | Chat risk | How UVU should meet it | Required proof |
|---|---|---|---|
| 1.1.1 Non-text Content | Uploaded images, generated diagrams, icons, charts, and formula images may have no text equivalent. | Name functional icons; mark decorative images as decorative; require a useful text description for meaningful visual output; never make an image the only version of code, data, or math. | Screen-reader inspection of uploaded and generated examples; compare the text equivalent with the visual meaning. |
| 1.3.1 Info and Relationships; 1.3.2 Meaningful Sequence | Visual bubbles, tables, headings, citations, and code may become an unstructured stream. | Use native headings, lists, links, tables, <pre><code>, message authors, and chronological DOM order. | Navigate by headings, landmarks, links, tables, and reading order with VoiceOver, NVDA, and JAWS. |
| 1.3.4 Orientation | A phone layout may force portrait mode. | Support portrait and landscape unless orientation is essential. | Rotate real iOS and Android devices during generation, upload, and dialogs. |
| 1.4.3 Contrast; 1.4.11 Non-text Contrast | Tokens, placeholder text, model badges, focus rings, errors, and disabled controls may disappear in themes. | fact threshold: normal text at least 4.5:1, large text at least 3:1, and essential controls/focus indicators at least 3:1. | Measure every theme and state, including forced colors, selected, disabled, error, hover, focus, and streaming. |
| 1.4.4 Resize Text; 1.4.10 Reflow; 1.4.12 Text Spacing | Long answers, code, tables, and the composer may clip or overlap. | fact thresholds: support 200% text resize and reflow at 320 CSS pixels, equivalent to a 1280-pixel viewport at 400% zoom. Confine necessary two-dimensional scrolling to a named table, code, or math region. | Complete a conversation at 400% zoom and with WCAG text-spacing overrides. |
| 2.1.1 Keyboard; 2.1.2 No Keyboard Trap; 2.1.4 Character Key Shortcuts | Hover-only controls, rich editors, upload widgets, menus, and code panes can trap users. | Every action must work without pointer gestures. Provide visible buttons for shortcuts. Permit character-key shortcuts to be disabled or remapped. | Keyboard-only and switch-scanning runs across every state; confirm escape and focus recovery from every overlay. |
| 2.2.1 Timing Adjustable; 2.2.2 Pause, Stop, Hide | Authentication expiry, long uploads, auto-scroll, streaming, animations, and disappearing notices can cause loss. | Warn before expiry, allow extension where security permits, preserve drafts, provide Stop generating, disable auto-scroll when the user moves away, and respect reduced motion. | Run timeout, slow-upload, network-loss, stop, resume, and reduced-motion cases. |
| 2.4.1 Bypass Blocks; 2.4.3 Focus Order; 2.4.6 Headings and Labels; 2.4.7 Focus Visible | Repeated navigation and a growing transcript can make the latest answer hard to reach. Focus can jump to new output. | Add skip links and “Jump to latest response.” Keep focus in the composer after Send. Move focus only after an explicit user action or for a blocking dialog, then restore it. | Record focus before and after every action. Navigate long conversations without a pointer. |
| 2.5.3 Label in Name | Visible labels may not match accessible names, breaking speech control. | Begin the programmatic name with the visible label. | Operate controls by their visible names using Voice Control or another speech-input tool. |
| 3.1.1 Language of Page; 3.1.2 Language of Parts | Spanish text may be pronounced with English rules. Code, names, and quoted material may be misidentified. | Set the interface language on the page and tag answer passages when language changes. Do not tag programming code as natural language. | Read English, Spanish, and code-switched conversations with each screen reader and voice. |
| 3.2.1 On Focus; 3.2.2 On Input | Focusing or selecting a model may unexpectedly submit, navigate, or replace the conversation. | Do not change context on focus. Tell the student before a selection changes context or clears work. | Tab and switch through every control without activating it. |
| 3.3.1 Error Identification; 3.3.2 Labels; 3.3.3 Error Suggestion; 3.3.4 Error Prevention | Upload, sign-in, prompt, export, and connection errors may rely on color, disappear, or lose the draft. | Identify the field or file in text, explain the correction, keep entered content, focus or link to the error summary, and confirm destructive actions. | Test invalid type, oversized file, malware rejection, expired session, rate limit, model failure, and failed export. |
| 4.1.2 Name, Role, Value | Custom buttons, toggles, menus, upload controls, and generation state may be unlabeled. | Prefer native HTML. Expose name, state, selection, expansion, and disabled status. | Inspect the accessibility tree and spoken output; test JAWS with AI-generated labels disabled. |
| 4.1.3 Status Messages | Generation, upload progress, search counts, save confirmation, and failure may be visual-only or overwhelmingly verbose. | Use a restrained status region. Announce “Generating response,” important progress, failure, and “Response ready” without moving focus. | Record exact speech, interruption, duplication, delay, and missing announcements across complete conversations. |
Sources: WCAG 2.1, WAI-ARIA 1.2, and status-message guidance.
WCAG 2.2 is not the named federal baseline, but UVU should design now for focus not obscured, non-drag alternatives, consistent help, accessible authentication, and fact minimum 24 × 24 CSS-pixel targets. This is cheaper than retrofitting them later. WCAG 2.2
Recommended streaming behavior
- Keep one stable semantic transcript. Each complete user or assistant message becomes one navigable item.
- Put visual streaming inside the pending assistant message, but do not recreate the accessible node for every token.
- Use a separate polite, atomic status for “Generating,” “Stopped,” “Connection lost,” and “Response ready.”
- Offer announcement choices: completion only, paragraph chunks, and manual read-aloud. Completion only should be the default.
- Keep composer focus after Send. Do not move focus into the answer automatically.
- Provide “Jump to latest response,” “Read response,” “Stop,” and “Copy” as named buttons.
- If the user navigates upward, stop automatic scrolling. Do not pull them back to the newest token.
- Reserve assertive alerts for urgent, actionable failures. Repeated alerts can erase or interrupt queued screen-reader speech.
- Mark the pending response busy only after cross-screen-reader testing; aria-busy support is not a substitute for an explicit completion message.
What mature products show
Slack documents adjustable message verbosity, replay of recent messages, reduced motion, simplified layouts, keyboard access, and support for major screen readers. GitHub publishes a current Copilot Chat ACR, yet still reports partial support for Name, Role, Value. A Gemini screen-reader user reported receiving only a completion notice and then having to hunt for the answer. A preliminary academic comparison of major generative-AI chats found continuing problems with navigation, labeling, feedback, and prompt handling for blind users. The evidence says mature teams manage announcement load and navigation, but maturity does not equal conformance. Slack accessibility, GitHub Copilot Chat ACR, blind-user AI-chat study.
Audit method
Use WCAG-EM with complete task paths:
- 1. Freeze the exact commit, configuration, themes, translations, identity provider, browser versions, and generated-content renderer.
- 2. Inventory every view and state, including failures and responsive variants.
- 3. Run automated checks as a regression smoke test, not as conformance proof.
- 4. Conduct criterion-by-criterion manual inspection.
- 5. Run complete conversations with the assistive-technology matrix below.
- 6. Test with disabled students using their normal devices and settings.
- 7. Audit generated HTML and every exported artifact separately.
- 8. Fix each blocking barrier and repeat the affected flow.
est test matrix: macOS Safari with VoiceOver; iOS Safari with VoiceOver; Windows with NVDA in Firefox and Chromium; Windows with JAWS in Chrome or Edge with AI Labeler disabled; Android Chrome with TalkBack; one-switch auto-scan; keyboard only; Windows Magnifier or ZoomText; macOS Zoom; 400% browser zoom; forced colors; reduced motion; speech input.
The conversation fixture must include sign-in, model selection, multiline prompt editing, valid and invalid uploads, a scanned document, upload progress, streaming, stop, retry, network loss, history search, rename, deletion confirmation, citations, headings, lists, links, a table, code, math, an image or diagram, read-aloud, Spanish, code-switching, session expiry, and PDF export.
The chosen front door: LibreChat
Current evidence state
| Evidence | Finding |
|---|---|
| Harvard accessibility work | FACT: Harvard states that it partnered with LibreChat to create and review a final Accessibility Conformance Report, or ACR. UNKNOWN: The report’s product version, scope, exceptions, test methods, findings, and remediation state were not publicly retrievable. Harvard Accessibility Summit |
| LibreChat release claim | fact — January 28, 2026: LibreChat v0.8.2 claimed a major accessibility overhaul covering screen readers, keyboard use, focus, and contrast. This is maintainer evidence, not an independent audit. v0.8.2 changelog |
| Open screen-reader report | fact — opened January 6, 2025 and still open at review: issue 5187 describes hover-only controls, unlabeled controls, missing focus movement, and inaccessible image-paste actions. Some reported areas may have changed later. Issue 5187 |
| Later reports | fact — April 2026: an enterprise user reported remaining missing labels and contrast errors in v0.8.4. fact — August 2026: six axe defects were fixed on the development branch with six bounded Chromium scenarios passing afterward. Neither result is a complete AT audit. Discussion 12532, issue 14978, PR 14979 |
| Upload history | FACT: a closed issue documented an attachment control that did not work with Enter or Space and a native input hidden from assistive technology. The issue page does not prove which stable release fixed it. Issue 4087 |
| Mobile risk | FACT: a v0.8.5 report described iOS WebKit killing the page after sidebar and orientation changes. The issue later closed, but it establishes mobile regression risk. Issue 12824 |
| Release drift | fact — September 3, 2026: the visible newest release was v0.8.8-rc2 and marked pre-release. It adds adaptive streams, attachment-only turns, artifacts, diagram export, and downloads—new accessibility surfaces. LibreChat releases |
| Public conformance proof | UNKNOWN: No public, downloadable LibreChat VPAT or ACR was located after searches of the official site, documentation, repository, releases, issues, and exact VPAT/ACR terms. This is not proof that none exists. |
Verdict: LibreChat remains a reasonable engineering starting point, not an approved accessible front door. UVU should use a pinned fork or pinned stable build and make accessibility evidence a release gate.
What UVU must add or prove
- Obtain the Harvard/LibreChat ACR and map it to UVU’s exact commit and enabled features.
- Reproduce the open and historical accessibility reports against the pinned build.
- Backport or verify the v0.8.2 work and PR 14979.
- Replace hover-only and drag-only actions with named controls.
- Implement the stable transcript and restrained status-channel pattern.
- Add component and full-flow regression tests for keyboard, focus, accessible names, state, errors, and reflow.
- Publish an accessibility statement with supported combinations, known limits, an access-barrier route, and a response owner.
- Require the same review when LibreChat, its Markdown renderer, math renderer, upload library, authentication, or theme changes.
Uploaded and generated content
| Content | Accessible handling |
|---|---|
| Upload control | Start with a labeled native file input. Put accepted type and size before the control. Drag-and-drop is only an enhancement. Announce selected file name, count, validation, progress, cancel, removal, failure, and retry. Preserve the prompt if upload fails. |
| Document ingestion | Show name, type, size, page count where available, extraction state, OCR use, and any unsupported or unreadable pages. Never say “document read” when extraction is partial. Provide the original and extracted accessible text. |
| Tables | Use semantic HTML with caption, header cells, and correct scope or headers relationships. Offer CSV. Use a named scrolling region for genuinely wide tables. Never export a table only as an image. W3C tables tutorial |
| Code | Use selectable <pre><code> text with a language label, line-wrap option, and named Copy button. Announce copy success politely. Avoid focus traps in editors and horizontal scrollers. |
| Math | Prefer native MathML. Keep accessible LaTeX or linear text available. Add spoken explanations for complex expressions. Test current VoiceOver, NVDA with MathCAT, and JAWS rather than assuming equivalent support. MathML Core |
| Images and diagrams | Require alt text or a structured text explanation. Mermaid or canvas output must have an equivalent list, table, or relationship description. |
| Generated PDF | Keep HTML as the primary accessible version. A PDF must have a descriptive title and filename, document language, tags, headings, lists, logical reading and tab order, real text, figure alternatives, table structure, descriptive links, contrast, bookmarks where useful, and OCR for scans. Test with PAC or Acrobat plus a real screen reader. Read Out Loud is not a screen reader. Section 508 PDF testing |
Alternatives if LibreChat falls short
- 1. Recommended: Build the smallest accessible web shell around UVU’s local OpenAI-compatible endpoints. Use native HTML, one transcript, one composer, one model control, and limited file support. Result: local models remain available and the surface is auditable. Risk: UVU owns the interface. Ease of undo: high because model APIs remain unchanged.
- 2. Use a procured cloud chat with a current, scoped ACR for the accessible lane. Result: faster front-door proof. Risks: different privacy, retention, cost, model-routing, and FERPA boundaries. Ease of undo: medium.
- 3. Switch to another open-source chat such as Open WebUI. Result: wider feature set. Risk: public reports show continuing transcript, labeling, and keyboard issues, so this is not a proven lower-risk swap. Ease of undo: medium. Open WebUI accessibility discussion, current transcript issue.
Voice and reading support
On-device speech-to-text
| Option | Evidence | Limits and recommendation |
|---|---|---|
| Apple SpeechAnalyzer / SpeechTranscriber | FACT: Apple describes on-device, low-latency streaming, meeting, distant-speech, and long-form transcription with downloadable assets. WWDC 2025 session | UNKNOWN: Apple publishes no comparable English/Spanish accent word-error rate or end-to-end latency. Use only after checking locale assets and on-device support at runtime. |
| Apple SFSpeechRecognizer | FACT: APIs expose whether on-device recognition is supported and can require it. Apple warns local recognition may be less accurate. | Do not call Apple Speech private merely because it runs on a Mac. Fail closed or disclose when the selected locale cannot stay local. |
| Whisper large-class model | FACT: Whisper large-v2 reported Spanish word-error rates of 4.2% on Multilingual LibriSpeech, 5.6% on Common Voice, 8.2% on VoxPopuli, and 3.0% on FLEURS. English results on the same datasets were 6.2%, 9.4%, 7.0%, and 4.2%. Whisper paper | These are dataset results, not UVU guarantees. Accented English reached fact 18.6% on the cited VoxPopuli evaluation. Course names, noise, disabilities, and code-switching require local tests. |
| Smaller Whisper models | FACT: On Multilingual LibriSpeech, tiny produced 15.7% English and 19.2% Spanish WER versus 6.2% and 4.2% for large-v2. | Smaller models may improve speed while materially reducing accuracy. Do not select them by latency alone. |
| WhisperKit | FACT: A published M3 Max test reported encoder latency improving from 612 ms to 218 ms after compression; tested WER rose from 1.93% to 2.25% on LibriSpeech-clean and from 11.55% to 12.85% on Earnings22. Streaming TIMIT mean interim latency was about 0.45 seconds and confirmed output about 1.7 seconds. WhisperKit paper | The evaluation did not prove Spanish parity or the same result on UVU’s target Macs. Interim words can change, so students must review before sending. |
| whisper.cpp | FACT: The project supports Metal and Core ML and reports more than 3× encoder speedup from Core ML compared with CPU-only operation. whisper.cpp | It is an implementation claim, not an accuracy or end-to-end latency study. |
Voice product rules
- Begin with push-to-talk dictation, not an always-listening conversational assistant.
- Show live partial text, clearly mark it as changing, and require review before Send.
- Provide Stop, Cancel, Undo, microphone state, input-language selection, and keyboard equivalents.
- Default to local processing. If local processing is unavailable, do not silently send audio to a cloud service.
- Delete transient audio after transcription unless the student explicitly saves it under an approved purpose.
- Never treat a transcript as authoritative. FACT: a study of Whisper transcripts found hallucinations in 1.7% of clips from participants with aphasia and 1.2% of control clips; fact 38% of hallucinated transcripts fell into harmful categories. The study was American English and does not establish current local-model rates. Careless Whisper
- est launch test: record at least 40 consenting speakers across English, Spanish, accents, speech disabilities, microphone types, quiet rooms, and ordinary campus noise. The count is a coverage target, not a prevalence sample.
- Report word-error rate, proper-name error rate, harmful insertions, language switches, interim corrections, time to first partial, time to stable text, and thermal/battery behavior by subgroup.
Text-to-speech and reading modes
Apple AVSpeechSynthesizer and macOS Spoken Content can operate on-device, select voices, adjust rate and pitch, pause, resume, stop, and highlight words or sentences. Additional voices may require downloads. UNKNOWN: Apple publishes no current Spanish-versus-English intelligibility or naturalness benchmark. Apple speech synthesis, macOS Spoken Content.
Provide:
- Read answer, pause, resume, previous paragraph, next paragraph, stop, speed, and voice controls.
- Synchronized highlighting that can be disabled.
- Standard, short answer, plain language, and step-by-step views.
- Original answer and source text beside every simplified version.
- A warning when simplification may remove detail.
- Downloadable text rather than audio-only export.
What helps, and what does not
- FACT: A meta-analysis of 22 studies with pooled fact n=2,942 found a small positive read-aloud effect, fact d=.35, fact 95% CI .14–.56; publication-bias adjustment reduced it to fact d=.24. Heterogeneity was very high. Read-aloud is worth offering, not promising. Text-to-speech meta-analysis
- FACT: A study with 170 children with dyslexia found no reading advantage for Dyslexie over Arial; most preferred Arial. A second sample contained fact 102 children with dyslexia and fact 45 controls and again found no benefit. Kuster et al.
- FACT: In another study, a roughly fact 7% Dyslexie reading-speed advantage disappeared after Arial’s spacing was matched. Adjustable spacing is more defensible than a branded font. Marinus et al.
- FACT: Extra-large spacing reduced errors and improved speed in a crossover study of fact 74 Italian and French children with dyslexia. This does not justify an extreme default for skilled adult readers. Zorzi et al.
- FACT: A small Spanish study found better comprehension for shorter words among readers with dyslexia, but automatic simplification studies have not shown reliable objective-comprehension improvement. Preserve the original and make simplification optional. Rello readability study, simplification study
- Evidence for decorative “Easy Read” pictures is mixed and can include confusion. Use visuals only when they add meaning.
Neurodivergent users
The evidence is stronger for calm, predictable, adjustable design than for diagnosis-specific AI modes. The service must not infer ADHD, autism, dyslexia, or anxiety from behavior.
| Need | Evidence-based design | Evidence boundary |
|---|---|---|
| Pacing | One task or choice at a time; Stop generation; visible progress; no forced typing animation; user-selected answer length. | W3C cognitive guidance is expert consensus, not a clinical treatment trial. |
| Structure | Descriptive headings, short paragraphs, numbered steps, summary first, stable control positions, and a visible conversation outline. | Helpful across conditions; not every user wants reduced detail. |
| Predictability | Explain waits and model changes, warn before destructive actions, keep navigation consistent, and avoid surprise context changes. | FACT: anxiety is strongly associated with intolerance of uncertainty, but interface predictability has not been proved to treat anxiety. |
| Interruption | Default notifications and reminders off; no streaks, guilt, autoplay, or unrelated suggestions; respect reduced motion. | FACT: in a randomized undergraduate study with fact n=221, active alerts increased reported inattention and hyperactivity with effects around fact d=.44–.45. |
| Memory aids | Autosave drafts, pin steps, show “where you left off,” keep a checklist, and allow optional reminders with fading. | FACT: external reminders can improve immediate prospective-memory performance but may reduce independent practice effects after removal. |
| Focus | Low-distraction view, hide optional sidebars, keep one primary action, and prevent the page from jumping during generation. | An adult ADHD experiment found larger distractor costs, but its laboratory manipulation does not justify making the interface cognitively demanding. |
| Control | Standard, brief, plain-language, and step-by-step modes; adjustable spacing and read-aloud; easy return to the original. | Diagnosis-based presets risk stereotyping and disclosure. |
| Human exit | Clear link to tutoring, Accessibility Services, Service Desk, and other appropriate people without forcing more AI dialogue. | AI-specific dependency evidence for disabled college students remains thin. |
Sources: W3C cognitive and learning-disability guidance, phone-notification experiment, prospective-memory reminder trial, and anxiety meta-analysis.
Dependency and distraction risks
- Endless follow-up suggestions can turn a short task into an unbounded session.
- Confident simplification can remove conditions or exceptions.
- Reminders can become interruption pressure.
- Personalization can expose or falsely infer disability.
- A student may rely on the assistant instead of practicing planning, reading, or help-seeking.
- Saved accessibility preferences can become sensitive profile data.
- A calm tone can make weak advice seem more trustworthy.
Mitigations include optional session goals, visible elapsed conversation length without pressure, “finish this task” mode, reminders default-off, periodic source checks, easy export to a human helper, and no disability inference.
UVU’s existing practice
FACT: UVU Accessibility Services already offers individualized accommodations, learning specialists, organization and test-taking support, peer mentoring, alternative formats, text-to-speech, speech-to-text, screen readers, enlargement, Braille, captioning, transcription, and distraction-reduced testing. Its Accessible Technology Center asks students to bring the devices they normally use. Accessibility Services, programs, testing services.
The AI service should complement those practices, not require disclosure or replace the interactive accommodation process. FACT: UVU says information provided to Accessibility Services is protected by FERPA. Accommodation rights
Ethical testing with disabled students
- Ask UVU’s IRB for a written determination before recruitment or recording. UVU states that pilot and feasibility studies can require review and that the IRB decides exemption.
- Accessibility Services may distribute a neutral invitation, but should not disclose its accommodation roster to the product team.
- Recruit by access method and functional need, not by demanding a diagnosis list.
- Make participation voluntary, unrelated to grades, employment, service access, or accommodations.
- Provide consent in plain English and Spanish, with text, audio, and asynchronous options.
- Permit a support person, camera-off participation, breaks, and the participant’s own assistive technology.
- Pay for preparation, testing, and follow-up. Prorate payment if someone stops early.
- Treat an inaccessible session as a product failure, not a participant failure.
- Do not publish identifiable quotations, audio, disability details, or chat content without explicit consent.
est coverage plan: 18–24 paid students across screen reader, low vision or magnification, motor or switch access, deaf or hard-of-hearing access, reading or cognitive access, neurodivergent workflows, Spanish or code-switching, and mobile or weak-network use. Overlap is welcome; this is a barrier-discovery sample, not a statistical prevalence sample.
Spanish and multilingual service
UVU population evidence
- fact — fall 2025: UVU reports 48,669 students, 14% Hispanic, and 40% first-generation on its current Key Indicators page. UVU Key Indicators
- FACT: The Common Data Set reports Hispanic/Latino counts of 821 of 5,104 degree-seeking first-year students, 4,054 of 28,625 degree-seeking undergraduates, and 6,660 of 47,518 total undergraduates. The calculated shares are fact 16.09%, fact 14.16%, and fact 14.02%. UVU Common Data Set
- FACT: UVU’s First-Generation Student Success Center reports 41% and uses a stated definition tied to parents or guardians lacking a United States bachelor’s degree. The difference from 40% likely reflects source timing or cohort. Do not average them. First-generation page
- fact — fall 2019 survey, n=1,028: 6.1% selected Spanish as a native language. fact — separate spring 2019 survey, n=939: 21.4% said they could speak Spanish. These old, self-selected surveys measure different things.
- UNKNOWN: No current public administrative census of preferred, home, or instructional language was located. Hispanic identity must not be used as a Spanish-preference proxy.
Candidate-model language evidence
| Model | Published result | What it establishes | Verdict |
|---|---|---|---|
| Qwen3.8-27B | fact — Spanish Arena snapshot September 2, 2026: score 1410 ±44 from 187 votes; rank 113 with a broad interval of 20–191. Spanish Arena | Pairwise preference with wide uncertainty, not curricular correctness. | unknown Spanish tutoring quality. |
| Qwen3.6-35B-A3B | FACT: maker reports SWE-bench Multilingual 67.2. Model card | Software repositories in multiple programming languages; it does not measure Spanish instruction. | unknown Spanish quality. |
| GLM-5.3-Flash | No exact Spanish result was located in its official card. Model card | Published agent and coding results do not establish natural-language Spanish quality. | unknown. |
| Mistral Small 4 | FACT: maker says Spanish is supported among dozens of languages but publishes no Spanish score in the card. Model card | Capability claim, not measured parity. | unknown. |
| gpt-oss-120B | fact — MMMLU Spanish: 80.6% at low, 84.6% at medium, and 85.9% at high reasoning. fact — fourteen-language averages: 74.1%, 79.3%, and 81.3%. fact — Spanish Arena: 1365 ±21 from 854 votes, rank 173. Model card, Spanish Arena | MMMLU measures translated multiple-choice knowledge; Arena measures preference. Neither measures UVU tutoring. | Best published Spanish evidence, but selection remains conditional. |
The apparent Qwen3.8-versus-gpt-oss contradiction is real but not actionable: Qwen3.8 leads on a small, volatile preference sample; gpt-oss has the stronger documented knowledge benchmark. The measures are not comparable.
Quantization is another gap. FACT: gpt-oss uses native post-training with 4.25-bit MXFP4 expert weights, so its official result should not be described as full precision. UNKNOWN: no Spanish evaluation was located for the Apple-quantized builds of the other candidates.
Bilingual tutoring and code-switching
- FACT: A 2026 meta-analysis of multilingual versus target-language-only teaching included 24 studies, 30 samples, and fact n=2,138; it reported between-group fact d=.49 and within-group fact d=1.63. It was language instruction, not AI tutoring. Zhang and Brown
- FACT: A peer-tutoring meta-analysis covering 14 studies reported fact g=.58, fact 95% CI .22–.94; samples were mainly Spanish-speaking and school-aged. Higher-education transfer is uncertain. Romero et al.
- FACT: A randomized online-science study with 50 Spanish-speaking eighth-grade students found benefits from bilingual support on several immediate and delayed measures. The sample was small and not college-level. Clark et al.
- FACT: A 2026 code-switching benchmark found that inserting non-English material into English consistently reduced comprehension and reasoning accuracy, while inserting English into non-English contexts often helped. Prompt-only mitigation was inconsistent. Lost in the Mix
- Evidence that AI code-switching improves learning is thin. A 2024 tutoring study used fact 400 simulated Chinese-English and Korean-English dialogues rather than Spanish learners or measured learning outcomes. SIGDIAL study
Language design
Keep three controls separate:
- Interface language: menus, buttons, errors, consent, privacy, and help.
- Content language: the source document or course material.
- Reply language: English, Spanish, both, or follow the current prompt.
Use human-reviewed Spanish for identity, privacy, safety, accessibility, and campus-service text. Keep canonical English course terms visible with optional Spanish definitions. If a student code-switches, preserve the mix or ask a short non-blocking language question; do not force translation. Never infer language from name, ethnicity, or browser locale alone.
Required UVU language evaluation
est benchmark: 300 paired tasks, built as five UVU-relevant domains with 60 tasks per domain: course explanation, writing support, quantitative reasoning, coding or technical help, and campus-service/safety information.
For each exact deployed quantization and prompt:
- Run matched English, Spanish, and natural code-switched variants.
- Use est two independent bilingual graders per answer, with adjudication.
- Score correctness, completeness, instruction following, terminology, source faithfulness, language fidelity, cultural assumptions, unsupported confidence, and safety.
- Blind graders to model identity.
- Separate first-answer quality from multi-turn recovery.
- Treat Spanish-only critical errors as blockers even if the English average is strong.
- est policy threshold: no more than a five-percentage-point paired correctness gap between English and Spanish for ordinary tutoring, plus no unresolved critical gap in campus, safety, financial, disability, or academic-policy tasks. This is a proposed governance threshold, not a published fact.
Mobile and low-bandwidth access
What is known
- UNKNOWN: No current UVU-specific public measure of phone-only, primary-phone, or home-broadband access was located.
- fact — national 2025 survey, n=5,022: 16% of United States adults were smartphone-dependent, including 27% of adults aged 18–29, 28% of Hispanic adults, and 34% of adults with household income below $30,000. These are national, not UVU, figures. Pew
- fact — one Hispanic-Serving Institution study, n=2,188: optimal smartphone-plus-computer access was 72% among first-generation students and 85% among non-first-generation students; unstable internet was reported by 30% and 28%. This supports risk planning, not a UVU estimate. Digital inequality study
- fact — multi-campus sample, n=2,913: 7% reported losing internet for inability to pay, 28% hit a mobile-data cap, and 18% lacked a working laptop for at least ten nonconsecutive days during the prior year. The study occurred during the pandemic. PLOS ONE
- FACT: UVU offers encrypted Eduroam on major desktop and mobile platforms. Its open Wolverine-WiFi does not fully protect device-to-access-point traffic.
- FACT: Guest Wi-Fi uses a captive portal and grants eight hours before re-registration.
- FACT: UVU lends laptops and hotspots free to current-semester students on a first-come basis. UNKNOWN: Public inventory and unmet demand were not located.
- FACT: UVU’s Service Desk has a Spanish chat link, but continuous ticket intake is not continuous staffed support.
Sources: UVU Eduroam, Wolverine-WiFi, guest instructions, equipment checkout, and Service Desk.
Phone and weak-network requirements
- Complete every core task at 320 CSS pixels in portrait and landscape.
- Test iOS Safari with VoiceOver and Android Chrome with TalkBack on real devices.
- Keep all primary actions visible without hover.
- Use a text-first page and lazy-load optional media.
- Do not autoplay audio, video, diagrams, or long read-aloud output.
- Show file size before upload and permit cancel, resume, and retry.
- Save unsent drafts locally with a privacy warning and clear-delete control.
- Make request retries idempotent so a reconnect does not submit twice.
- Expose Offline, Stale, Syncing, Synced, Upload paused, and Generation lost states in text and to assistive technology.
- Cache the application shell, help, and authorized static material. Do not claim that model answers are available offline when the inference server is unreachable.
- Preserve the completed portion of an answer after a dropped stream and clearly mark whether generation can resume.
- Provide accessible HTML or text downloads that are smaller than PDF when possible.
- UNKNOWN: Data cost per conversation remains unmeasured. Instrument payload bytes without collecting prompt content, publish upload size before transfer, and offer a Wi-Fi-only upload option.
Plan
Prioritized backlog
| Rank | Pilot work | Evidence required |
|---|---|---|
| P0-A | Freeze a stable LibreChat commit or select the small accessible shell. Name one accessibility owner. | Exact commit, configuration, ownership, and supported-browser record. |
| P0-B | Implement the stable transcript, restrained status region, stable composer focus, Stop, jump/read-latest, and no forced auto-scroll. | Complete VoiceOver, NVDA, JAWS, TalkBack, keyboard, and switch runs. |
| P0-C | Fix names, states, focus, contrast, reflow, timeouts, errors, reduced motion, and language tagging. | Manual WCAG 2.1 A/AA audit plus automated regression checks. |
| P0-D | Ship English and human-reviewed Spanish interface, privacy, errors, safety, and help. | Bilingual review and language-tag testing. |
| P0-E | Run the exact deployed models through the paired UVU English/Spanish evaluation. | Versioned prompts, answers, rubrics, grader agreement, adjudication, and release decision. |
| P0-F | Recruit and pay disabled, bilingual, and low-bandwidth student testers after UVU’s IRB determination. | Consent, compensation record, de-identified findings, fixes, and retests. |
| P1-A | Add accessible upload state and extracted accessible HTML/text. | Valid, invalid, oversized, scanned, partial-extraction, retry, and cancellation tests. |
| P1-B | Add semantic tables, code, citations, and math. | AT and 400%-zoom conversation fixture. |
| P1-C | Add on-device dictation and system read-aloud. | Locale, privacy, WER, latency, harmful-insertion, and control tests. |
| P1-D | Add PDF export only after tagged-PDF proof; otherwise offer HTML, DOCX, or text. | PDF structure inspection and screen-reader task completion. |
| P2-A | Add broader multilingual support only after separate language evaluations and human-reviewed interface text. | Per-language evidence record. |
| P2-B | Maintain an ACR and accessibility statement for each supported release. | Updated scope, exceptions, methods, known limits, and remediation dates. |
Testers and cost
| Item | Planning basis | Cost |
|---|---|---|
| Accessibility remediation and test automation | est 600–1,200 staff hours, based on six major workstreams at roughly 100–200 hours each. | unknown dollars: UVU loaded labor rates and current LibreChat condition were not supplied. |
| Independent expert audit | Manual WCAG and AT audit of production-like flows plus retest. | est $20,000–$50,000 placeholder; not a quote. Scope and procurement rates can change it materially. |
| Disabled and bilingual student testing | est 18–24 students × two paid hours × $40/hour. | est $1,440–$1,920, plus unknown interpreter, transport, attendant, equipment, and accommodation costs. |
| Spanish model evaluation | est 300 tasks × two graders × ten minutes = 100 grading hours, plus est 40 hours adjudication and setup. | est $4,900 at an assumed est $35/hour; practical budget est $5,000–$9,000. |
| Speech evaluation | est 40 participants, recruitment, transcription references, and analysis. | est $4,000–$12,000; basis is paid participation plus specialist review, not a vendor quote. |
| Scale operations | Release review, regression maintenance, issue triage, student feedback, and annual independent check. | est 0.5–1.0 accessibility FTE plus est $15,000–$40,000 external review per year. |
Launch-gate checklist
Launch only when all applicable items are PASS:
- Exact production commit, configuration, themes, translations, renderer, authentication, and model routes are frozen.
- WCAG 2.1 Level A and AA manual audit covers complete processes.
- No unresolved defect blocks independent sign-in, prompt entry, response reading, Stop, error recovery, history, or sign-out.
- VoiceOver, NVDA, JAWS, TalkBack, keyboard, switch, magnification, forced colors, text spacing, and reduced-motion flows pass.
- Status announcements are neither missing nor duplicated or continuously interrupting.
- Focus never moves merely because tokens arrive.
- Every visible action has a matching accessible name and exposed state.
- The service reflows at 320 CSS pixels and remains operable at 400% zoom.
- English and Spanish interface content has human review and correct page/part language metadata.
- The exact deployed model and quantization pass the UVU paired language gate.
- Upload failure never loses the prompt; extraction limits are disclosed.
- Tables, code, math, diagrams, citations, and links work in the complete AT conversation.
- PDF export is disabled unless the produced file passes its independent audit.
- Dictation is disabled unless locale, local-processing, transcript-review, accuracy, latency, and deletion gates pass.
- Weak-network retry does not submit twice or hide whether a message was received.
- Paid disabled-student testing has been completed and blocking barriers have been fixed and retested.
- Accessibility statement, barrier-report route, support owner, and rollback path are live.
- Each later release repeats affected checks before deployment.
Cross-domain learnings
- Air-traffic and operations consoles: status is separated from the operator’s working focus. Chat should announce state without stealing focus.
- Captioning and speech recognition: partial output is visibly useful but unstable. Final text should be confirmed before it becomes an instruction, record, or submitted prompt.
- Localization systems: interface locale, source-content language, and output language are separate controls. Human-approved terminology is retained for high-risk text.
- Mobile banking and resilient forms: autosave, idempotent retry, clear receipts, and explicit pending states prevent duplicate or lost actions.
- Document publishing: structured source content produces more reliable accessible outputs than repairing PDFs after generation.
- Participatory disability research: standards inspection and lived-experience testing answer different questions. Disabled users are paid experts, not free final-stage validators.
- Learning support: optional scaffolds can help, but permanent prompts and reminders may reduce independent practice. Controls should be adjustable and fadeable.
- Safety-critical alarms: frequent alerts become noise. Screen-reader announcements need priority, acknowledgement, and restraint.
Devil’s advocate
The strongest case against this recommendation is to avoid a custom campus front door altogether.
LibreChat’s public conformance proof is incomplete, dynamic chat is unusually difficult to audit, accessible document conversion is expensive, Spanish benchmark coverage is poor, and voice creates new privacy and accuracy risks. UVU already provides Read&Write, screen readers, alternative formats, accessible testing, learning specialists, and other human support. A custom interface could consume scarce funds while duplicating mature cloud tools and still failing students at the edges.
Under that view, UVU would procure one established cloud chat with a current scoped ACR, restrict it to low-risk content, retain existing accessibility services, and delay local-model access until an accessible interface exists.
That is a credible alternative. Its weakness is that it changes the plan’s privacy, model-control, cost, and local-compute goals, while a vendor ACR still does not prove the complete UVU service.
What would change the recommendation
- A current, independently reviewed LibreChat ACR covering UVU’s exact release and complete workflows could reduce the need for a separate shell.
- Reproduction showing that LibreChat’s open screen-reader reports are fixed in a stable release could lower remediation scope.
- A different open-source front door passing the complete UVU AT audit could replace LibreChat.
- Current UVU device and language research showing little phone, Spanish, or weak-network need could change priority, though not ADA obligations.
- Exact Apple-locale tests showing strong private English and Spanish speech performance could move dictation earlier.
- Exact deployed-model testing showing reliable Spanish parity could permit broader Spanish launch.
- A federal rule change could alter the compliance date or standard; it would not remove existing equal-access duties.
- Insufficient staffing or budget should narrow the pilot to accessible text chat, not lower the gate.
What the plan should change
- 1. Replace “WCAG 2.1 AA by launch” with the complete process, AT matrix, tester, evidence, and release-gate requirements above.
- 2. Change LibreChat from “chosen front door” to “conditional implementation candidate pending exact-build proof.”
- 3. Make accessible text chat the first release; gate uploads, math, diagrams, PDF, and voice separately.
- 4. Add independent interface language, content language, and reply-language controls.
- 5. Add the paired UVU English/Spanish benchmark for every exact deployed model and quantization.
- 6. Fund paid disabled and bilingual student participation and obtain UVU’s IRB determination before recruitment.
- 7. Add mobile, captive-portal, weak-network, interrupted-upload, and data-use tests.
- 8. Require accessible HTML as the primary generated document; treat PDF as an additional audited artifact.
- 9. Assign an accessibility owner and maintain a release-specific ACR, known-issues statement, and barrier-response process.
- 10. Reserve the metered frontier pool as an accessible quality fallback only after its privacy, language, and front-door path passes the same gates.
Source log
Source numbers are identifiers, not measured claims.
- 1. fact date — April 8, 2024; updated for April 20, 2026 IFR: DOJ Title II web and mobile rule fact sheet.
- 2. fact date — accessed September 3, 2026: DOJ first steps for public entities.
- 3. fact date — June 5, 2018: W3C WCAG 2.1.
- 4. fact date — December 12, 2024: W3C WCAG 2.2.
- 5. fact date — June 6, 2023: WAI-ARIA 1.2.
- 6. fact date — May 11, 2026 update: Understanding Status Messages.
- 7. fact date — July 23, 2026: WCAG Evaluation Methodology 2.0.
- 8. fact date — accessed September 3, 2026: W3C guidance on involving disabled users.
- 9. fact date — May 24, 2022 update: UK Government assistive-technology testing.
- 10. fact date — October 29, 2024 update: UK Government accessibility testing.
- 11. fact date — July 2025 update: U.S. ICT Testing Baseline.
- 12. fact date — June 4, 2026 event: Harvard Accessibility Summit LibreChat ACR session.
- 13. fact date — January 28, 2026: LibreChat v0.8.2 changelog.
- 14. fact date — September 3, 2026 review: LibreChat releases.
- 15. fact date — January 6, 2025: LibreChat issue 5187.
- 16. fact date — November 25, 2025 merge: LibreChat PR 10607.
- 17. fact date — April 2, 2026: LibreChat discussion 12532.
- 18. fact date — August 18, 2026: LibreChat issue 14978.
- 19. fact date — August 20, 2026 merge: LibreChat PR 14979.
- 20. fact date — September 17, 2024: LibreChat upload issue 4087.
- 21. fact date — April 25, 2026: LibreChat mobile issue 12824.
- 22. fact date — August 18, 2026 update: U.S. Web Design System file input.
- 23. fact date — February 16, 2023 update: W3C tables tutorial.
- 24. fact date — June 24, 2025: MathML Core publication history.
- 25. fact date — March 2026 review: U.S. Section 508 PDF testing.
- 26. fact date — December 11, 2025: WCAG2ICT document guidance.
- 27. fact date — July 2, 2026 update: GitHub Copilot Chat ACR.
- 28. fact date — accessed September 3, 2026: Slack accessibility.
- 29. fact date — 2025 publication: Preliminary evaluation of generative-AI tools for blind users.
- 30. fact date — August 29, 2026: Open WebUI transcript issue 29211.
- 31. fact date — July 4, 2026: Open WebUI university accessibility remediation discussion.
- 32. fact date — 2023 publication: Whisper paper.
- 33. fact date — 2025 publication: WhisperKit paper.
- 34. fact date — accessed September 3, 2026: whisper.cpp repository.
- 35. fact date — WWDC 2025: Apple SpeechAnalyzer session.
- 36. fact date — accessed September 3, 2026: Apple on-device recognition support.
- 37. fact date — accessed September 3, 2026: Apple required on-device recognition.
- 38. fact date — accessed September 3, 2026: Apple speech synthesis.
- 39. fact date — 2024 publication: Careless Whisper speech hallucination study.
- 40. fact date — 2017 online publication: Text-to-speech meta-analysis.
- 41. fact date — December 4, 2017: Dyslexie font study.
- 42. fact date — June 4, 2012: Extra-large spacing study.
- 43. fact date — May 19, 2016: Font-spacing comparison.
- 44. fact date — 2013 publication: Spanish readability study.
- 45. fact date — 2013 publication: Dyslexia text-simplification study.
- 46. fact date — April 29, 2021: W3C cognitive accessibility guidance.
- 47. fact date — April 3, 2019: AASPIRE participatory research guidelines.
- 48. fact date — February 19, 2025: Accessible consent guidance.
- 49. fact date — 2013 publication: Complex prospective memory in adults with ADHD.
- 50. fact date — 2014 publication: Distraction and adult ADHD experiment.
- 51. fact date — 2016 publication: Phone-notification experiment.
- 52. fact date — December 27, 2021: External-reminder randomized trial.
- 53. fact date — October 31, 2019: Intolerance-of-uncertainty meta-analysis.
- 54. fact date — accessed September 3, 2026: UVU Accessibility Services.
- 55. fact date — accessed September 3, 2026: UVU Accessibility Services programs.
- 56. fact date — accessed September 3, 2026: UVU accommodations and FERPA statement.
- 57. fact date — accessed September 3, 2026: UVU IRB application process.
- 58. fact date — September 30, 2019: HHS participant-payment guidance.
- 59. fact date — fall 2025 reporting: UVU Common Data Set.
- 60. fact date — fall 2025 data; accessed September 3, 2026: UVU Key Indicators.
- 61. fact date — accessed September 3, 2026: UVU First-Generation Student Success Center.
- 62. fact date — fall 2019: UVU Student Opinion Survey.
- 63. fact date — August 1, 2023: UVU Wolverine-WiFi instructions.
- 64. fact date — accessed September 3, 2026: UVU Eduroam.
- 65. fact date — accessed September 3, 2026: UVU guest Wi-Fi instructions.
- 66. fact date — accessed September 3, 2026: UVU laptop and hotspot checkout.
- 67. fact date — August 5, 2025: gpt-oss model card.
- 68. fact date — August 2026 release; accessed September 3, 2026: Qwen3.8-27B model card.
- 69. fact date — April 2026 release: Qwen3.6-35B-A3B model card.
- 70. fact date — September 2, 2026: GLM-5.3-Flash model card.
- 71. fact date — March 16, 2026: Mistral Small 4 announcement and model card.
- 72. fact date — September 2, 2026 snapshot: Spanish Arena leaderboard.
- 73. fact date — July 2026: Lost in the Mix code-switching evaluation.
- 74. fact date — September 2024: Conversational tutoring code-switching study.
- 75. fact date — July 25, 2026: Multilingual versus monolingual teaching meta-analysis.
- 76. fact date — January 8, 2026: Pew digital-divide report.
- 77. fact date — February 8, 2022: College-student digital inequality study.
- 78. fact date — February 10, 2021: Multi-campus connectivity study.
NOT_RUN
- NOT_RUN: No UVU or vendor contact.
- NOT_RUN: No live UVU or LibreChat deployment inspection; exact build and configuration remain unknown.
- NOT_RUN: No VoiceOver, NVDA, JAWS, TalkBack, switch, magnification, keyboard, or PDF task run.
- NOT_RUN: No accessibility defect reproduction against LibreChat.
- NOT_RUN: No download or review of the Harvard/LibreChat ACR because no public copy was located.
- NOT_RUN: No Apple-Silicon speech benchmark, local model run, quantization comparison, or Spanish tutoring benchmark.
- NOT_RUN: No UVU student recruitment, recording, survey, compensation, or IRB submission.
- NOT_RUN: No campus Wi-Fi coverage, latency, outage-history, hotspot-inventory, phone-share, or data-cost measurement.
- NOT_RUN: No deployment, publication, procurement, account change, or external send.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/46-accessibility-language.md in the research pack.
Roadmap and timing
Appendix Z in five lines
- QuestionShould the university buy machines now or wait for the next chip, and how should later purchases be staged?
- AnswerBuy four machines only after one delivered unit passes every test, buy later waves only when use proves the need, and do not commit to 26 now.
- Deciding numbersabout 15 months of delayest390 fleet device-monthsfactabout $8,916 all-inest
- What the plan doesUse a delivered-machine gate, buy in three waves, retest model fit, compare local and cloud costs, and review replacement in year three.
- Still unknownThe new machines have not been tested, the 512-gigabyte price and state compute share are UNKNOWN, and live security, access, and recovery checks are NOT_RUN.
A finance committee will ask "why not wait for the next chip?" and "what if the models outgrow the boxes?" This appendix answers with dated evidence: every Mac mini and Mac Studio chip since 2020 with bandwidth, memory, and launch price; Apple's stated AI direction and the credible reporting on what comes next; the best open model that fit each memory tier, quarter by quarter; the cost of waiting in device-months and stream-months; an options-style purchase plan in gated waves; what could break the trajectory; and a dated 24-month watchlist with the decision each event triggers. Research cutoff September 3, 2026; 150-plus dated sources. fact release facts and arithmetic · est thresholds and wave sizes · unknown rumors, so labeled.
What this changes in the plan
1. "Buy after September 22" becomes a delivered-unit acceptance gate. Shipment is availability, not proof: one delivered machine must load the selected model, complete the 1k/8k/32k-token tests, sustain at least four simultaneous streams at the 10-token-per-second floor, pass batch 1/2/4/8, cold-load, peak-memory, and 24-hour soak tests, plus identity, logging, isolation, accessibility, and recovery checks. 2. Buy in waves, not 26 at once: wave one, four minis (about $8,916 all-in) for a 30-faculty cohort; wave two, up to eight more, only after half the pilot participants return at their next natural task within 30 days, measured demand uses at least 60% of capacity for four teaching weeks, support and incidents stay inside the staffed plan, and no model release has changed the best tier; wave three, up to the planned 26, after spring results, observed queue times, and a refreshed local-versus-cloud comparison. Do not reserve 26 for symmetry. 3. Waiting has a price. The rumored next Pro-class chip implies about 15 months of delay, forgoing 60 validation device-months and 390 fleet device-months; a hypothetical machine 25% faster repays no more than 9.6 months of waiting inside a 48-month service life. Delay only if a must-have requirement fails, Apple officially announces within 90 days a machine with the required memory or 35% more measured throughput for no more than 20% more cost, a dated grant award within 60 days would cover a material share, or accessibility or security testing blocks the service. 4. Memory rule: buy for models that create useful service now; the biggest flagships already exceed 512 GB and belong in the metered pool. 5. The 512 GB line stays unpriced until Apple posts the configuration and at least two approved research workloads need a model that cannot fit 256 GB — and only at a price no higher than about six all-in minis. 6. Year three is a review with real resale quotes; year four a refresh only where measured gains justify it; older machines become development, spare, or low-priority capacity. 7. Dual-runtime acceptance (MLX and llama.cpp interchangeable behind one service contract) and a quarterly open-model checkpoint on one benchmark version. The 24-month calendar of Apple events, model checkpoints, grant deadlines, state milestones, and the April 26, 2027 accessibility deadline is at the end. est
Prepared independently by Utlyze. Not a publication of, endorsed by, or affiliated with Utah Valley University. As of September 3, 2026, MDT.
Evidence labels: fact means a dated source or arithmetic from sourced figures. est means a planning estimate whose basis is stated. unknown means the needed evidence was not public or could not be verified. Source-log numbers are reference identifiers, not quantitative claims.
Executive verdict — ten lines
- 1. Buy the four-machine validation set after the September 22, 2026 shipment date, provided one delivered machine passes the full service test.
- 2. Do not buy the full 26-machine fleet at once; preserve the option to stop, expand, or change tiers after real use data.
- 3. Waiting for the rumored next Pro-class chip would likely mean about 15 months of delay est: named reporting says late 2027.
- 4. That delay would forgo 60 validation device-months and 390 fleet device-months fact: arithmetic.
- 5. It would also forgo 240–480 validation stream-months and 1,560–3,120 fleet stream-months est: four-plan versus eight-tested streams per machine.
- 6. The historical Mac mini interval is 22.4 months at the median fact, but the most recent Pro memory ceiling did not rise beyond 64 GB fact.
- 7. A hypothetical 25% faster next machine needs about 9.6 months of waiting to repay itself inside a 48-month service horizon est inputs; FACT arithmetic; a 15-month wait is too long.
- 8. Open models are improving quickly, but the largest flagships already exceed 512 GB; UVU should buy useful service tiers, not try to own every frontier model.
- 9. Review resale and workload evidence in year three; refresh in year four only where measured service gains justify it.
- 10. The plan should replace its fixed four-stream assumption with delivered tests and make adoption, accessibility, price, and model releases explicit expansion gates.
# 1. Apple silicon cadence
Mac mini and Mac Studio releases, 2020–2026
Dates below are availability dates. Prices are original US retail starting prices, not UVU’s configured or education prices.
| Machine and chip | Available | Unified-memory bandwidth | Maximum memory | Launch price | Evidence |
|---|---|---|---|---|---|
| Mac mini, M1 | Nov. 17, 2020 fact | 68.25 GB/s fact-secondary; Apple did not publish it on the launch page | 16 GB fact | $699 fact | Apple launch, Apple specifications |
| Mac Studio, M1 Max | Mar. 18, 2022 fact | 400 GB/s fact | 64 GB fact | $1,999 fact | Apple Studio launch, M1 Max |
| Mac Studio, M1 Ultra | Mar. 18, 2022 fact | 800 GB/s fact | 128 GB fact | $3,999 fact | Apple M1 Ultra |
| Mac mini, M2 | Jan. 24, 2023 fact | 100 GB/s fact | 24 GB fact | $599 fact | Apple M2 mini |
| Mac mini, M2 Pro | Jan. 24, 2023 fact | 200 GB/s fact | 32 GB fact | $1,299 fact | Apple M2 mini |
| Mac Studio, M2 Max | June 13, 2023 fact | 400 GB/s fact | 96 GB fact | $1,999 fact | Apple Studio launch, M2 Ultra specifications |
| Mac Studio, M2 Ultra | June 13, 2023 fact | 800 GB/s fact | 192 GB fact | $3,999 fact | Same Apple sources above |
| Mac mini, M4 | Nov. 8, 2024 fact | 120 GB/s fact | 32 GB fact | $599 fact | Apple M4 mini |
| Mac mini, M4 Pro | Nov. 8, 2024 fact | 273 GB/s fact | 64 GB fact | $1,399 fact | Same Apple source above |
| Mac Studio, M4 Max | Mar. 12, 2025 fact | 410–546 GB/s fact: chip option | 128 GB fact | $1,999 fact | Apple 2025 Studio |
| Mac Studio, M3 Ultra | Mar. 12, 2025 fact | 819 GB/s fact | 512 GB fact | $3,999 fact | Same Apple source above |
| Mac mini, M6 | Sept. 22, 2026 fact: announced availability | 153–170 GB/s fact: configuration | 32 GB fact | $899 fact | Apple 2026 mini |
| Mac mini, M5 Pro | Sept. 22, 2026 fact: announced availability | 307 GB/s fact | 64 GB fact | $1,699 fact | Same Apple source above |
| Mac Studio, M5 Max | Sept. 22, 2026 fact: announced availability | 460–614 GB/s fact: chip option | 128 GB fact | $2,499 fact | Apple 2026 Studio |
| Mac Studio, M5 Ultra | Sept. 22, 2026 fact: announced availability | 1,200 GB/s fact | 512 GB fact | $5,499 fact | Same Apple source above |
The table covers every released Mac mini generation and every Mac Studio chip option through the announced September 2026 shipment wave. Apple did not sell an M1 Pro Mac mini, an M3-series Mac mini, an M4 Ultra Studio, or an M5 base mini fact.
What the cadence says
Mac mini availability intervals were:
- M1 to M2: 26.2 months fact: arithmetic
- M2 to M4: 21.5 months fact: arithmetic
- M4 to the September 2026 generation: 22.4 months fact: arithmetic
- Median: 22.4 months; mean: 23.4 months fact: arithmetic
Mac Studio intervals were 14.9, 21.0, and 18.4 months fact: arithmetic, giving an 18.4-month median and 18.1-month mean fact: arithmetic.
The intervals are irregular enough that “wait one more quarter” is not a reliable purchasing strategy.
Maximum memory has risen in large steps over the full history, but it often stays flat for a generation:
- Pro mini: 32 GB to 64 GB to 64 GB fact
- Max Studio: 64 GB to 96 GB to 128 GB to 128 GB fact
- Ultra Studio: 128 GB to 192 GB to 512 GB to 512 GB fact
The next chip therefore cannot be assumed to bring a larger memory tier.
Bandwidth trend
Generation-to-generation maximum-bandwidth changes were:
- Base mini: +46.5%, +20.0%, then +27.5% using the lower M6 configuration or +41.7% using its highest configuration fact: arithmetic.
- Pro mini: +36.5%, then +12.5%; median gain 24.5% fact: arithmetic.
- Max Studio: 0%, +36.5%, then +12.5% using the highest chip options fact: arithmetic.
- Ultra Studio: 0%, +2.4%, then +46.5% fact: arithmetic.
A 25% next-generation Pro gain is therefore a reasonable central planning case est: recent Pro median rounded. A 12%–37% range better represents the evidence est: bounded by the two observed Pro transitions.
Bandwidth is important because model weights and the growing attention cache must be read from memory during generation. It is not the sole speed measure: software support, prompt length, active parameters, quantization and batching also matter.
Apple’s stated AI direction
Apple’s M5 announcement put a neural accelerator in every GPU core and claimed as much as 3.5 times M4 performance on selected AI workloads fact: Apple claim, not independent testing. The M5 Pro and Max continued that design and Apple claimed up to four times peak GPU AI compute fact: Apple claim. M6 and M5 Ultra continued the same direction while increasing memory bandwidth to as much as 1,200 GB/s fact. See Apple’s M5 direction, M5 Pro and Max and M6/M5 Ultra.
Apple’s “4.8 times” LM Studio statement for the new mini concerns prompt processing fact: Apple claim. It is not proof of 4.8-times text-generation speed or eight simultaneous production sessions.
What reportedly comes next
Named reporting by Bloomberg’s Mark Gurman, summarized by MacRumors on June 25 and July 12, 2026, says:
- Base M7 machines may arrive in the first half of 2027 [RUMOR].
- M7 Pro and M7 Max may arrive near the end of 2027 [RUMOR].
- M7 Ultra may follow in 2028 [RUMOR].
- Apple may skip M6 Pro and M6 Max [RUMOR].
Sources: June report and July report.
Those reports correctly anticipated parts of the August 2026 product line, which raises their credibility, but the remaining dates and specifications are still rumors. They are not suitable as purchase commitments.
Apple has announced a September 9, 2026 event fact, but it has not promised a Mac announcement for that event unknown. See Apple’s event notice.
Memory-price pressure and Apple’s price changes
TrendForce forecast conventional PC DRAM contract prices rising more than 100% quarter over quarter in the first quarter of 2026 fact: Feb. 2 forecast. In July, it projected another 13%–18% quarter-over-quarter conventional DRAM increase and a 10%–15% NAND flash increase in the third quarter fact: July 3 forecast. Forecasts can be wrong, but the direction matches Apple’s actions. Sources: TrendForce February and TrendForce July.
In March 2026, Apple temporarily removed the 512 GB Studio option and raised the 256 GB upgrade from $1,600 to $2,000 fact: contemporaneous reporting. In June, Apple attributed broader increases to memory and storage costs driven by AI data-center demand fact: reported Apple statement.
Compared with the previous generation’s original starting price, the September line is higher by:
- Base mini: $300, or 50.0% fact: arithmetic
- Pro mini: $300, or 21.4% fact: arithmetic
- Max Studio: $500, or 25.0% fact: arithmetic
- Ultra Studio: $1,500, or 37.5% fact: arithmetic
The amount of each increase caused specifically by memory is unknown; processors, storage configurations and product positioning also changed.
The 512 GB M5 Ultra price is due in late October 2026 fact: Apple timing, but its exact day and final configured price are unknown.
# 2. Open-model trajectory
Method and warning
The following table asks a narrow question: what was the highest-scoring downloadable open-weight model released by each quarter’s end whose four-bit weights could fit inside 48, 64, 128, 256 or 512 GB?
“Four-bit” means roughly half a gigabyte per billion parameters. Where an official artifact size was unavailable, the fit estimate is:
total parameters × 0.5 GB × 1.10 overhead est.
This is only a weight-fit test. It does not reserve memory for macOS, the model’s working cache, long prompts, simultaneous users or the server. The service plan should retain its 15% operating reserve est from the machine-matrix method.
Scores are the currently reported Artificial Analysis Intelligence Index v4.1.1 scores observed September 3, 2026 fact: provider-reported, sometimes model-estimated. Applying one current score version avoids joining incompatible index versions. The winning-model choice remains est, because no public catalog proves it contains every downloadable model and every four-bit artifact.
Best open model by weight-only tier
| Quarter | 48 GB | 64 GB | 128 GB | 256 GB | 512 GB |
|---|---|---|---|---|---|
| 2024 Q1 | Solar Mini — 6 | Solar Mini — 6 | Solar Mini — 6 | Solar Mini — 6 | Solar Mini — 6 |
| 2024 Q2 | Qwen2 72B — 6 | Qwen2 72B — 6 | Qwen2 72B — 6 | Qwen2 72B — 6 | Qwen2 72B — 6 |
| 2024 Q3 | Qwen2.5 72B — 9 | Qwen2.5 72B — 9 | Qwen2.5 72B — 9 | Qwen2.5 72B — 9 | Qwen2.5 72B — 9 |
| 2024 Q4 | QwQ Preview 32B — 9 | QwQ Preview 32B — 9 | QwQ Preview 32B — 9 | QwQ Preview 32B — 9 | DeepSeek V3 — 14 |
| 2025 Q1 | QwQ 32B — 13 | QwQ 32B — 13 | QwQ 32B — 13 | QwQ 32B — 13 | DeepSeek R1 — 19 |
| 2025 Q2 | QwQ 32B — 13 | QwQ 32B — 13 | QwQ 32B — 13 | QwQ 32B — 13 | DeepSeek R1-0528 — 20 |
| 2025 Q3 | gpt-oss-20b — 15 | gpt-oss-120b — 24 | gpt-oss-120b — 24 | gpt-oss-120b — 24 | gpt-oss-120b — 24 |
| 2025 Q4 | Qwen3 Next 80B — 17 | gpt-oss-120b — 24 | gpt-oss-120b — 24 | GLM-4.7 — 34 | GLM-4.7 — 34 |
| 2026 Q1 | Qwen3.5 35B — 30 | Qwen3.5 35B — 30 | MiniMax M2.5 — 34 | Qwen3.5 397B — 34 | GLM-5 — 41 |
| 2026 Q2 | Qwen3.6 27B — 38 | Qwen3.6 27B — 38 | Qwen3.6 27B — 38 | MiniMax M3 — 45 | GLM-5.2 — 53 |
| 2026 Q3 through Sept. 3 | Qwen3.8 27B — 52 | Qwen3.8 27B — 52 | Qwen3.8 Flash Next — 56 | GLM-5.3 Flash — 57 | GLM-5.3 — 60 |
Every displayed score is fact: current index report; every “best” selection is est: documented method.
Important operating corrections:
- gpt-oss-120b’s 62.4 GB artifact fits 64 GB by weight only fact, but is not a sound 64 GB production choice after operating reserve and prompt cache.
- MiniMax M2.5’s estimated 126.5 GB footprint barely fits 128 GB est; it is not a practical 128 GB campus service.
- The current deployable choices from the machine matrix are Qwen3.8 27B on 48/64 GB, Qwen3.8 Flash Next on 128 GB, GLM-5.3 Flash on 256 GB and GLM-5.3 on 512 GB.
- Kimi K3’s official four-bit package is about 1.56 TB fact; DeepSeek V4 Pro is about 850 GB fact. Neither fits a 512 GB Mac.
- UVU does not need every frontier model locally. It needs useful local tiers plus a controlled cloud path for models that exceed them.
Compact-model progress
The strongest approximately 64 GB-or-smaller series in the existing research was:
- Qwen2.5 72B, September 2024: score 9 fact
- Qwen3 Next 80B, September 2025: score 17 fact
- Gemma 4 31B, April 2026: score 30 fact
- Qwen3.8 27B, August 2026: score 52 fact
A simple four-point regression gives 20.62 index points per year, with an R² of 0.825 fact: arithmetic. It projects 65.12 in September 2027 and 85.78 in September 2028 est.
Those point forecasts should not be used in a budget. The sample has only four models fact; Artificial Analysis changed its index repeatedly; the scale may have a ceiling; and model selection creates survivor bias. A safer September 2027 compact-tier planning range is 58–66 est: regression center with a wide judgment discount.
Open versus closed
| Observation date | Evidence | Gap |
|---|---|---|
| Nov. 4, 2024 | Epoch broad capability review | About 12 months est from study |
| Oct. 30, 2025 | Epoch ECI comparison | About 3 months and 7 ECI points est from study |
| May 29, 2026 | Epoch update | About 4 months; 90% interval 7–11 ECI points; strict crossing about 6 months est from study |
| May 1, 2026 | NIST evaluation of DeepSeek V4 Pro | About 8 months on its particular suite est from NIST analysis |
| Sept. 3, 2026 | Current best open versus leading closed index result | 60 versus 66, a 6-point gap fact: current index |
| Sept. 3, 2026 | Chatbot Arena comparison | 1,489±5 versus 1,507±5, an 18-point gap fact: current leaderboard |
Epoch’s broad “best open” estimate and NIST’s named DeepSeek result are not mutually exclusive: they measure different model sets and tasks. Together they support a present planning range of roughly 3–8 months est, not one exact lag.
Artificial Analysis changed from index v1 to v2 in February 2025, then through v2.1, v2.2, v3, v4, v4.1 and v4.1.1 by August 2026 fact. Scores from those versions must not be plotted as one uninterrupted historical line.
Release cadence by maker
| Maker | Observed releases | Cadence conclusion |
|---|---|---|
| Qwen | At least 7 material families or updates from Sept. 2024 through Aug. 2026 fact | Roughly one material release every four months est, with bursts |
| Z.ai | GLM-4.5, 4.6, 4.7, 5, 5.1, 5.2 and 5.3 from July 2025 through Aug. 2026 fact | Median observed gap about 63 days, or 2.1 months fact: arithmetic |
| DeepSeek | V3, R1 and several major revisions through V4 Pro fact | Irregular; about 3–6 months between major capability steps est |
| Moonshot | K2 through K3 across six named releases from July 2025 through July 2026 fact | Median observed gap about 82 days, or 2.7 months fact: arithmetic |
| OpenAI open weights | One family with 20B and 120B sizes on Aug. 5, 2025 fact | Future cadence unknown |
| Mistral | Mistral 3 on Dec. 2, 2025 and Small 4 on Mar. 16, 2026, plus specialists fact | General-model cadence is irregular and unknown |
| Meta | Llama 3, 3.1, 3.3 and 4 from Apr. 2024 through Apr. 2025 fact; Muse Spark remained a private preview in 2026 fact | No new downloadable general family for about 17 months by Sept. 2026 fact; restart date unknown |
Official sources include Qwen2.5, Qwen3, Z.ai release notes, DeepSeek updates, Moonshot’s release index, gpt-oss, Mistral 3, Mistral Small 4, Llama 4 and Muse Spark.
Plausible next 12–24 months
These are planning judgments, not maker promises:
- At least two meaningful new models that fit the 128 GB weight tier by September 2027: 75% likelihood est: Qwen, Z.ai and Moonshot cadence.
- A 48/64 GB model reaching approximately 58–66 on the current index by September 2027: 60% likelihood est: compact-model trend, discounted for benchmark changes.
- At least one high-profile new flagship still exceeding 512 GB: 80% likelihood est: Kimi K3, DeepSeek V4 Pro and Qwen flagship direction.
- Tool use, image input and longer contexts becoming normal in compact models: 80% likelihood est: current maker roadmaps and releases.
- A fully permissive license becoming standard across the strongest models: below 50% likelihood est: current mixture of Apache, MIT and custom licenses.
The planning implication is important: smaller models are likely to make today’s machines more useful, even while the largest models remain too large. Model progress does not automatically make the hardware obsolete.
# 3. Cost of waiting, in numbers
Capacity forgone each month
The current plan assumes four simultaneous streams per 48 GB mini est. Delivered-model research estimates eight streams at at least ten generated tokens per second for the current 27B Qwen model est: not yet tested on delivered M5 Pro hardware.
| Delay | Four-machine validation set | 26-machine fleet |
|---|---|---|
| One month | 4 device-months fact; 16–32 stream-months est | 26 device-months fact; 104–208 stream-months est |
| 15 months to a rumored late-2027 Pro generation | 60 device-months fact; 240–480 stream-months est | 390 device-months fact; 1,560–3,120 stream-months est |
| 22.4-month historical mini median | 89.6 device-months fact; 358.4–716.8 stream-months est | 582.4 device-months fact; 2,329.6–4,659.2 stream-months est |
A stream-month is one model-serving slot available for one month. It is a capacity measure, not proof of demand or student benefit.
The plan prices four machines at $8,916 and 26 machines at $57,954 est: education configuration plus $90 network upgrade per unit. Spread over 48 months, those amounts are $185.75 and $1,207.38 per calendar month est inputs; FACT arithmetic.
Those amounts are not the economic loss from waiting. UVU retains the cash when it waits. The monetary value of missed work is unknown until the pilot measures demand, successful tasks, time saved and outcome value.
Adoption momentum
The behavioral findings show that access does not equal repeat use:
- CSU reported at least 250,000 activations by spring 2026 fact, but activation is not proof of sustained use.
- One reported measure found student training at about 0.7% and faculty training at 16% fact: different populations and denominators.
- CSU Bakersfield recorded 37 completions among 43 participants offered a $500 incentive, or 86% fact: arithmetic, but it had no control group.
- Virginia Tech reported typical weekly use among 78% of 425 surveyed users fact; this is a selected-user result, not a campus rate.
- Manchester’s reported 90% measure used a 30-day definition that was not public in enough detail to reproduce fact rate; UNKNOWN definition.
For a 30-faculty cohort, a one-month delay removes 30 faculty-months of scheduled exposure est: arithmetic under a fixed cohort. Missing the roughly four-month spring teaching period removes about 120 faculty-months est. It does not prove 120 lost completions or any exact learning loss.
The sound adoption measure is whether a participant returns at the next natural teaching task within 30 days est recommendation, not whether an account was activated or used every week.
September shipment and course preparation
The machines become available September 22, 2026 fact: Apple announcement.
UVU’s spring 2027 schedule says faculty return January 4, classes start January 11 and the term ends May 5 fact. UVU says instructors receive a Canvas shell for assigned classes fact, but a university-wide public date for creating those shells was not found unknown. New employees may not receive access until a class is within 30 days fact: UVU faculty guidance.
Therefore:
- September through November is a defensible course-design and recruitment window est.
- November 30 is a proposed internal content-freeze date est, not an official UVU deadline.
- Waiting past that window risks turning a spring teaching pilot into a technical demonstration rather than a course-integrated service est.
Sources: UVU spring dates and UVU faculty Canvas resources.
Does a faster future machine repay the wait?
For a 48-month service horizon, a future machine that is proportionally faster by g repays a wait of at most:
maximum wait = 48 × g ÷ (1 + g) fact: algebra.
| Future service gain | Maximum break-even wait |
|---|---|
| 12% est low case | 5.1 months fact: arithmetic |
| 25% est central case | 9.6 months fact: arithmetic |
| 40% est high case | 13.7 months fact: arithmetic |
A roughly 15-month wait for a rumored late-2027 Pro machine misses even the 40% case. A normal 22.4-month mini interval misses it by a wider margin.
This calculation assumes immediate full use, identical price and no resale value. Actual adoption will ramp more slowly, which makes the result less decisive. Conversely, missing a fixed spring cohort makes delay more costly. Those effects should be measured, not guessed.
Plain decision rule
Buy the four-machine validation set after September 22 if all of these conditions hold:
- 1. One delivered machine loads the selected model and completes the 1,000-, 8,000- and 32,000-token tests est test lengths.
- 2. It sustains at least four simultaneous service streams at the plan’s ten-token-per-second floor est acceptance rule.
- 3. It passes batch sizes 1, 2, 4 and 8, cold-load, peak-memory and 24-hour soak tests est acceptance protocol.
- 4. Identity, logging, network isolation, accessibility and administrator recovery tests pass.
- 5. Apple has not officially announced, for shipment within 90 days est decision window, a machine that provides required memory or at least 35% more measured service throughput for no more than 20% additional configured cost est thresholds.
Delay only when one of these is true:
- No current configuration passes a must-have requirement.
- An official product inside that 90-day window crosses the test above.
- A dated state, federal or shared-compute award within 60 days est window would cover a material part of the same need.
- Accessibility or security testing blocks the intended service.
Do not delay merely because another chip will eventually exist.
# 4. Refresh and resale, forward
Observed retained value
| Prior machine | Age at observation | Observed value versus original price |
|---|---|---|
| M2 Pro mini, 32 GB/1 TB | About 3.6 years fact | $924.95 versus $1,899; 48.7% fact: transaction and arithmetic |
| M2 Max Studio | About 3.3 years fact | $979–$1,249 versus $1,999; 49.0%–62.5% fact |
| M2 Ultra Studio | About 3.3 years fact | $2,049–$2,369 versus $3,999; 51.2%–59.2% fact |
| M1 Ultra Studio | About 4.5 years fact | Approximately 40%–55% fact: market sample |
| Base M1 mini | About 5.8 years fact | Swappa average sold price $402 versus $699; 57.5% fact: Sept. 2026 snapshot and arithmetic |
The M1 mini result may be temporarily inflated by new-memory scarcity est. Samples also differ by memory, storage, condition, fees and sale channel. Net resale proceeds after fees, labor and warranty risk are unknown.
Conservative planning values remain:
- Year-three mini residual: 50% est
- Year-three Studio residual: 50% est
- Year-three Ultra residual: 55% est
- Year-four mini residual: 40% est
These are budget assumptions, not guaranteed sale prices.
The money analysis calculated five-year cash outlay of $90,341 for keeping the fleet, $122,438 for a year-three refresh and $128,233 for a year-four refresh est. Those cases used different terminal fleet values, so they do not prove that one refresh date has the lowest economic cost.
Recommended option-style purchase plan
An option-style plan buys enough evidence to make the next choice without locking in the whole fleet.
Wave 1 — four machines
After September 22, buy four configured M5 Pro 48 GB minis for an estimated $8,916 all-in est, subject to the delivered-unit gate.
Purpose:
- Validate actual throughput and stability.
- Operate a 30-faculty cohort est cohort size.
- Establish demand, support effort and repeat-use evidence.
- Preserve the right not to buy the remaining 22 machines.
Wave 2 — up to eight more
Release an additional order only after:
- At least 50% of pilot participants return at their next natural task within 30 days est adoption gate.
- Measured service demand uses at least 60% of available capacity during meaningful teaching windows for four weeks est operating gate.
- Support load, accessibility and incident handling remain within the staffed plan.
- No current model release materially changes the best hardware tier.
Eight additional machines would make 12 total est option size. The exact number should be set from queue and concurrency evidence, not this placeholder.
Wave 3 — up to the planned 26
Release the final order only after spring-course results, observed queue time and a refreshed local-versus-cloud cost comparison. Fourteen additional machines would reach 26 total est option size; FACT arithmetic.
Do not reserve a 26-machine quantity simply to preserve architectural symmetry.
Special triggers
- Delivered-unit test: A failure changes the configuration or pauses the purchase; it does not lower the gate.
- 512 GB price: Consider a 512 GB Ultra only if at least two approved research workloads need a model that cannot fit 256 GB est gate, delivered tests confirm the benefit, and the configured price is no more than the cost of approximately six all-in minis—$13,374 est threshold; FACT arithmetic.
- Next Apple event: Reopen only unplaced orders. Do not idle machines already producing evidence.
- Model release that changes a tier’s leader: Re-run the narrow model/machine test. A better model does not automatically require new hardware.
- Year three: Obtain real resale quotes, compare service throughput per dollar and decide which machines merit replacement.
- Year four: Refresh where the measured gain clears the same waiting rule. Keep useful older machines as development, test, spare or low-priority capacity rather than forcing a uniform replacement.
# 5. What could break the trajectory
The percentages below are subjective 24-month planning priors est, not measured frequencies.
| Risk | Likelihood | Signal to watch | Plan response |
|---|---|---|---|
| US–China policy or geopolitics slows open-weight releases | Medium, 35% est | New BIS restrictions involving model weights or training; PRC export controls; repository removal | Keep approved model artifacts and license records; maintain several makers; retain a metered cloud route |
| Licenses tighten | Medium, 40% est | New commercial-use, redistribution, data-use or geographic clauses | Review every version before deployment; never auto-upgrade; keep Apache/MIT fallbacks |
| Apple changes memory prices again | High, 70% est | TrendForce contract-price reports; configurator changes; removal of high-memory options | Buy the four-machine option now; set price caps for later waves; do not budget the 512 GB machine before its real quote |
| Serving leadership shifts between MLX and llama.cpp | High, 65% est | Unsupported architecture, crash or throughput regression in either stack | Keep one common service interface; allow either runtime; test both for each selected model |
| A frontier lab releases a large permissive model | Medium-low, 30% est | Downloadable model above 300B parameters with a usable commercial license est signal threshold | Re-test 256/512 GB value and cloud pricing before a Studio order |
| Hosted open-model prices collapse | Medium, 45% est | Delivered cloud cost per successful task stays at least 30% below fully loaded local cost for two quarters est gate | Move variable and oversized work to the hosted pool; retain local privacy and continuity capacity |
Policy and license evidence
NTIA’s July 2024 open-weight report found benefits and risks but did not find enough evidence then for immediate restrictions; it recommended continued monitoring fact. The July 2025 White House AI Action Plan supported open models fact. BIS later changed advanced-chip and training controls fact. Policy currently supports access, but it is not permanent. Sources: NTIA, AI Action Plan and BIS policy statement.
“Open-weight” does not mean “open-source.” gpt-oss uses Apache 2.0 fact; several DeepSeek and Z.ai releases use permissive licenses fact; Meta Llama and some newer models use custom terms fact. UVU needs a version-specific license record.
Serving-stack evidence
MLX supports affine two- through eight-bit quantization plus newer four-bit formats fact. But an April 2026 issue reported unsupported DeepSeek V4 architecture, and August reports described a Metal-buffer leak and an invisible server-thread failure fact: issue reports, not universal defects. llama.cpp supports Apple Metal and native formats for several families fact.
The response is not to select one permanent winner. The service boundary should allow either runtime while keeping authentication, logs and user behavior stable.
Hosted-price evidence
Together currently lists GLM-5.3 Flash at about $0.15 per million input tokens and $0.50 per million output tokens fact: Sept. 2026 listed price. Its April DeepSeek V4 Pro offer was $2.10 input and $4.40 output per million tokens fact. These prices can fall or rise.
No responsible local-versus-cloud comparison is possible until UVU knows:
- Prompt and output volume per completed task.
- Cache-hit rates.
- Peak demand.
- Local support and electricity cost.
- The privacy class of each request.
All are unknown before the pilot.
# 6. Triggers and calendar
| Date | Event | Decision |
|---|---|---|
| Sept. 9, 2026 fact event date | Apple event | Reopen only unplaced orders if an official Mac ships within 90 days and crosses the hardware rule |
| Sept. 9, 2026 fact meeting date | Utah Board of Higher Education meeting | Check agenda and later minutes for AI compute allocation and task-force action |
| Sept. 22, 2026 fact announced date | M6/M5 Pro mini and M5 Studio shipment wave | Accept and test; do not scale until the delivered gate passes |
| Sept. 30, 2026 est checkpoint | Quarterly model review | Re-run only tiers whose leading model changed |
| Oct. 1–2, 2026 fact meeting dates | USHE meeting | Check allocation, credential and task-force decisions |
| Oct. 31, 2026 est operating deadline | Late-October 512 GB price expected | Compare real quote and tested unique workload against six-mini threshold |
| Nov. 4, 2026 fact deadline | NSF AI Infrastructure Hubs deadline | Submit only through an approved consortium and normal institutional approval |
| Nov. 19, 2026 fact meeting date | USHE meeting | Check state compute and task-force milestones |
| Nov. 30, 2026 est internal date | Proposed spring course-content freeze | Decide whether spring launch is course-integrated or a smaller technical pilot |
| Jan. 4, 2027 fact | UVU faculty return | Finish faculty orientation and accessibility checks |
| Jan. 11, 2027 fact | UVU spring classes begin | Start measured pilot only if launch gates pass |
| Jan. 20, 2027 fact deadline | NSF IUSE deadline | Decide whether measured pilot evidence supports an institutional proposal |
| Mar. 25–26, 2027 fact meeting dates | USHE meeting | Check state allocation and task-force progress |
| Mar. 31, 2027 est checkpoint | Spring Apple/model review | Reopen unplaced hardware options only |
| Apr. 26, 2027 fact federal date | Federal web/mobile accessibility deadline for covered state entities with populations at least 50,000 | Treat WCAG 2.1 AA compliance as a launch requirement; UVU counsel should confirm the exact legal classification unknown legal determination |
| May 13, 2027 fact meeting date | USHE meeting | Review recommendations, allocations and credential implementation |
| June 30, 2027 est checkpoint | Midyear model and Apple review | Compare compact-model improvement with delivered fleet demand |
| July 21, 2027 fact deadline | Second NSF IUSE deadline | Apply only if evidence and university approvals support it |
| Sept. 30, 2027 est checkpoint | One-year hardware/model review | Decide whether Wave 3 or a changed tier is justified |
| Dec. 31, 2027 [RUMOR window] | Reported possible M7 Pro/Max window | Compare actual product, if any, with year-one evidence |
| Mar. 31, June 30 and Sept. 3, 2028 est checkpoints | Quarterly reviews through the 24-month horizon | Review models, prices, grants and runtime support; avoid calendar-driven replacement |
USHE announced a statewide AI task force on May 1, 2026 and named UVU representation fact, but no public deadline for its final recommendations was found unknown. USHE also reported a $15 million one-time state allocation for system AI compute capacity fact; UVU’s share, access rules and award timing are unknown. Sources: USHE task force, legislative update and meeting calendar.
NAIRR provides shared research-compute allocations fact, but the next allocation date and capacity available to this service are unknown. It should be watched as a complement, not used as an assumed fleet replacement. See NAIRR.
The Department of Justice’s April 2026 rule change moved the relevant federal accessibility date to April 26, 2027 for state and local entities serving at least 50,000 people fact. Public universities and vendor-supplied content are within the rule’s described scope fact; UVU should obtain counsel’s classification rather than infer it from student enrollment. See the DOJ compliance guide.
# Cross-domain learnings
Software delivery: canary releases
Site-reliability teams send a change to a small part of the system, compare it with the existing service and stop the rollout when the evidence is bad. Google’s published canary method formalizes that approach fact.
UVU equivalent: four machines are the canary. The other 22 are not a promised follow-on order.
Source: Google SRE canarying.
Clinical trials: preplanned adaptation
Adaptive clinical trials allow enrollment or allocation to change at interim reviews, but the rules are written before investigators see the results fact.
UVU equivalent: approve throughput, adoption, accessibility and price triggers now. Do not invent a favorable success rule after the pilot.
Source: FDA adaptive-design guidance.
Capital programs: stage gates
GAO’s technology guidance separates discovery, development, production and deployment decisions fact. Each gate asks for stronger proof.
UVU equivalent: delivered-unit performance, faculty-pilot behavior and campus expansion are different gates. SSH access or a loaded model is not campus readiness.
Source: GAO technology transition guidance.
Public health: tools do not create behavior by themselves
A bundled hand-hygiene program in Geneva reported compliance rising from 48% to 66% and infections falling from 16.9% to 9.9% fact: uncontrolled before/after evidence. A mandatory surgical checklist in Ontario later showed no significant population-level improvement fact.
The general lesson is not that checklists work or fail. The surrounding training, feedback, local leaders and workflow matter. UVU should fund adoption work, not count machines or activated accounts as success.
# Devil’s advocate
The strongest case against buying even four machines is this:
Hosted open models are already inexpensive, and their prices may fall faster than local hardware costs. The largest open models exceed Apple’s 512 GB ceiling. Apple memory cannot be upgraded after purchase. Current throughput is estimated, not delivered proof. Local service also creates staffing, patching, accessibility, incident-response and physical-continuity costs. USHE’s $15 million allocation, NAIRR or another shared service may cover the same need. A rumored Pro-class Apple update could arrive near the end of 2027.
On that view, UVU should run the first faculty pilot entirely through free or metered hosted tools, gather demand through spring 2027 and make no hardware purchase until state allocations and M7 products are clear.
This is a strong argument against an immediate 26-machine fleet. It is weaker against four machines if UVU’s goals include private local work, predictable service, instruction in local AI operations or protection against provider changes. The four-machine purchase is best understood as the price of learning whether those benefits are real.
# What would change the recommendation
The recommendation should change from “buy four, then stage” to “wait or use cloud” if any of these occurs:
- The delivered M5 Pro cannot sustain four simultaneous ten-token-per-second sessions under the required context and safety settings.
- An official machine shipping within 90 days supplies required memory or at least 35% more measured service throughput for no more than a 20% configured-price premium est rule.
- USHE or NAIRR provides recurring, usable capacity before UVU must commit the order.
- Hosted service remains at least 30% cheaper per successful task for two quarters after privacy, support and egress are included est rule.
- Fewer than 40% of pilot participants return at their next natural task within 30 days est stop/rework threshold.
- Accessibility, identity or incident-recovery requirements cannot be met without a material redesign.
- A new license prevents the intended campus use.
- The final 512 GB price is unusually low and at least two approved workloads demonstrate a unique need for the full GLM tier; that would add a Studio option, not justify replacing the mini validation set.
# What the plan should change
- 1. Replace “buy after September 22” with a delivered-unit acceptance gate. Shipment is availability, not performance proof.
- 2. Keep the four-machine purchase; remove any implied commitment to all 26. Use four, then up to eight, then up to fourteen more est wave sizes.
- 3. Replace the fixed four-stream capacity constant. Report four as the acceptance floor and eight as an unverified upside for the current 27B model.
- 4. Add the waiting equation. A 25% improvement repays no more than 9.6 months of waiting inside a 48-month horizon est input; FACT arithmetic.
- 5. State the memory rule plainly. Buy for models that create useful service now; use the metered pool when a flagship exceeds local memory.
- 6. Make adoption a purchase gate. Measure return at the next natural task within 30 days, support burden and completed work—not activation.
- 7. Move the year-three review and year-four refresh from policy to default options. Refresh only where actual throughput, resale and workload data justify it.
- 8. Add dual-runtime acceptance. MLX and llama.cpp should be interchangeable behind one campus-facing service contract.
- 9. Add the dated 24-month watchlist. Include USHE meetings, grant deadlines, Apple events, quarterly model checkpoints and the April 26, 2027 accessibility deadline.
- 10. Mark the 512 GB line as unpriced. Do not budget or recommend it until Apple posts the final configuration and UVU demonstrates a unique workload.
- 11. Update the open-model history quarterly using one benchmark version. Preserve old snapshots separately; do not splice changing index versions.
- 12. Add a per-successful-task local/cloud comparison. Token price alone does not include failed work, support, privacy, peak capacity or staff time.
# Source log
- 1. Sept. 3, 2026 fact access date — Current UVU program.
- 2. Sept. 3, 2026 fact access date — Mac generations findings.
- 3. Sept. 3, 2026 fact access date — Model-progress findings.
- 4. Sept. 3, 2026 fact access date — Model/machine matrix.
- 5. Sept. 3, 2026 fact access date — Money findings.
- 6. Sept. 3, 2026 fact access date — Adoption findings.
- 7. Nov. 10, 2020 fact — Apple introduces M1 Macs.
- 8. Nov. 2020; accessed Sept. 3, 2026 fact — Apple M1 mini specifications.
- 9. Oct. 18, 2021 fact — Apple M1 Pro and M1 Max.
- 10. Mar. 8, 2022 fact — Apple Mac Studio launch.
- 11. Mar. 8, 2022 fact — Apple M1 Ultra.
- 12. Jan. 17, 2023 fact — Apple M2 and M2 Pro mini.
- 13. June 5, 2023 fact — Apple M2 Studio.
- 14. June 5, 2023 fact — Apple M2 Ultra.
- 15. Oct. 29, 2024 fact — Apple M4 mini.
- 16. Mar. 5, 2025 fact — Apple M4 Max and M3 Ultra Studio.
- 17. Oct. 15, 2025 fact — Apple M5 AI-compute direction.
- 18. Mar. 3, 2026 fact — Apple M5 Pro and M5 Max.
- 19. Aug. 25, 2026 fact — Apple M6/M5 Pro mini.
- 20. Aug. 25, 2026 fact — Apple M5 Max/M5 Ultra Studio.
- 21. Aug. 25, 2026 fact — Apple M6/M5 Ultra architecture.
- 22. Feb. 2, 2026 fact publication date — TrendForce DRAM forecast.
- 23. July 3, 2026 fact publication date — TrendForce third-quarter memory forecast.
- 24. Mar. 5, 2026 fact publication date — 512 GB Studio option removal.
- 25. Mar. 6, 2026 fact publication date — Studio memory upgrade-price report.
- 26. June 29, 2026 fact publication date — Apple product price-change report.
- 27. June 25, 2026 fact report date; RUMOR content — Reported M7 schedule.
- 28. July 12, 2026 fact report date; RUMOR content — Reported M6 Pro/Max omission.
- 29. Aug. 2026 fact notice date — Apple September 9 event notice.
- 30. Nov. 4, 2024 fact publication date — Epoch open-model report.
- 31. Oct. 30, 2025 fact publication date — Epoch open/closed comparison.
- 32. May 29, 2026 fact publication date — Epoch ECI gap update.
- 33. May 1, 2026 fact publication date — NIST DeepSeek V4 Pro evaluation.
- 34. Sept. 19, 2024 fact — Qwen2.5.
- 35. Apr. 29, 2025 fact — Qwen3.
- 36. Accessed Sept. 3, 2026 fact — Z.ai release notes.
- 37. Accessed Sept. 3, 2026 fact — DeepSeek update history.
- 38. Accessed Sept. 3, 2026 fact — Moonshot release index.
- 39. July 16, 2026 fact — Kimi K3.
- 40. Aug. 5, 2025 fact — OpenAI gpt-oss.
- 41. Dec. 2, 2025 fact — Mistral 3.
- 42. Mar. 16, 2026 fact — Mistral Small 4.
- 43. Apr. 5, 2025 fact — Meta Llama 4.
- 44. Apr. 8, 2026 fact — Meta Muse Spark.
- 45. Accessed Sept. 3, 2026 fact — Artificial Analysis open-model leaderboard.
- 46. Accessed Sept. 3, 2026 fact — Artificial Analysis Qwen3.8 27B record.
- 47. Accessed Sept. 3, 2026 fact — Artificial Analysis gpt-oss-120b record.
- 48. Accessed Sept. 3, 2026 fact — Artificial Analysis GLM-5 record.
- 49. Accessed Sept. 3, 2026 fact — Artificial Analysis GLM-5.2 record.
- 50. Accessed Sept. 3, 2026 fact — MLX quantization documentation.
- 51. Apr. 30, 2026 fact issue date — MLX DeepSeek V4 support issue.
- 52. Accessed Sept. 3, 2026 fact — llama.cpp project.
- 53. July 30, 2024 fact — NTIA open-model-weight report.
- 54. July 2025 fact — America’s AI Action Plan.
- 55. May 13, 2025 fact — BIS AI-model-training policy statement.
- 56. Accessed Sept. 3, 2026 fact — Meta Llama 4 license.
- 57. Accessed Sept. 3, 2026 fact — Together pricing.
- 58. Apr. 2026 fact — Together DeepSeek V4 Pro pricing.
- 59. Updated June 30, 2026 fact — UVU spring schedule.
- 60. Accessed Sept. 3, 2026 fact — UVU faculty Canvas resources.
- 61. Mar. 10, 2026 fact — USHE legislative update and AI compute allocation.
- 62. May 1, 2026 fact — USHE AI task force.
- 63. Accessed Sept. 3, 2026 fact — USHE meeting calendar.
- 64. Apr. 2026 fact — DOJ accessibility compliance guide.
- 65. 2026 solicitation fact — NSF AI Infrastructure Hubs.
- 66. Accessed Sept. 3, 2026 fact — NSF NAIRR.
- 67. Accessed Sept. 3, 2026 fact — Google SRE canary releases.
- 68. Dec. 2019 fact — FDA adaptive-trial guidance.
- 69. July 2006 fact — GAO technology transition guidance.
- 70. Accessed Sept. 3, 2026 fact — Swappa M1 mini market.
# NOT_RUN
- not run — Vendor, UVU, USHE, NSF or regulator contact, as prohibited.
- not run — Purchase, reservation, quote request or configurator checkout.
- not run — Delivered M5 Pro/M6/M5 Studio testing; machines do not ship until September 22, 2026.
- not run — 512 GB M5 Ultra price comparison; final price is not yet public.
- not run — Live UVU identity, network, accessibility, security or recovery testing.
- not run — Legal determination of UVU’s exact federal accessibility classification.
- not run — Verification of a public Utah AI workforce RFP; none was located.
- not run — Exhaustive machine-readable reconstruction of every historical open-model leaderboard. The tier winners are documented estimates using the stated fit rule.
- not run — Economic-value calculation for missed usage; UVU demand, task value, support cost and successful-task volume are unknown.
The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/47-roadmap-timing.md in the research pack.
Method note
Every appendix row traces to a dated public source or a measured result in the Utlyze working papers (forty-eight research lanes, an adversarial verification pass, a three-persona red team, and a live benchmark of the exact serving stack on our own 512GB-class hardware — method in the guide, §13). Claims are graded fact where a cited source states them, est where we computed or judged, unknown where the public record is silent. Nothing in these appendices required contacting UVU — by design, so this document arrives owing nothing.
Continue: the guide · the budget simulator · the campus map. This page prints clean — File → Print.