UTLYZE
Appendices · evidence date September 3, 2026

The evidence,
in full.

This page holds the evidence behind the plan. Use Appendices A–Z to check the order the colleges are proposed in, what the national data actually measures, what each college needs built differently, every free or discounted program worth claiming, the pre-decided plan-B branches, the questions only UVU can answer, which model runs on which machine, how the service is engineered, duty of care and resilience, Utah law and what peer universities do, the money angles, assessment and integrity, learning outcomes, trust and sources, accessibility and language, and roadmap and timing. Each claim should show a source and say whether it is fact (checked against that source), estimate (a planning calculation) or unknown (not yet confirmed). Where a claim carries no chip, treat it as an estimate and check it against the linked source before relying on it. Two words recur in the tables and in the five-line summaries: a fact nobody has been able to establish yet is UNKNOWN, and a test or a request the plan has written down but has not yet performed is NOT_RUN.

Prepared independently by Utlyze. Not a publication of, endorsed by, or affiliated with Utah Valley University.

Read the five lines first. Every appendix opens with the same card: the question it answers, the answer, the numbers the answer turns on with their fact est unknown grades, what the plan does about it, and what is still unknown. Read those five lines, and open the tables under them only where you want to check the work.

APPENDIX ACollege sequencing & receptivityThe wave order on the campus map, shown with its work APPENDIX BDiscipline usage — the national dataEvery 2024–26 study that splits AI use by field APPENDIX CWhat each college needs builtUses, capabilities, controls, and success measures APPENDIX DDepartment by departmentAll 51 catalog units: demand, uses, model route, constraints APPENDIX EWhat UVU already hasFull asset inventory, gap ledger, and the ask-once data packet APPENDIX FThe free-and-discounted catalogPrograms UVU can claim before spending a dollar APPENDIX GThe contingency playbook28 pre-decided branches; the three that matter most APPENDIX HOpen items & the ask-once listWhat remains unknown, who owns it, what we ask Dr. Burns APPENDIX IThe campus map data53 building pins, college homes, regions, imagery rights APPENDIX JWho is already doing itNamed champions, leadership statements, friction, students, adjuncts APPENDIX KWhat we plug intoIntegration surfaces, compute reality, and the five remaining questions APPENDIX LThe free path, honestly28 providers, the terms, the capacity math, Muse Spark, harnesses, verdicts APPENDIX MThe harness market171 tools: adoption, growth, funding, stability, dead-or-declining, fit APPENDIX NAdoption, auditedSix-source audit, 30 cross-domain cases, the $20k first wave, ranked changes APPENDIX OThe two Mac lineups, side by sideNewest vs previous: launch prices, today's availability, speed evidence, the fleet priced both ways APPENDIX PEvery angle, and where it livesThirty-eight angles graded COVERED, THIN, or OPEN, with where each lives APPENDIX QModel × machine: the matrixTen models × six machines; two models per box; same money both ways; the 512GB decision APPENDIX RService engineeringWorkload classes, prefill limits, batching and caching, service levels, agents, the queue model APPENDIX SDuty of care and resilienceMinors, HIPAA, crisis handoff, twenty graded security controls, failover, recovery targets, incident runbook APPENDIX TUtah law, state policy, licenses, and peersUtah AI disclosure law, ADA Title II deadline, FERPA and open records, USHE, license-fit table, six Utah peers APPENDIX UMoney anglesFive-year cases, resale values, lease vs buy, who pays, student staffing, grants, electricity APPENDIX VAssessment redesign and academic integrityIntegrity data, detector evidence, designs that hold up, Secure/Open labels, seven college redesigns APPENDIX WDoes it improve learning? The evidence and the measurementTrials with effect sizes, mechanisms, tutor modes, a pre-registered study, the learning-outcomes contract APPENDIX XTrust, sources, and error handling: the interfaceError rates by task, trust calibration, exact-passage citations, repair cases, refusals, the pilot test plan APPENDIX YAccessibility and languageWCAG criteria that bite, LibreChat's evidence, voice, neurodivergent design, Spanish quality, mobile, launch gates APPENDIX ZRoadmap and timingChip cadence with prices, best open model per tier by quarter, cost of waiting, gated waves, 24-month watchlist
A

College sequencing & receptivity

The wave order used on the campus map — derived, not asserted

Appendix A in five lines

  1. QuestionWhich colleges should go first?
  2. AnswerSmith Engineering and Technology and Woodbury Business go first; Education and Humanities and Social Sciences follow in wave two.
  3. Deciding numbersscores of 12 for Smith, Woodbury, and Educationest10 for Humanities and Social Sciences and Health and Public Serviceest5 for Artsest
  4. What the plan doesStart Smith and Woodbury, add Education and Humanities and Social Sciences in wave two, then Science with a research lane; Health and Public Service runs on a separate safety lane, and Arts is a later pilot once rights controls exist.
  5. Still unknownNo public college-level use survey exists, so every score here reads visible behavior rather than measured demand — and a high score does not by itself buy an earlier wave.

No public UVU survey reports AI use by college unknown — so receptivity here measures visible behavior: AI courses in the 2026–27 catalog, named faculty leads, college policies, Kahlert curriculum grants, training completions, and applied projects. That evidence, college by college — cross-checked against what deans and faculty have said in public (Appendix J) — produces the scores below. The scoring rule is printed with the table so it can be argued with: priority = 2×receptivity + demand + story value − complexity est.

A score is not a place in the queue. The score says how ready a college looks. The wave says when the service can actually be built for it, and only two colleges can be stood up at once. Education scores level with Engineering and Business and still opens wave two; Health & Public Service scores above Science and still runs on a lane of its own, because its safety controls have to exist first. The sequence the plan follows is the one in §07 of the guide and on the campus map: wave one Engineering & Technology and Woodbury Business, wave two Education and Humanities & Social Sciences, Science following with the shared research service, Health & Public Service on the safety lane, Arts later once rights controls exist. est sequence

Score rankCollegeReceptivityDemandComplexityStoryScoreWaveWhy, in one line
1Smith Engineering & TechnologyHIGH545121Deepest AI curriculum (MS-AAI, AI certificate, IS Applied-AI emphasis), the new Smith building's AI lab, the NVIDIA agreement, named faculty leads.
2Woodbury School of BusinessHIGH534121Highest national student use (70% weekly+); live AI courses in foresight, automation, strategy; the institute's director teaches here.
3School of EducationHIGH545122AI Academy completions, UNESCO AI-education committee seats, and Utah's K-12 AI rollout (HB 273 + statewide Gemini) make teacher prep time-sensitive. Starts with no live P-12 student data.
4Humanities & Social SciencesHIGH433102The only college with published AI guiding principles and a faculty resource guide; AI courses in English, Communication, Political Science; a Kahlert grant. Largest writing load.
5Health & Public ServiceHIGH, controlled55410separate safety laneRaised from MED after the public-voice sweep (Appendix J): named nursing champions have already built virtual-patient bots; a $35M AI-equipped campus is funded. Clinical work cannot precede partner, privacy, and safety controls — simulation first.
6College of ScienceMED4448follows wave 2 · shared research serviceReal applied work (ML in CHEM 3300, the student-built DIFFRAX device) but concentrated in pockets; needs reproducibility and lab-safety design first. Its heaviest demand is the metered research pool.
7School of the ArtsMED3425later · rights-aware pilotMultimodal value is real, but rights controls — likeness consent, provenance, licensing — must exist before a creative pilot.

How to read "story": the Utah-facing narrative each college can carry. Education ranks first there — "UVU prepares Utah teachers to use AI safely before they enter classrooms" lands statewide while HB 273 and the state's Gemini rollout are live. Engineering's NVIDIA-agreement story and Woodbury's employer-evidence story follow. Sources: UVU 2026–27 catalog, Kahlert Institute pages, UNESCO Chair rosters, Ethics Awareness Week programs, college policy documents, Utah HB 273, the June 2026 Google–USBE announcement — full URL log in the working papers fact (behaviors) / est (scores).

B

Discipline usage — the national data

Every recent study that splits AI use by field, with what each one actually measures

Appendix B in five lines

  1. QuestionWhich fields use artificial intelligence most often?
  2. AnswerBusiness, technology, and engineering students use it most often, but every field still needs support.
  3. Deciding numbersBusiness 70%factTechnology 68%factEngineering 65%fact
  4. What the plan doesThe program starts Engineering and Business, then adds Humanities and Social Sciences with writing and source-check work.
  5. Still unknownThe 94,060-person California State University survey publishes no open field table, so the two largest datasets are logged as checked but unusable.

These rates are not interchangeable — "use" ranges from daily coursework habit to once-a-semester. Read the measure column before comparing rows. The pattern that survives across all of them: business, technology, and engineering lead student frequency; faculty use is above 70% in every field; writing-heavy disciplines feel the most assessment strain.

Study & populationWhat it measuredRates by fieldGrade
Lumina–Gallup, 3,801 U.S. students, Oct 2025 (pub. Mar 2026)Daily + weekly AI use in courseworkBusiness 70% · Technology 68% · Engineering 65% · Vocational 58% · Healthcare 56% · Social sciences 53% · Natural sciences 49% · Humanities 43%fact
College Board, 3,482 U.S. faculty, summer 2025 (pub. Feb 2026)At least one AI use in the faculty roleBusiness & comms 88% · Health sciences 86% · CS 84% · Engineering 83% · Social sciences 81% · Math 74% · English 73% · Arts & music 71%fact
Berkeley/SERU, 95,513 undergrads at 20 public research universities, spring 2024Regular use (monthly+)Computer science 62% · Mathematics 53% · Business 51% · Arts 24%fact
Middlebury, 634 students across 43 majors, Dec 2024–Feb 2025Any academic use during the semesterNatural sciences 91% · Social sciences 85% · Humanities 75% · Arts 73% · Languages 57% · Literature 49%fact
Anthropic telemetry, 574,740 academic conversations, Apr 2025Share of conversations (intensity, not adoption)CS 38.6% of conversations vs 5.4% of U.S. degrees — a 7× intensity ratio; business under-indexes (8.9% vs 18.6% of degrees)fact
Michigan Engineering, n=182, Jun 2025Ever used GenAIEngineering students 96%fact
Grenoble (JMIR), n=388 health students, Apr 2026Used GenAI during a clinical placementPharmacy 64% · Nursing 54% · Medicine 48% · Physiotherapy 46% · Midwifery 26%fact
Kogod Business School, 483 students over 3 waves, Aug 2026Used AI for academic work, prior 6 monthsBusiness students >80%; 29% use it 11+ times weeklyfact

What this means for the plan: capacity follows the frequency leaders — Engineering & Technology and Woodbury Business take wave 1 and its hardware — but service design cannot skip the "low" fields. Humanities' 43% weekly is still nearly half the college, and its faculty carry the heaviest redesign load. That is why wave 2 brings Humanities & Social Sciences in with citation-verification and writing-process tooling rather than raw capacity, and why how often a field uses AI does not by itself set its place in the queue (Appendix A). Sources linked by name; the two large gated datasets (CSU's 94,060-respondent survey, Digital Education Council) publish no open field tables unknown and are logged as checked-but-unusable.

C

What each college needs built

Same engine, seven different services — uses, controls, and what success means

Appendix C in five lines

  1. QuestionWhat different artificial intelligence service does each college need?
  2. AnswerOne shared service needs different tools, controls, and success tests for each college.
  3. Deciding numbersseven different servicesestplanned target of zero incidents with personally identifiable informationestplanned target of zero incidents with protected health informationest
  4. What the plan doesChange the tools by college, start Health with made-up cases, and wait on Arts until consent and license rules exist.
  5. Still unknownNo live build or result is reported for any college service; these checks are NOT_RUN.

"What is in a school stays with that school" is a budget rule and a design rule. Money and use cases stay local; security, identity, and data rules stay central. This table is the design contract per college est, with the binding compliance anchors fact.

CollegeHighest-value usesBuild differencesHard controlsSuccess means
Engineering & TechCode explanation & debugging, cyber labs, technical-image help, simulation, project docsCode execution sandboxes, model choice, technical vision, long contextLearning sandboxes separated from production; secret scanning; human review for aviation/safety-critical outputsBetter project rubrics, testable code, faster feedback, AI credential completions
Woodbury BusinessSpreadsheet automation, SQL & financial modeling, case analysis, market researchTables/charts, structured data, cited web research, document generationLicensed case material stays out of general models (Harvard Business Publishing requires written permission for GenAI uploads); client-data audit trailsVerified model accuracy, employer-scored case decisions, honest disclosure
EducationStandards-aligned lesson plans, rubrics, differentiation, practicum rehearsal, licensure prepGrounded retrieval, reading-level control, tutor mode, teacher-controlled student experiencesUtah HB 273: approved tools only, educator judgment preserved, no AI in high-stakes decisions; no student records or PII in unapproved tools (FERPA)Candidates can explain safe vs unsafe use; less prep time; zero PII incidents
Humanities & Soc SciWriting feedback, citation verification, multilingual tutoring, qualitative coding, deepfake analysisSource-linked retrieval, corpus search, citation checking, transcriptionPreserve student voice and process evidence; never accept invented citations; CHSS's own published AI principles govern; no detector-based accusationsCitation accuracy up, documented writing process, better source evaluation
ScienceMath/stats tutoring, notebooks, literature synthesis, spectrum/microscope image interpretationMath + vision + reproducible notebooks, uncertainty reportingNo unsupervised lab-safety decisions; reproducible calculations; faculty review; synthetic or approved research data onlyConcept mastery, reproducible outputs that survive faculty checking
Health & Public ServiceSynthetic clinical cases, anatomy tutoring, licensure prep, documentation practice, emergency simulationsHigh-reliability grounded answers, strict refusals, full audit, human sign-offSynthetic/de-identified first (HHS Safe Harbor); clinical partners' rules govern placements; separate high-assurance environment before any real clinical workflowLicensure and simulation results up; every consequential output reviewed; zero PHI incidents
ArtsIdeation, previsualization, storyboards, variants, portfolio critique, production planningMultimodal generation/editing, asset history, provenance metadataLikeness and voice consent recorded; source licenses tracked; human authorship preserved (Copyright Office: prompts alone create no copyright)Faster iteration at equal or better jury quality; portfolios still show the student's judgment
D

Department by department

All 51 catalog units — demand tier, uses, model route, and the constraint that governs each

Appendix D in five lines

  1. QuestionWhat route and safety rule does each department need?
  2. AnswerAll 51 catalog units need access, but the model route and safety rule change with the work.
  3. Deciding numbers51 child unitsfact1.8× load for computer, information technology, and cyber workest0.8× text load for culinary and trades workest
  4. What the plan doesKeep budgets by college and set each department’s route, tools, and stop rules.
  5. Still unknownThe type of the Reserve Officers’ Training Corps unit and direct use rates for several fields are UNKNOWN.

The college is the budget unit; the department is where work actually differs. This appendix maps every unit in UVU's 2026–27 catalog — four colleges and three schools, 51 child units fact — to a demand tier, its highest-value uses, a model route, and the constraint that must be honored. Tiers rest on the strongest available evidence per field (a mechanical-engineering cohort study, a 91%-use mathematics sample, clinical-placement rates by health profession, the 38.6% computer-science share of academic AI conversations) and are graded est where triangulated. Notably: no family lands in a LOW tier — every department shows enough recurring task overlap to need access, guidance, and a course-level policy.

The ten routes

RouteMeaning
R1 — local generalFast local models for routine scale; the long-document local model for writing and multilingual work; the vision-capable local model when images help.
R2 — coding hybridStrong local code models first; frontier escalation for hard multi-file work.
R3 — quantitative/science hybridLocal for tutoring and checking; frontier for hard proofs and advanced research.
R4 — controlled health/socialApproved local retrieval first; secure frontier escalation; no clinical or counseling decisions, ever.
R5 — multilingual/speechLocal text translation; a separate approved speech service for audio, transcription, pronunciation.
R6 — creative multimodalLocal analysis; approved frontier image/audio/video services with rights controls.
R7 — safety-critical technicalSource-grounded local support; frontier for review only; human authority always controls.
R8 — business analyticsLocal routine analysis; frontier for complex models and current external facts.
R9 — writing/researchLocal long-document model with retrieval; frontier for difficult synthesis.
R10 — educationLocal lesson and material creation; multimodal or accessibility escalation when needed.

Which models fit which family — and when frontier is genuinely needed

Department familyCapabilities that matterLocal fitFrontier triggerRoute
CS/IT/cyberCode, logs, repositories, tools, long interactionQwen3.8 or GLM-5.3; Qwen3.6 for routine helpMulti-repository work, difficult debugging, advanced defensive securityHYBRID est
Mechanical/civil/electrical engineeringMath, code, diagrams, documents, CAD-adjacent visionQwen3.8 or GLM-5.3-FlashNovel designs, difficult simulation, high-stakes reviewHYBRID est
Math/statisticsStepwise reasoning, code, symbolic checkingQwen3.8/Qwen3.6 plus a calculator or symbolic systemAdvanced proofs, exploratory research, persistent hard reasoningHYBRID; GPT-5.6 preferred for the hardest tier est
Physics/chemistry/biologyMath, diagrams, literature, code, long papersQwen3.8 or GLM-5.3-Flash for learning and preprocessingNovel synthesis, advanced research, complex chemistry or biologyHYBRID/FRONTIER est
Nursing/healthRetrieval, documents, case simulation, privacyApproved local model with curated retrievalComplex evidence review in an approved secure tenancyCONTROLLED HYBRID est
Accounting/financeMath, spreadsheets, tables, current documentsQwen3.8/GLM-5.3 for routine analysisComplex models, regulations, current markets, high-stakes reviewHYBRID est
Marketing/managementDrafts, presentations, research, analysisQwen3.6, Mistral Small 4, or Qwen3.8Large multimodal campaigns or deep market synthesisLOCAL-FIRST est
English/writingLong documents, revision, voice, feedbackMistral Small 4 or Qwen3.8Large-source synthesis or high-stakes publicationLOCAL-FIRST HYBRID est
History/philosophyLong documents, source comparison, argumentMistral Small 4/Qwen3.8 with retrievalLarge archives, difficult synthesis, current political factsHYBRID est
LanguagesMultilingual text, translation, speechMistral Small 4 for text; pilot actual UVU languagesReal-time speech, transcription, pronunciation, 70+ language translationHYBRID; Gemini speech stack est
Psychology/sociologyLong documents, qualitative coding, statisticsLocal secure model with retrievalSensitive-data research or complex literature synthesisCONTROLLED HYBRID est
EducationLesson creation, differentiation, accessibility, documentsQwen3.6/Mistral/Qwen3.8 are sufficient for routine workLarge multimodal curriculum packages or speech accessLOCAL-FIRST est
Art/design/mediaVision, layout, images, video, iterative reviewQwen3.8 or GLM-5.3-Flash for analysisNative generation, audio/video, polished multimodal artifactsHYBRID; Gemini media services where approved est
Music/dance/theatreAudio, video, scripts, translation, rehearsalLocal text models for scripts and planningSpeech, transcription, music/audio/video processingHYBRID; separate media stack est
Aviation/public safety/tradesManuals, current rules, scenarios, visionLocal retrieval system for instruction and searchComplex synthesis only; never autonomous operational controlLOCAL-FIRST HYBRID est

Model names reflect the current local portfolio (the Qwen, Mistral, and GLM classes the guide's §11 details) and current frontier services. Benchmark caution carried from the working papers: providers publish their own numbers under different scaffolds — no clean cross-vendor ranking exists, and a declared million-token context is a ceiling, not a promise fact. The upgrade rule stands: models are the cartridge, not the console — routes persist while names change quarterly.

Relative compute load per family

For capacity planning: relative text-token load per active user against a mixed-campus average of 1.0×, with practical working-context ranges est. Vision and speech are flagged separately — token counts understate their compute.

FamilyLoadWorking contextWhy
CS/IT/cyber1.8×32K–128KRepositories, logs, repeated debugging, tests, and many-turn sessions
Mechanical/civil/electrical engineering1.5×16K–64KMulti-step problems, standards, code, diagrams, and reports
Math/statistics1.4×8K–32KPrompts are often short, but reasoning and checking outputs are long
Physics/chemistry/biology1.5×32K–128KPapers, lab context, equations, data, diagrams, and scientific reasoning
Nursing/health1.2×16K–64KCases, study material, documentation, and retrieved clinical guidance
Accounting/finance1.4×16K–64KTables, spreadsheet logic, regulations, explanations, and revisions
Marketing/management1.3×32K–64KResearch packets, campaigns, cases, plans, and presentation iteration
English/writing1.5×32K–128KFull drafts, readings, multiple revisions, and detailed feedback
History/philosophy1.4×32K–128KLong readings, primary sources, argument maps, and comparison
Languages1.2× text8K–32KMany short translation and dialogue turns; speech cost is separate
Psychology/sociology1.3×32K–64KLiterature, qualitative records, survey design, and statistics
Education1.4×32K–64KLesson sets, differentiation, rubrics, examples, and feedback
Art/design/digital media1.2× text; about 1.8× compute-equivalent16K–64KImage and video input make token counts a poor compute proxy
Music/dance/theatre1.1× text8K–32KScripts and plans are moderate; audio/video processing is separate
Aviation1.0×16K–64KEpisodic manual, regulation, scenario, and technical-writing work
Culinary/trades0.8× text8K–32KShort instructional sessions with occasional image-heavy bursts
Public administration/public safety1.2×16K–64KPolicy documents, reports, scenarios, and current-source retrieval

The master table — all 51 units

UVU unitCollege/schoolDemand tier & basisPrimary usesRouteGoverning constraints
Allied HealthHealth & Public SvcHIGH est; health-program evidence, 2026Study aids, terminology, evidence search, case simulation, documentation practiceR4No patient identifiers; approved evidence; clinician/instructor signoff
Criminal JusticeHealth & Public SvcMED est; social-science and writing task fitCase comparison, policy, reports, scenarios, statisticsR4/R9Current law; sensitive records; bias and civil-rights review
Emergency ServicesHealth & Public SvcMED est; health plus safety-critical task fitScenario practice, protocols, reports, public educationR7No operational reliance; current approved protocols; no incident identifiers
Health SciencesHealth & Public SvcHIGH est; 83.2% health-education use, 2026Literature, study support, data, public-health materialsR4Privacy, evidence quality, no diagnosis
NursingHealth & Public SvcHIGH est; 12 uses/student/semester, 2025Exam prep, cases, communication, documentation practiceR4No PHI; clinical judgment remains human
Physician AssistantHealth & Public SvcHIGH est; clinical evidence baseCases, literature, differential-reasoning practice, patient education draftsR4No unsupervised diagnosis or treatment; source and faculty check
Public Administration Graduate ProgramHealth & Public SvcMED est; writing/policy/business task fitPolicy memos, budgeting, public records, stakeholder analysisR8/R9Current law and policy; records privacy; citation check
Public HealthHealth & Public SvcHIGH est; health, data, and communication task overlapEpidemiology support, data, literature, outreachR4/R3Population bias, privacy, current evidence, no clinical claims without review
Utah Fire and Rescue AcademyHealth & Public SvcMED est; vocational evidence, 2026Protocol search, scenarios, study aids, report practiceR7Approved manuals only; no live operational authority
CommunicationHumanities & Soc SciHIGH est; business/communication faculty use 88%Drafting, audience adaptation, media analysis, research, presentationsR1/R9Attribution, source verification, likeness and media rights
EnglishHumanities & Soc SciHIGH est; writing study, 2026Brainstorming, revision, feedback, close reading, rhetoricR9Preserve student voice and process evidence; verify citations
History and Political ScienceHumanities & Soc SciMED est; history direct evidence, 2025Primary-source comparison, timelines, argument, policy/current affairsR9Primary-source fidelity, current facts, bias, provenance
Integrated StudiesHumanities & Soc SciMED est; mixed task compositionCross-field synthesis, planning, writing, reflectionR9Apply each contributing course’s rules; avoid false cross-field certainty
Languages and CulturesHumanities & Soc SciMED est; 54.2% regular ChatGPT use, 2025Translation, conversation, grammar, pronunciation, cultural comparisonR5Dialect/cultural review; audio privacy; voice consent
Marriage and Family Therapy Graduate ProgramHumanities & Soc SciHIGH est; health plus psychology task fitRole-play, theory review, documentation practice, literatureR4No client data; no therapy; supervisor and ethics review
Philosophy and HumanitiesHumanities & Soc SciMED est; 54.5% philosophy-reading use, 2024Argument maps, counterarguments, reading support, writing feedbackR9Text fidelity, authorship, explicit uncertainty
Psychology and CounselingHumanities & Soc SciMED est; 54% psychology use, 2026Concept tutoring, study design, statistics, literature, role-playR4No client/participant data; no diagnosis or counseling
Social and Behavioral SciencesHumanities & Soc SciMED est; social-science proxies; sociology direct rate unknownQualitative coding, surveys, theory, statistics, literatureR4/R9IRB and participant privacy; bias; source verification
BiologyScienceMED est; 50% small-sample use, 2026Literature, bioinformatics, diagrams, lab preparation, tutoringR3Lab safety, dual-use review, source checking
ChemistryScienceMED est; grouped science evidenceCalculations, spectra explanation, literature, lab preparationR3Chemical safety and dual-use controls; verify units and methods
Earth ScienceScienceMED est; science plus visual/data task fitMaps, field notes, geospatial code, diagrams, literatureR3/R6Map/data provenance; field and hazard decisions remain human
Exercise Science and Outdoor RecreationScienceMED est; health and science task fitAnatomy, program-design practice, data, risk scenarios, public educationR4No personal health prescriptions; safety and privacy checks
Mathematical and Quantitative ReasoningScienceHIGH est; 91% math use, 2026Stepwise tutoring, alternative explanations, practice, checkingR3Protect independent skill formation; verify every calculation
MathematicsScienceHIGH est; same direct math evidenceProof critique, modeling, statistics, code, research preparationR3Formal validation; frontier required for hard research; AI is not proof
PhysicsScienceMED est; n=804 physics evidence, 2025Problem explanations, simulation, lab analysis, literature, diagramsR3Verify equations, assumptions, units, and experimental interpretation
Elementary EducationEducationHIGH est; preservice-teacher evidence, 2025Lesson plans, differentiation, examples, family communicationR10No minor data; developmental suitability; district and course rules
Graduate EducationEducationHIGH est; education output-creation share 74.4%Curriculum, research, leadership, policy, assessment designR10/R9Research citations, student data, and policy must be verified
Secondary and Special EducationEducationHIGH est; teacher-education direct evidenceSubject lessons, accommodations, rubrics, simulationsR10No IEP or minor identifiers; accessibility and disability-bias review
Student Leadership and Success StudiesEducationHIGH est; education and social-science task fitCoaching practice, study aids, student communication, program materialsR10/R4No student records; do not replace human advising
Art and DesignArtsMED est; design-student evidence, 2024Ideation, critique, visual analysis, layout, portfoliosR6Copyright, provenance, style appropriation, disclosure
DanceArtsMED est; direct rate unknownMovement analysis, rehearsal planning, grants, promotion, video reviewR6Performer consent, likeness, choreography rights
MusicArtsMED est; direct rate unknownTheory tutoring, transcription, rehearsal, program notes, promotionR5/R6Composition and recording rights; voice/performer consent
Theatrical Arts for Stage and ScreenArtsMED est; direct rate unknownScripts, dramaturgy, design, production planning, captionsR6/R9Script rights, performer likeness, disclosure and provenance
Applied Engineering and Transportation TechnologiesEngineering & TechHIGH est; engineering plus vocational evidenceDiagnostics, manuals, calculations, reports, code, instructionR7/R2Approved manuals; physical safety; no automated equipment control
Architecture and Engineering DesignEngineering & TechHIGH est; engineering and visual task fitDrawings, code checks, CAD assistance, specifications, presentationsR6/R7Licensed professional review; building codes; design provenance
Aviation ScienceEngineering & TechMED est; aviation course evidence, 2024Manuals, regulation study, scenarios, weather explanation, writingR7Current FAA/approved material; never operational authority
Computer ScienceEngineering & TechHIGH est; 38.6% Claude share, 2025Coding, debugging, architecture, tutoring, testingR2Course-specific authorship rules; secrets and sandbox controls
Construction TechnologiesEngineering & TechHIGH est; engineering/trades task fitEstimating, plans, codes, scheduling, safety instructionR7/R8Current code and manufacturer sources; field signoff
Culinary Arts InstituteEngineering & TechMED est; n=313 culinary evidence, 2024Recipe scaling, costing, substitutions, menu writing, instructionR1/R6Food safety, allergy verification, cultural attribution
Digital MediaEngineering & TechHIGH est; code plus multimodal task fitImage/video analysis, storyboards, prototypes, code, accessibilityR6/R2Copyright, likeness, training-data provenance, disclosure
Electrical and Computer EngineeringEngineering & TechHIGH est; Mines daily-user clusterCircuits, embedded code, signals, diagrams, reportsR2/R3/R7Validate calculations and hardware behavior; lab safety
Information Systems and TechnologyEngineering & TechHIGH est; CS/IT/cyber direct evidenceCode, databases, cloud, networking, business systems, cyber labsR2Authorized systems only; no secrets; sandbox offensive-security work
Mechanical and Civil EngineeringEngineering & TechHIGH est; ME direct evidence, 2024Calculations, simulation, code, drawings, standards, reportsR3/R7Engineer review; current standards; physical and public safety
Technology Management and MechatronicsEngineering & TechHIGH est; engineering, code, and business overlapControls, automation, code, operations, project planningR2/R7/R8Physical-system safety; isolated testing; management decisions remain human
AccountingWoodburyHIGH est; direct classroom evidence, 2024Reconciliation, explanations, spreadsheets, memos, standards researchR8Current standards; audit trail; confidential financial data
Air Force and Army ROTCWoodburyMED est; organizational type partly unknownPublic doctrine study, writing, history, leadership scenariosR7/R9No controlled, operational, or personal data; authority check
Business Graduate StudiesWoodburyHIGH est; 78% program integration, 2024Cases, strategy, finance, operations, research, presentationsR8Confidential employer/client data; current sources; decision audit
Finance and EconomicsWoodburyHIGH est; broad business 70% weekly useModeling, data, policy, markets, explanations, researchR8/R3Current market and regulatory data; no unchecked financial advice
MarketingWoodburyHIGH est; 27% of business syllabi added AI tasks, 2026Research, segmentation, campaigns, content, tests, mediaR1/R6/R8Brand control, copyright, privacy, bias, disclosure
Organizational LeadershipWoodburyHIGH est; management and writing task fitCases, coaching practice, communication, change plansR1/R8Personnel privacy; no automated employment decisions
Strategic Management and OperationsWoodburyHIGH est; business and quantitative task fitProcess maps, forecasts, cases, supply chains, presentationsR8Current data, confidential operations, human decision ownership

Unit names and college placement are fact from the 2026–27 catalog (each unit links to its catalog page); tier-basis links go to the underlying study. Two structural notes from the catalog sweep: Dental Hygiene and Respiratory Therapy nest under Allied Health; the ROTC unit's organizational type is partly unknown. Sequencing among colleges stays in Appendix A — this table is about what to build for whom, not who goes first.

E

What UVU already has

The complete public-record inventory — so nothing is bought twice, and nothing is asked twice

Appendix E in five lines

  1. QuestionWhat does Utah Valley University already have, and what should it not buy twice?
  2. AnswerThe public record shows an employee gateway, paid seats, a course helper, device management, computer nodes, and an active program office.
  3. Deciding numbers154 Copilot licensesfact51 ChatGPT licensesfact120 Artificial Intelligence Academy completersfact
  4. What the plan doesMeasure current seat use and reuse existing systems before adding another product or program.
  5. Still unknownCurrent seat use and prices, Canvas settings, internal contracts, and exact hardware are UNKNOWN.

Before this plan recommends a single purchase, it inventories everything UVU already owns, licenses, runs, or can reach — from board minutes, budget documents, IT service pages, procurement postings, and vendor announcements. The plan's stance toward all of it: assume none of it is available. Every existing system, seat, compute node, and staff hour is already allocated to the people using it today, so the service is designed to stand entirely on its own. This inventory exists so we know what we are building on top of — the integration surfaces (identity, LMS, device management, network, data rules) — and what we must never needlessly duplicate. It is a map of the house, not a list of things to borrow. This is a public-web inventory fact — it cannot see inside UVU's contract ledger, which is exactly why the ask-once packet at the end exists.

The headline assets

  • An employee AI Gateway is already live — launched July 1, 2026 after a 360-person pilot: ChatGPT, Claude, Gemini, and local open-source models behind one door, about $5 in tokens per employee per week. Students not yet included. Our service keeps its own front door for its own users and can offer the Gateway a local-engine hook — a connection point, not a dependency, and never a competing campus-wide portal pitch.
  • Paid seats already exist and are the experiment — 154 Copilot and 51 ChatGPT licenses deployed January 2026. Their usage data should decide seat expansion, not vendor pitches.
  • Wilson, UVU's own course assistant — 44 courses and ~2,000 students by February 2026, funded ($163k donated services + $96k Academic Affairs). Our service is not a course chatbot and does not replace it; a Wilson-style assistant can call our engine if UVU chooses.
  • The management layer is done — Jamf Pro already auto-enrolls every university Apple device. The Mac fleet this plan proposes needs zero new management software.
  • Compute today is CPU-shaped — three UVU-owned Notchpeak nodes (~224 cores, 680 GB RAM, no disclosed GPU) plus guest routes to shared Utah GPU systems (80 H200s on One-U, preemptible; the Redtail supercomputer exists but no UVU allocation is public). Shared, preemptible access is not the same as owned capacity — and none of it is assumed available to this plan, which is exactly why the plan owns its floor.
  • The program office exists and is funded — the Kahlert Applied AI Institute ($5.2M gift), AI Academy (120 completers), twice-monthly Power Hours, a faculty coaching pilot, curriculum grants, and a paid AI apprenticeship. New work routes through it.

Full inventory

GradeAssetScope and current evidenceWhat it means for the plan
FACTMicrosoft 365Office 365 is available to current students and employees. Basic Microsoft Copilot is included for all students and staff. A current UVU page says most employees have M365 A5 and receive Power BI Pro. UNKNOWN: the full A1/A3/A5 mix, especially for students.
Software for Students, AI Toolkit, and Power BI licensing, undated; accessed 2026-09-03.
Do not buy a second general office, collaboration, or basic chatbot suite without a proven need.
FACTPaid Copilot and ChatGPT seatsJanuary 2026 employee deployment: 154 Copilot licenses and 51 ChatGPT licenses. Copilot had grown 175% since July 2024. UNKNOWN: SKUs, assigned/active seats, utilization, prices, owners, or renewals.
Board of Trustees minutes, 2026-01-29, pp. 7–8.
The paid-seat model is already proven. Measure these seats before expanding or buying another product.
FACT; secondary detailUVU AI GatewayFree to all employees; no install required. Reported launch: 2026-07-01 after a 360-user pilot. Reported models: ChatGPT, Claude, Gemini, and local open-source models. Reported features include multi-model comparison, projects, knowledge bases, prompt chains, and agent export. Reported allowance: about $5 in tokens per employee per week, resetting Mondays. Students were not yet included.
UVU employee resources, accessed 2026-09-03; TechBuzz report using UVU executive statements, published 2026-07-14.
Treat the Gateway as the front door. Build governance, retrieval, local models, and routing behind it instead of creating a rival portal.
FACTWilson AI / Ask WilsonUVU-developed website and Canvas assistant. Fall 2025 board snapshot: 26 courses and more than 1,800 students. February 2026 CIO interview: 44 courses and 2,000 students. Current UVU page says “select” courses. UNKNOWN: the 2026-09-03 count, models, hosting, retention, and active-use rate.
Board minutes, 2026-01-29; CIO interview, 2026-02-11; Wilson AI, accessed 2026-09-03.
Expand and measure Wilson before buying a second general course chatbot.
FACTWilson development fundingEarlier development received $163,000 in donated services and $96,000 from Academic Affairs for FY2025. The pilot had 14 classes in the October 2024 budget presentation.
Digital Transformation PBA presentation, 2024-10-30, pp. 4 and 14.
Wilson is an existing funded product, not a concept-stage purchase.
FACT; EST for tier entitlementCanvas by InstructureCanvas is UVU’s only approved and supported LMS. Instructure says every current Canvas customer is now Canvas Core, with five included IgniteAI capabilities. EST: UVU therefore has Core entitlement. UNKNOWN: UVU feature enablement and any Plus/Next purchase.
UVU Canvas technology, accessed 2026-09-03; Instructure Q2 2026 tier announcement, April 2026.
Test and enable included Core functions before buying substitutes.
FACTCopyleaksUVU’s current OTL page lists Copyleaks among the supported course technologies. An indexed UVU guide said its plagiarism and AI detection tool was integrated with Canvas and available to all instructors. Faculty Senate recorded false-positive concerns. UNKNOWN: current term, cost, and detector configuration.
OTL, accessed 2026-09-03; indexed integration guide, published circa 2025, now returning 404; Faculty Senate agenda, 2025-04-29.
Do not buy a duplicate detector until Copyleaks’s accuracy, use, and renewal need are reviewed.
FACTCanvas-associated toolsTeams, Kaltura, Qualtrics, Honorlock, and OneDrive are approved. Honorlock is UVU’s official online proctor after the move from Proctorio. Approved publisher LTIs include Cengage, Macmillan, McGraw-Hill, Pearson, VitalSource, and WileyPLUS. UNKNOWN: AI add-ons and contract terms.
Canvas technologies and employee resources, accessed 2026-09-03.
Use the supported Canvas ecosystem. Do not assume publisher AI products are included merely because their LTIs are approved.
FACTZoom postureZoom is not approved or supported for employee or student use; Teams is UVU’s standard. An isolated purchase cannot be ruled out.
Zoom at UVU, accessed 2026-09-03.
Do not plan around Zoom AI Companion. Use Teams unless UVU changes its standard.
FACTAdobe Creative CloudAvailable to every active student, employee, and faculty member. UNKNOWN: Adobe Firefly entitlement, credit pool, admin controls, or paid premium features.
Adobe Creative Cloud Suite, accessed 2026-09-03.
Do not buy a duplicate creative suite. Confirm Firefly credits before buying a separate image service.
FACT / UNKNOWNGrammarlyThe free version is installed in open labs. UNKNOWN: any paid institutional Grammarly contract.
EUTS software inventory, accessed 2026-09-03.
Do not describe Grammarly as a paid campus license without internal proof.
FACTLucidchart and LucidsparkFree education upgrade for students and employees.
Software for Students, accessed 2026-09-03.
Reuse for diagrams and collaborative ideation before adding another whiteboard product.
FACTLinkedIn LearningFree access is publicly stated for students and employees. Board data records 161 AI-course uses in Spring 2025 and 301 in Fall 2025. UNKNOWN: current contract term, cost, and whether counts represent unique people.
Online Student Hub, accessed 2026-09-03; Board minutes, 2026-01-29, p. 11.
Use the existing catalog before buying generic AI-awareness training.
FACT / UNKNOWNPluralsightAt least some institutional licenses exist. The Art & Design page says unlimited access for all current students, faculty, and staff; the central learning page describes narrower Dx and Academic Affairs coverage.
Art & Design labs and Dx Learning, accessed 2026-09-03.
Confirm scope once. Do not budget new technical-course licenses until the conflict is resolved.
FACT / UNKNOWNGoogle services and GeminiUVU Status lists Google Suite as operational. Broad `MY.UVU.EDU` Google Workspace access closed on 2023-08-11, while curriculum-specific `UVU.EDU` use may remain. Gemini is proven through the employee Gateway, not as a separate campus-wide license.
UVU Status, checked 2026-09-03; Google Transition, closure 2023-08-11.
Treat Gateway Gemini as the proven route. Confirm any remaining Workspace or Gemini Education contract before relying on it.
FACTMicrosoft storageOneDrive and SharePoint are official storage for all faculty, staff, and students. Box and Dropbox have been discontinued for most users. Planned limits effective 2026-12-31 are 100 GB per person and 500 GB per Teams/SharePoint site.
File Storage, accessed 2026-09-03.
Build governed AI retrieval against Microsoft stores. Do not center new AI work on Box or Dropbox. Account for the coming quotas.
FACTCivitas LearningUVU’s predictive student-success platform: Course Insights, Persistence Insights, Completion Insights, and Inspire. Access is available by request to relevant faculty, staff, administrators, advisors, and support teams.
Civitas access article, published 2023-10-19; current page checked 2026-09-03.
Reuse existing risk and persistence analytics rather than creating a second student-risk model without need.
FACT / UNKNOWNBI and data platformsCurrent public evidence supports Power BI, Power BI Pro for most A5 employees, Tableau dashboards, Business Objects, MyDataHub, a university data warehouse, a data lake, and renewed Collibra. Canvas, Jira, finance-budget, and EvolveFM data had been moved into the data lake by October 2024. UNKNOWN: lake vendor and current architecture.
Power BI, Tableau dashboard example, and Institutional Reports, accessed 2026-09-03; Dx PBA, 2024-10-30.
Start from the governed data stack. Do not assume Databricks or Snowflake is needed.
UNKNOWNDatabricks and SnowflakeExact official-domain searches found course, résumé, and employment references but no institutional deployment.
Public-web search completed 2026-09-03; no qualifying UVU deployment source found.
Do not list either platform as an owned asset or proposed need until the internal architecture is checked.
FACTEnterprise software estate and funding contextDx reported 211 maintained titles and a $5,998,301 software-expense snapshot in October 2024. It requested $600,000 ongoing for one year of M365 after five prepaid CARES-funded years; that figure is a request, not verified final spend. Final FY2025 allocations included $150,000 for AI Institute startup and a bundled $200,000 for background checks, LMS, AwardCo, and LinkedIn Learning increases.
Dx PBA, 2024-10-30; FY2025 final allocations, FY2025.
The public list is only a small part of UVU’s software estate. Use the internal catalog and contract ledger for purchase checks.
FACTJamf ProMDM for university-owned Macs, iPhones, iPads, and Apple TVs. New university Apple devices auto-enroll; enrollment is mandatory. UNKNOWN: device count and coverage rate.
Jamf Pro, accessed 2026-09-03.
No separate Apple MDM purchase is needed.
FACTLabs and MyAppsPhysical labs provide Adobe, Microsoft, SPSS, creative software, and other tools. MyApps gives all students virtual access to Adobe, ArcGIS, Avid, SPSS, NVivo, Power BI Desktop, RStudio, SAS, Stata, VS Code, Mathematica, and other software.
EUTS software inventory, accessed 2026-09-03.
Pilot software and modest models through existing labs/virtual delivery before buying a new general lab.
FACT / UNKNOWNApple labsThe Fall 2026 Art & Design page lists six Apple labs; EUTS lists four rooms and a different fourth-room number. Models, station counts, Apple silicon, RAM, and GPUs are not public.
Art & Design labs and EUTS software, accessed 2026-09-03.
The Mac fleet exists, but a hardware census is needed before sizing local AI work.
FACT; EST totalsThree UVU-owned Notchpeak CPU nodesOne 96-core/380 GB node and two 64-core/150 GB nodes. UVU account holders receive priority submission through SLURM. EST: 224 cores and 680 GB RAM total. No GPU is disclosed on these nodes.
UVU High Performance Computing Resources, undated; accessed 2026-09-03.
Use this capacity for CPU preprocessing, simulation, inference, and teaching. Do not count it as GPU capacity.
FACT / ESTShared CHPC GPU accessCHPC supports UVU accounts and has general GPU resources that do not require allocations. One-U Responsible AI has ten nodes with eight H200s each; all CHPC users may submit guest/preemptible jobs. EST: active UVU CHPC users can use guest capacity. Priority eligibility for an external UVU group is unknown.
CHPC accounts, updated 2026-08-26; CHPC GPU guide, updated 2026-06-04; One-U AI resource notice, January 2026 issue.
Benchmark queue time and usable GPU hours before buying hardware. Shared/preemptible access is not the same as a dedicated cluster.
FACT statewide; UNKNOWN for UVURedtail supercomputerStatewide system: 33 nodes, 264 NVIDIA H200 SXM5 GPUs, 3,696 CPU cores, 66 TB RAM, and about 1 PB usable scratch. Early-access proposals were open to USHE faculty and staff. No public UVU allocation was found.
Redtail, updated 2026-07-08; launch report, 2026-07-23.
Confirm whether UVU received access before including Redtail in a capacity plan.
FACT / EST / UNKNOWNSmith Engineering BuildingThe nearly 200,000-square-foot building opened 2026-01-22. Post-opening evidence confirms a large drone lab with cameras, workstations, and drones. Pre-opening documents name AI, VR, computer-science, manufacturing, and prototyping spaces. A brochure planned a 30-person interactive AI classroom. Installed GPU and workstation details are unknown.
Ribbon cutting, 2026-01-22; building preview and named-spaces brochure, 2025-04-14/undated.
Use the building’s real labs, but do not claim a GPU cluster until the as-built inventory is supplied.
FACT service / UNKNOWN hardwareSCaFL and Business Resource Center AI labCurrent page advertises AI inference training, simulation, big-data/data-science work, fabrication, and a UTOPIA “10GB” connection. Independent use requires consultation, certification, and fees. Its former server-detail page now returns 404.
SCaFL, accessed 2026-09-03.
Reuse the lab for pilots after confirming server specifications and availability.
FACTNVIDIA partnershipThree-year voluntary collaboration with no exchange of funds. Includes DLI ambassador training, teaching kits, workshops, materials, advanced tools, and GPU-accelerated cloud workstations used for DLI activity. UNKNOWN: persistent credits, GPU hours, cloud tenant, research allocation, or donated hardware.
NVIDIA announcement, 2025-03-10; UVU partnership page, accessed 2026-09-03.
Use the training and workshop rights. Do not book “NVIDIA cloud credits” as compute capacity until quantified.
FACTKahlert Applied AI Institute funding$5.2 million Kahlert Foundation gift announced 2025-10-07. Earlier final PBA allocation supplied $150,000 in startup funds.
UVU-issued gift announcement, 2025-10-07; FY2025 allocations.
Use the funded institute as the central program office instead of creating a parallel AI office.
FACTAI Academy and trainingBoard snapshot: 120 AI Academy completers, 22 Advanced AI participants, and more than 3,000 AI learning opportunities. AI in Action counts were 349, 204, 484, and 1,875 for Fall 2024 through Fall 2025.
Board minutes, 2026-01-29, pp. 0 and 11.
The base training layer exists. Focus new spend on role-specific practice and measured outcomes.
FACTPower Hours and faculty coachingPower Hours run twice monthly for faculty and staff. The faculty coaching pilot is free and confidential, uses three sessions per semester, and is limited to a small founding cohort.
AI Institute and faculty coaching, accessed 2026-09-03.
Scale the existing coaching model if results support it; do not buy generic coaching first.
FACT2026 Faculty Summer Institute AI trackMay 5 build lab, two-week sprint, May 29 showcase, and $1,500 stipend after deliverables. Faculty built reusable course agents and AI-ready assignments.
Applied AI Fluency in the Classroom, 2026 program.
Reuse its templates, deliverables, and trained faculty.
FACT lower bound / UNKNOWN totalCurriculum Integration GrantsAt least two distinct 2026 awards are public. One Cao/Tang project is explicitly $5,000; Devin Gilbert lists a separate award with no public amount. Total awards, dollars, courses, and results are unknown.
Cao/Tang profile and Gilbert CV, accessed 2026-09-03.
Treat two projects as the proven floor, not the full grant count.
FACTApplied AI ApprenticeshipPaid, registered, full-time, 12 months; current partners SchoolAI and AskElephant. First cohort began August 2025. Completers earn industry certifications and a stackable business-analysis certificate. UNKNOWN: cohort size, completions, placements, and Clarion’s current role.
AI Apprenticeship, accessed 2026-09-03.
Expand the existing employer pathway rather than creating a separate apprenticeship structure.
FACTAcademic AI programsCurrent offerings include a 30-credit MS in Applied AI, an 18-credit AI Graduate Certificate, two 18-credit undergraduate AI certificates, and a 120-credit Information Systems BS with an Applied AI emphasis.
MS, graduate certificate, Applied AI certificate, Applied AI in Organizations, and BS emphasis, 2026–27 catalog.
New credentials should fill a measured gap, not duplicate this stack.
FACTUNESCO Chair and AI committeeUNESCO Chair on AI and Environmental Stewardship for Sustainable Futures, established in 2025, Chair ID 2025US3042. Current UVU AI and Education committee has 12 named members.
UNESCO registry, 2026-01-12; UVU committee, accessed 2026-09-03.
Use this as an ethics, education, sustainability, and international network—not as evidence of funded compute.
FACT announcement / UNKNOWN deliveryUSHE AI Task Force and statewide credentialBarclay Burns is one of 11 listed task-force members. USHE announced a free AI Workforce Credential for more than 50,000 eligible 2025–27 graduates, beginning 2026-07-01. A June 2026 state briefing still described an August RFP. No public launch, platform, enrollment, or UVU delivery record was found.
USHE announcement, 2026-05-01; legislative briefing, 2026-06-17.
UVU students may be eligible, but do not count UVU as the platform or credential operator until confirmed.
FACT / UNKNOWNAI and data governancePolicy 445 governs institutional data. Policies 446, 447, and 452 cover privacy, security, and accessibility. Policy 441 on AI was proposed in 2025 but is absent from the approved manual as of the cutoff.
Policy 445, effective 2024-03-28; Policy Manual, checked 2026-09-03; Policy 441 executive summary, 2025-09-25.
Build on the approved data controls. Confirm the binding AI standard before adding another governance layer.

Canvas: what the LMS tier already includes

Instructure moved every current customer to Canvas Core, which includes five AI functions at no extra cost — but inclusion is not enablement, and the free preview of the premium tier ended June 30, 2026. First move: turn on and test what's already paid for.

Canvas capabilityTierAvailability & controlUVU state
IgniteAI Summaries for DiscussionsCoreGenerally available 2026-03-21; disabled by default; admin opt-in.UNKNOWN: no public UVU enablement proof.
IgniteAI Search for Courses, formerly Smart SearchCoreInstitution opt-in; semantic course-content search; admins may delegate teacher control.UNKNOWN: no public UVU enablement proof.
IgniteAI Translations for Discussions, Inbox, and AnnouncementsCoreIncluded for all tiers; feature control remains with institution/educator.UNKNOWN: no public UVU enablement proof.
Question Authoring Assistance for QuizzesCoreIncluded without an upgrade.UNKNOWN: no public UVU enablement proof.
Accessibility Remediation for CoursesCoreIncluded in the Course Accessibility Checker.UNKNOWN: no public UVU enablement proof.
Discussion Insights, Rubric Generator, and Grading AssistancePlus or NextPaid higher tier.UNKNOWN: no public proof UVU has Plus or Next.
IgniteAI Agent, Ask Your Data, and Study ToolsNextPaid highest tier.UNKNOWN: no public proof UVU has Next.

The gap ledger — "already have, therefore…"

Read the "therefore" column as a duplication guard, not a borrowing plan: it says what we must not pitch or buy twice. The service still brings its own front door, engine, operations, and budget.

Already haveTherefore the plan does — or does not — need
FACT: M365 and basic Copilot for all students/staff; AI Toolkit, accessed 2026-09-03.No second broad productivity suite. Start with adoption, role fit, and data boundaries.
FACT: 154 paid Copilot and 51 ChatGPT seats; Board minutes, 2026-01-29.No automatic seat expansion. First obtain assignment, use, renewal, and outcome data.
FACT: Employee AI Gateway; UVU employee page, accessed 2026-09-03.No competing campus-wide portal pitch. The service keeps its own front door for its own users and offers the Gateway a local-engine hook.
FACT: Wilson AI reached 44 courses/2,000 students; CIO interview, 2026-02-11.No generic course-chatbot purchase until Wilson’s measured limits are known.
EST: Canvas Core includes five AI functions; Instructure, April 2026.Enable and test summaries, search, translation, quiz assistance, and accessibility remediation before buying duplicates.
FACT: Copyleaks is supported in the Canvas environment; OTL, accessed 2026-09-03.No second detector without an accuracy or contract reason. Detection results should not be treated as sole proof of misconduct.
FACT: Teams is standard; Zoom is not approved; Canvas technology, accessed 2026-09-03.No Zoom AI Companion plan. Use Teams and its included capabilities.
FACT: Adobe Creative Cloud is campus-wide; Adobe KB, accessed 2026-09-03.No replacement creative suite. Confirm Firefly credits and controls first.
FACT: Power BI, Tableau, Civitas, Business Objects, MyDataHub, a warehouse/data lake, and Collibra exist; Dx PBA, 2024-10-30.No assumed Databricks/Snowflake purchase. Map the current stack and missing workload first.
FACT: OneDrive/SharePoint are official; Box/Dropbox are mostly discontinued; File Storage, accessed 2026-09-03.Put governed retrieval on Microsoft stores, not on a new Box/Dropbox knowledge layer.
FACT: Jamf manages university Apple devices; Jamf, accessed 2026-09-03.No Apple MDM purchase. Use Jamf inventory to find AI-capable Macs.
FACT: Physical labs and MyApps already deliver software; EUTS, accessed 2026-09-03.Pilot on the existing fleet before buying a general AI lab.
FACT: Three dedicated CPU nodes plus general CHPC access; UVU HPC, accessed 2026-09-03.Use CPU nodes for preprocessing and suitable work; measure shared GPU queues before GPU capital spending.
EST: UVU users can reach preemptible One-U H200 capacity; CHPC, January 2026.The plan still may need dedicated GPU capacity if queueing, privacy, or uptime requirements fail; shared access is not ownership.
UNKNOWN: Redtail allocation; Redtail, updated 2026-07-08.Confirm UVU’s allocation before treating Redtail as available capacity.
FACT: NVIDIA training and cloud-workstation access for DLI activity; NVIDIA, 2025-03-10.Use activated DLI benefits. Do not count unspecified “cloud credits” in capacity planning.
FACT: AI Academy, Advanced AI, AI in Action, LinkedIn Learning, Power Hours, coaching, and faculty institutes exist; Board minutes, 2026-01-29.No new generic AI literacy catalog. Fund role-based practice, coaching capacity, and outcome measurement.
FACT: $5.2 million gift, startup funds, grants, and apprenticeship structure exist; gift release, 2025-10-07.No parallel AI program office. Route new work through the institute and disclose how current funds are committed.
FACT: Data governance, privacy, security, and accessibility controls exist; Policy Manual, checked 2026-09-03.No duplicate governance framework. The missing item is a confirmed binding AI policy and operating decision record.
FACT: UNESCO and USHE relationships exist; UNESCO, 2026-01-12; USHE, 2026-05-01.Use them for standards, reach, and coordination. Do not count them as software, funding, or compute without a specific award.

The full data ledger — deferred diligence, not a precondition

The twelve items below were the original public-record gaps. A second public sweep (Appendix K) closed most of what matters for building and reduced the real ask to five interface questions. This ledger is kept for later commercial diligence — request it as an export when useful, never as a gate on the pilot.

  1. Software and spend: Export UVU’s current AI-relevant license ledger with product, SKU, vendor and reseller, contract/PO, owner, funding unit, term, renewal, purchased seats, assigned seats, 30/90-day active users, and annual cost. Include OpenAI, Anthropic, Microsoft, Google, Adobe, Instructure, Copyleaks, Grammarly, Turnitin, Zoom, Qualtrics, Kaltura, Honorlock, LinkedIn Learning, Pluralsight, Tableau, Civitas, ServiceNow, Salesforce, AWS, Azure, GCP, Databricks, and Snowflake.
  2. Microsoft: Confirm the current student and employee A1/A3/A5 mix; the exact 154 Copilot SKU; the 51-seat ChatGPT tier; Power BI Pro/Premium capacity; Copilot Studio/Azure OpenAI access; and the outcome of the October 2024 $600,000 M365 request.
  3. Gateway: Provide its current model/provider list, hosting diagram, API contracts, weekly-budget rule, total and per-model spend, paid-tenant rules, active users, student-launch decision, retention, provider-training terms, logging, data classes, and approved connectors.
  4. Wilson AI: Confirm the 2026-09-03 course and section count, unique students, active users, current name, models/providers, retrieval and Kaltura design, hosting, cost, retention, provider-training terms, assessment results, and whether the Faculty Senate oversight group was formed.
  5. Canvas: Confirm Core/Plus/Next contract status, term and cost; current feature flags for all IgniteAI functions; data-processing terms; Wilson, Copyleaks, Honorlock, Kaltura, and publisher LTI status; and whether Khanmigo or another tutoring pilot exists.
  6. Compute: Confirm all three UVU Notchpeak nodes are online and provide CPU model, hostnames, usable cores/RAM, storage, partition/QoS, priority terms, and FY2026 use. Confirm UVU eligibility and limits for general GPUs, One-U H200 priority, and any Redtail allocation.
  7. Campus hardware: Supply the current Jamf and IT asset export for Macs, GPU workstations, servers, Citrix GPU profiles, Smith Building AI/VR/drone/robotics rooms, SCaFL servers, DGM/XR equipment, and any surviving departmental GPU servers. Include model, count, room, owner, access rule, and operational state.
  8. NVIDIA: Provide the executed or redacted MOU, exact term, activated deliverables, current ambassadors, completed certifications/workshops, learner counts, cloud tenant, instance type, GPU hours, dollar credits, expiry, and any donated equipment.
  9. People and funding: Give one dated roster for AI Academy, Advanced AI, Power Hours, coaching, Summer Institute, curriculum grants, and apprenticeships. Include unique people, completions, outcomes, every grant/project/amount, apprenticeship employers and participants, credentials, placements, and the committed versus uncommitted Kahlert budget.
  10. Data and agreements: Confirm the warehouse/data-lake architecture and current roles of Collibra, Power BI, Tableau, Business Objects, MyDataHub, Civitas, Databricks, and Snowflake. List all executed AI MOUs, DPAs, DUAs, and data-sharing agreements with data classes, processors, retention, model-training use, IP, and expiry.
  11. State and UNESCO: Confirm whether the AI Workforce Credential actually launched, who owns the platform, UVU’s delivery role, enrollment/completions, and RFP status. Provide the current UNESCO Chair and committee charter, funding, partner agreements, and AI/data projects.
  12. Public-payment reconciliation: Produce UVU vendor payments for FY2024–FY2026 for the named vendors, including payments through VLCM, CDW, SHI, Carahsoft, cooperative contracts, marketplaces, and other resellers. Reconcile them to contracts and the license ledger.

Search coverage note: procurement search fields showed no events for the major AI vendors — which proves only that the searchable fields are empty, not that no purchases exist (reseller, cooperative, and low-value purchases don't appear there). The state transparency portal's vendor search loads through a client-side form that automated public access could not operate; the reconciliation request above covers it fact.

F

The free-and-discounted catalog

Everything worth claiming before spending a dollar — with the catch printed next to each

Appendix F in five lines

  1. QuestionWhat can Utah Valley University claim before it buys more?
  2. AnswerUse the $0 programs first, get a paid-seat quote in week 1, and let the small pilot set the hardware need.
  3. Deciding numbersMicrosoft Copilot Chat $0factpilot Mac $2,139factseat benchmark $1.60/mofact
  4. What the plan doesClaim the $0 rows, seek research awards for later growth, and let measured pilot use set the owned hardware size.
  5. Still unknownA binding university seat quote is NOT_RUN, and several program amounts, access rules, and renewal terms are UNKNOWN.
ProgramWhat UVU getsCostThe catchGrade
Microsoft Copilot ChatEnterprise-protected chat for everyone on the existing M365 A-license — already deployed at UVU$0No campus control of models or data locality; per-user paid tier is the upsell pathfact
Google Gemini for EducationGemini with enterprise data protection free on an edu tenant; Utah's K-12 system already adopted statewide (June 2026)$0Requires tenant setup and policy review; free tier terms can change at renewalfact
OpenAI researcher accessFrontier-model seats for named researchers (institutional program, ~5 per institution)$0Named individuals only — not a campus servicefact
Anthropic science programScientist seats plus API credits for approved research projects — published ceiling up to $50k credits per project$0Application-gated, project-scoped, renewal not guaranteedfact
NAIRR pilotNational AI Research Resource compute allocations for research projects$0Competitive applications; batch research compute, not a campus assistantfact
NVIDIA–UVU agreementTraining, certification, tools, and cloud resources under the existing 3-year no-funds MOU$0No hardware included; workforce/curriculum scopefact
Apple education pricing on NASPO ValuePointMini M5 Pro 48GB at $2,139 (vs $2,299 retail) on the competed Utah PA4282 cooperative contract — no new bid required through June 2027−7%Education price verified for the pilot config; volume quotes may improve it (ask)fact
Utah AI Moonshot$5M state program; UVU's research office lists it with October deadlinesgrantNotice of intent was due Aug 28 — first question for Dr. Burns: was one filed?fact
HERFP research fund$45M pool; HB 373 reserves 15–25% of research money for regional institutions like UVUgrantCompetitive; proposal work requiredfact
HB2 AI compute ($15M)Statewide AI data-center appropriationpartnerSits under University of Utah Item 81 — UVU access is a partnership conversation, not an entitlementfact
Hosted open-weight APIsBurst overflow on open models (gpt-oss-120b class) at commodity metered rates — roughly $0.17 per user-year at campus scale in our demand modelmeteredZero-data-retention terms must be verified in writing per providerest
Negotiated edu seats (benchmark)Cal State's renewal: ~$19.26/user·yr at 675k users; Colorado ~$20; Maine ~$22.64 per billed FTE$1.60/moBenchmarks, not quotes — UVU-scale pricing must be quoted in week 1 (the plan's F1 branch takes the deal if it beats the floor)fact

Education and research programs — the wider sweep

A second pass on September 3 checked every education-specific program a campus could claim, including several missed above. Personal benefits (a student's own Azure credit, a verified student's free GitHub Copilot) are real and worth publicizing; none is campus capacity.

ProgramCurrent benefitLimit or catchCampus value
Google Gemini for EducationNo added charge through qualifying Education Fundamentals accounts; institution controls and enterprise protections. ProgramProduct limits remain; not equivalent to an unrestricted Gemini API pool.High. Use as a primary entitlement.
Microsoft Copilot ChatNo added charge for Microsoft 365 A1/A3/A5 faculty, staff and higher-ed students 13+. Current education pageFull Microsoft 365 Copilot is $18/user/month; agents and Graph integration differ.High. Use existing tenant controls.
GitHub Copilot StudentFree to verified students; teachers and some maintainers can receive Copilot Pro. EligibilityCoding assistant, not a general campus inference API. Eligibility checked regularly; automatic model selection applies.High for coding courses.
OpenAI Academic ResearchersUp to five managed seats for 12 months; business protections and Pro-level limits. FAQ, updated Aug. 2026Selective; waitlist; no API credit.Research teams only.
OpenAI Researcher AccessUp to $1,000 API credits for 12 months; quarterly review. Program FAQResponsible-AI research, not operations.Apply project by project.
OpenAI Codex for StudentsVerified U.S./Canada university students receive $100 in Codex credits. Terms, updated 2026-09-02Personal ChatGPT/Codex credit, not general API credit.Useful for student coding.
Anthropic Claude for EducationInstitution plan with Learning Mode, SSO, training and faculty research API credits. ProgramPaid; public seat price and API-credit amount unknown.Procurement candidate, not free entitlement.
Anthropic AI for ScienceNewer 2026-08-27 announcement offers up to $50,000 per project. AnnouncementSelective scientific research. Older help material still says $20K.Strong grant target.
Anthropic External Researcher AccessTypically $1,000 API credit for qualifying safety/alignment research. ProgramSelective; not operational capacity.Apply when research fits.
Perplexity Education ProVerified students/educators: $10/month; institutional Enterprise Pro: $300/user/year. Pricing, updated 2026-09-02Discounted, not free. Personal plan lacks institutional controls.Optional personal benefit.
Notion EducationFree Plus workspace for individual students/teachers; verified student organizations can receive shared workspaces. Education planNotion AI is only an undisclosed limited trial on Free/Plus/Education.Productivity benefit, not AI capacity.
NVIDIA Teaching KitsFree teaching material; approved courses can receive codes worth up to $90/course/student for labs. Teaching KitsCourse-specific approval; not an unrestricted hosted-NIM grant.Good for GPU/AI courses.
NVIDIA Inception/DGX creditsInception is free; qualified startups have seen offers up to $100K DGX Cloud credit.Startup program, not a university program. No public general renewal.Only for qualifying spinouts.
AWS Educate/AcademyFree learning content and Academy labs. Research-credit proposals accept faculty and students. Research FAQStudent research awards capped at $5,000; faculty cap not public. Credits generally expire in one year. Educate no longer provides the old blanket cloud credits.Course labs and funded research.
AWS general Free Tier$100 at signup plus up to $100 earned.New customers; free-plan account lasts at most six months.Onboarding only.
Azure for Students$100 for 12 months, no card, renewable annually while eligible; Azure OpenAI is within the service catalog. OfferOne account/customer; education, teaching and noncommercial research only; nontransferable.One of the strongest legitimate personal cloud benefits.
Google Cloud teaching creditsUp to $100 per teaching staff member and $50 per student; generally usable for 12 months after course start. Faculty programCourse-specific, nontransferable. Google’s general $300 trial cannot fund Gemini Developer API charges after March 2026.Good course sandbox.
Google Cloud research creditsFaculty researchers and PhD candidates may apply.Amount unknown; project award, not campus entitlement; expires after about one year.Apply per project.
Hugging Face Academia HubSSO/admin/audit plan with $2 monthly compute credit per seat. ProgramPaid: starts at $10/seat/month, minimum 250 annual seats.Possible managed teaching/research hub.
Hugging Face ClassroomsFree teaching workspaces and resources. ClassroomsCompute/API allowance unknown.Course organization, not capacity promise.
NAIRRSelective access to federal and industry compute, data, models and training. NSF NAIRRAllocation-based; not an automatic university entitlement.Continue pursuing for research workloads.

Order of operations: claim the $0 rows first (they kill the "we can't afford anything" objection and generate the adoption data every later decision needs), quote the seat benchmarks in week 1, and let the measured pilot decide how much owned floor to build. The research-grant rows fund scale, not the pilot — the pilot is deliberately small enough to need no grant.

G

The contingency playbook

28 pre-decided branches — so no failure requires an improvised meeting

Appendix G in five lines

  1. QuestionWhat does the plan do when a main price, speed, or use test fails?
  2. AnswerEach failure has a set response: switch what is bought, stop adding computers, or freeze growth.
  3. Deciding numbers≤$30/user·yrestunder 70% of projected speedestunder 10% active after 30 days and under 20% return at month 3est
  4. What the plan doesUse seats if they pass the price and safety gates; stop growth and compare other paths if speed or use misses.
  5. Still unknownThe hardware-speed and user-return checks are NOT_RUN.

The architecture dossier carries a full branch matrix: eight bureaucratic (B1–B8), seven technical (T1–T7), four financial (F1–F4), five demand (D1–D5), and four external (E1–E4) branches, each with a trigger, a pre-decided response, and an owner. Eight weekly tripwires watch the triggers. The three highest-leverage pre-decisions, in full:

BranchTriggerPre-decided response
F1 — the quote flipA binding, all-in seat quote at ≤$30/user·yr that passes security, accessibility, retention, and exit gatesTake it. Seats become the faculty front door for Public-tier work; the owned fleet contracts to a 4-Mac Sensitive-data enclave plus evaluation and teaching. Gateway, governance, and adoption work carry over unchanged. The plan is not loyal to hardware — it is loyal to the floor price.
T1 — the benchmark missShipped M5 minis measure under 70% of projected throughput on two independent runsStop at 4 boxes. Run a same-corpus bake-off: Studio tier vs Linux/NVIDIA vs hosted APIs vs seats. Lowest passing 3-year cost wins. No sunk-cost defense — $12.7k is the price of knowing.
D2 — the adoption floorUnder 10% 30-day active and under 20% return rate at month 3Freeze scaling. Twenty user interviews, two funded department workflows, re-measure at month 6. If still under, end the general service and keep the enclave — the honest outcome beats a zombie service.

Five decision memos are pre-written for the slower-burning cases: a major model release, a central-IT takeover offer, agent/automation traffic growth, retrieval workloads crushing prefill, and one department dominating usage. The full matrix travels with the architecture dossier rather than this page — it is an operating document, not a pitch.

H

Open items & the ask-once list

What we could not verify from outside — each asked precisely, once

Appendix H in five lines

  1. QuestionWhat is still open, and who must answer it?
  2. AnswerThe pilot can start, but campus growth waits on room, service, insurance, access, records, money, sponsor, and buying decisions.
  3. Deciding numbersfour decisions only Dr. Burns can makeestfive technical questions for the information technology teamest$30 per user per yearest
  4. What the plan doesAsk each question once, name an owner, and close the needed items before buying, first use, or campus growth.
  5. Still unknownRoom approval is absent; insurance and the access report are UNKNOWN; records work is shallow; Dr. Burns’s answers remain open.

A plan that hides its unknowns isn't one. These are the items public evidence could not settle unknown, each with its next action and its owner-to-be. Nothing on this list blocks the validation pilot; several block campus scale.

ItemStateNext actionOwner-to-be
Facilities, network, and continuity sign-off for the pilot roomAbsentOne-page named-room sign-off before the purchase orderFacilities + Dx
Interim service terms (acceptable use + incident taxonomy)AbsentCounsel-approved one-pager before the first user (Maine and UT-Austin templates exist)Counsel + service owner
Insurance posture (property, cyber, E&O for AI advice)UnknownWritten confirmation from risk managementUVU risk office
Asset accounting, tagging, and disposal pathMostly policy-resolvedExpense code + custodian + retirement path named in the purchase packetController
Chat-platform accessibility report (VPAT/ACR, WCAG 2.1 AA per Policy 452)Unknown availabilityDemand the exact report and version; test the deployed build regardlessAccessibility office
Records retention / GRAMA / legal holds vs user deletion rightsShallowRetention-table reconciliationRecords officer + counsel
Concurrent-enrollment minorsExcluded by designSeparate consent and safety design before any phase-2 inclusionProgram council
Utah AI Moonshot / HERFP notice-of-intent statusUnknownAsk in week 1 — deadlines are OctoberDr. Burns

The four decisions only Dr. Burns can make — asked once, here, so no one asks twice (the five technical questions for UVU's IT side are in Appendix K):

  1. Was a Moonshot or research-fund notice of intent filed before the August 28 deadline?
  2. Is there a budget number — or should the simulator's dial stay the planning instrument?
  3. Which college sponsors the pilot, and does Dx co-sponsor identity and security review from day one?
  4. If a seat vendor beats $30 per user per year all-in (branch F1), should the flip be taken automatically or brought back for a decision?
I

The campus map data

Buildings, college homes, regions, and imagery rights behind the planning map

Appendix I in five lines

  1. QuestionWhat place data and image rights make the campus map safe to use?
  2. AnswerUse the checked college homes and building pins with public federal aerial photos, not Utah Valley University’s map art.
  3. Deciding numbers53 building pinsestmedian offset near 19 metersfactone flagged pin ~880 m from its mapped buildingfact
  4. What the plan doesPut each college at its checked home, show split sites, and keep the federal photo credit on the map.
  5. Still unknownExact roles at four outlying sites and Payson’s construction schedule are UNKNOWN.

The planning map plots colleges on real overhead imagery. This appendix is its ground truth: every building pin on UVU's live campus map, the verified college→building mapping, every UVU location beyond the Orem core, and the licensing that makes the imagery reusable. Coordinates are UVU's own map pins rounded to four decimals — label and entrance points, not surveyed centroids (a footprint cross-check put the median offset near 19 meters) fact.

Where each college lives

Verified against deans' offices, advising pages, and department locations — not assumed from building names. Two findings worth pinning: the engineering core moved to the new Smith building in January 2026 (older maps still label the CS building as its home), and Health & Public Service is genuinely split — administration in the Losee Center, health teaching across the pedestrian bridge on West Campus, public safety in EN.

College / schoolBldgRoleEvidence
Smith College of Engineering & TechnologySEPRIMARY; dean, Computer Science, Digital Media, Mechanical/Civil Engineering, Electrical/Computer Engineering, labs factSE building page, 2026 opening
Smith College of Engineering & TechnologyCSTRANSITIONAL_SECONDARY; some advising/support pages still use CS after the SE move estCET advising, college directory
Smith College of Engineering & TechnologySASATELLITE; automotive and transportation technology factTransportation Technology
Woodbury School of BusinessKBPRIMARY; dean, advising, Accounting, Finance, Economics, Marketing, Management and Operations factWoodbury directory, advising
College of Humanities & Social SciencesCBPRIMARY; dean, advising and all principal department offices factCHSS directory, CHSS advising
School of EducationMEPRIMARY; expressly identified as the school’s main hub for dean/faculty, advising and classrooms factEducation locations
School of EducationNBSATELLITE; autism classrooms, outreach and research factEducation locations
School of EducationLCSATELLITE; Student Leadership and Success Studies factEducation locations
College of ScienceSBPRIMARY_ADMIN; dean and Biology factCollege overview, Biology
College of SciencePSPRIMARY_TEACHING_LABS; Chemistry, Earth Science, Physics, advising and planetarium factChemistry, Planetarium
College of ScienceLASATELLITE; Mathematics factMathematics
College of ScienceRLSATELLITE; Exercise Science factExercise Science
College of Health & Public ServiceLCPRIMARY_ADMIN; dean’s office in LC 414 factCHPS locations
College of Health & Public ServiceHPWEST_HEALTH_HUB; remaining health teaching and labs, but older program lists are partly stale estUVU locations, CHPS locations, current Dental location, current Respiratory location
College of Health & Public ServiceENPUBLIC_SAFETY_CORE; Criminal Justice and Forensic Science activity factCHPS locations
College of Health & Public ServiceMESPECIALIST; forensic laboratory factCHPS locations
College of Health & Public ServiceH6SPECIALIST; Center for National Security Studies factCHPS locations
School of the ArtsNCPRIMARY; dean, Music, Theatre, performance and rehearsal facilities factArts directory, Noorda facilities
School of the ArtsGTART_DESIGN_CORE; Art & Design office, studios, labs and galleries factArt & Design, facilities
School of the ArtsLADANCE_SATELLITE; Dance office and rehearsal scheduling factDance

Every UVU location beyond the Orem core

LocationStatusCoordinatesRole & programs
Orem West Campuscurrent40.2802, -111.7293Health Professions and Utah National Guard; HP used as anchor point
UVU locations, HP marker, NG marker
Vineyard / Genevacurrent fields future academic40.3071, -111.7424Current soccer camps, intramurals and recreation; planned 225+ acre health, wellness and innovation campus with future CHPS and CHSS space
UVU Fields, Vineyard master plan
Provo Airportcurrent40.2178, -111.7178Aviation Sciences plus Utah Fire & Rescue Academy and emergency-service training
Aviation visit page, Utah Fire & Rescue Academy, UVU locations
Lehi / Thanksgiving Pointcurrent40.4283, -111.8959TG and TH; MBA, MS Cybersecurity, general education, Dental Hygiene, Respiratory Therapy, Paramedic and Police Academy
UVU locations, Dental, Respiratory Therapy, Police Academy
Wasatch Campus, Hebercurrent40.5455, -111.4136Associate degrees, Elementary Education, graduate education, community education and WARM hospitality/resort-management pathway
Wasatch Campus, academics
Canyon Park (Bldg L)current40.3232, -111.6794Culinary Arts Institute, Canyon Park Building L
UVU locations, UVU marker
Capitol Reef Field Stationcurrent38.1852, -111.1795Field research, creative work and engaged learning across disciplines; capacity is 24 overnight or 40 day visitors, not enrollment
Field Station, plan a trip
Museum of Art, Lakemountcurrent40.2650, -111.7007UVU Museum of Art at Lakemount; public museum and visual-art academic resource
Museum visit page, UVU marker
Payson (land held)planned not operatingUVU owns 38.7 acres northeast of the I-15 Main Street interchange; construction schedule remains undetermined
UVU land announcement, Vision 2030
Eagle Mountainpartner or instructional point40.3800, -111.9719Current official marker; exact 2026 program and ownership scope not established
UVU marker, staff guide
Salempartner or instructional point40.0610, -111.6701Current official marker; exact 2026 program and ownership scope not established
UVU marker, staff guide
Santaquinpartner or instructional point39.9753, -111.7911Current official marker; exact 2026 program and ownership scope not established
UVU marker, staff guide
Spanish Forkpartner or instructional point40.1107, -111.6623Current official marker; exact 2026 program and ownership scope not established
UVU marker, staff guide

No current official source publishes campus-by-campus enrollment — UVU's Fall 2025 total of 48,669 is institution-wide and is not assigned to any location here fact. The Lehi program pages override the older CHPS location list (dental hygiene and respiratory therapy now teach at Thanksgiving Point).

The full building roster — all 53 pins

CodeBuildingPin (lat, lon)Zone
AMUCAS Gym40.2838, -111.7180North
ASUtah County Academy of Science40.2832, -111.7175North
AXAuxiliary Annex40.2738, -111.7318West (outlying)
BABrowning Administration40.2771, -111.7141Core
BBUCCU Ballpark40.2766, -111.7170South
BCNUVI Basketball Center40.2778, -111.7169Core
BRBusiness Resource Center40.2737, -111.7153South
C2Continuing Education 240.2778, -111.7051East
CBClarke Building40.2818, -111.7175North
CSComputer Science40.2790, -111.7110Core
DXDigital Transformation40.2785, -111.7328West (outlying)
ECUCCU Events Center40.2787, -111.7169Core
ENEnvironmental Technology40.2777, -111.7145Core
FAFaculty Annex40.2775, -111.7107East
FCFacilities Complex40.2803, -111.7052East
FGBrandon D. Fugal Gateway Building40.2771, -111.7135Core
FLFulton Library40.2810, -111.7164Core
GTGunther Technology40.2782, -111.7110Core
H2H2; no current expanded name published40.2748, -111.7083East
H3H3; no current expanded name published40.2747, -111.7078East
H4Army ROTC contested pin40.2805, -111.7136Core
H6Center for National Security Studies40.2802, -111.7093East
H71112 South 400 West40.2770, -111.7052East
H8Forensic Science Facility40.2766, -111.7052East
H91140 South 400 West40.2763, -111.7052East
H10GEAR UP40.2781, -111.7057East
H11TRIO Upward Bound40.2784, -111.7058East
H12Facilities South40.2751, -111.7066East
H131052 South 400 West40.2783, -111.7052East
H14Student Alumni40.2751, -111.7071East
H15Executive Events40.2749, -111.7071East
HFHall of Flags40.2779, -111.7144Core
HPHealth Professions40.2802, -111.7293West Campus
KBScott C. Keller Building40.2765, -111.7125Core
LALiberal Arts40.2801, -111.7167Core
LCLosee Center40.2784, -111.7125Core
MEMcKay Education40.2832, -111.7185North
NBMelisa Nellesen Center for Autism40.2831, -111.7195North
NCThe Noorda Center40.2777, -111.7091East
NGNational Guard40.2805, -111.7284West Campus
PSPope Science40.2780, -111.7150Core
RLRebecca D. Lockhart Arena40.2789, -111.7155Core
SASparks Automotive40.2776, -111.7118Core
SBScience Building40.2783, -111.7159Core
SCSorensen Center40.2786, -111.7143Core
SEScott M. Smith Engineering Building40.2761, -111.7073East
SLStudent Life and Wellness40.2795, -111.7150Core
SPSchool Community University Partnership40.2845, -111.7237North
WBWoodbury Building40.2774, -111.7131Core
WEWee Care Center40.2766, -111.7058East
WHSchool of the Arts Warehouse / Scene Shop40.2868, -111.7286West (outlying)
WSWolverine Service Center40.2820, -111.7225West Edge
YAYoung Alumni40.2830, -111.7205North

Source: UVU's live interactive campus map, accessed September 3, 2026, cross-checked against the Fall 2025 printable map and a 2023 building directory. Known name drift preserved in the working papers (Gunther "Technology" vs "Trades"; the CS building's pre-Smith label). One pin — H4, Army ROTC — sits ~880 m from its mapped footprint and is flagged rather than silently corrected fact.

Imagery rights

The map's aerial layer is NAIP imagery (USDA Farm Service Agency) served through the U.S. Geological Survey's National Map — U.S. public domain, with the credit line printed on the map page. UVU's own campus-map artwork is not an open layer (personal, non-commercial terms): it is used here as a factual reference for codes, names, and pin locations, and its artwork is not reproduced. Building-footprint data, if ever added, would come from OpenStreetMap under ODbL with its own credit fact.

J

Who at UVU is already doing it — the public record

Named champions, leadership statements, the friction people describe in public, student voice, and the adjunct facts

Appendix J in five lines

  1. QuestionWho at Utah Valley University has publicly shown what they think or built with artificial intelligence?
  2. AnswerThe public record gives the program someone to nominate in every college, but it also shows trust and part-time access problems.
  3. Deciding numbersemployee use rose from 61% to 76%fact75% expressed cautionfact62% of instructors are part-timeest
  4. What the plan doesUse peer picks from every college, pay part-time teachers for required work, and keep a path that does not require artificial intelligence.
  5. Still unknownUVU does not say whether concurrent-enrollment teachers sit inside the part-time count; adjunct access between terms and a formal student-government view are also UNKNOWN.

Adoption is the sponsor's first concern, so this appendix answers a plain question from public sources only: who at UVU has already said, in public, what they think and what they've built? Every row is a professional statement its author published — a news interview, a faculty profile, a conference program, a Senate minute — with the link beside it. Nothing here infers a private view; "silent" means the public record is silent, not that a person has no opinion. Sweep date: September 3, 2026.

What leadership has said

PersonPublic statementSignalWhere / when
Jon Anderson, President since August 10, 2026
University
“Artificial intelligence is a tool we are actively integrating to shape the future of education.”Champion at prior institution fact unknown2025-11-20, Pittsburgh Quarterly; then PennWest president · Statement; UVU profile
F. Wayne Vaught, Provost and Senior VP
Academic Affairs
“UVU is taking a leadership role in applied artificial intelligence.”Champion fact2025-03-20, UVU NVIDIA announcement · UVU
Christina Baum, VP Digital Transformation/CIO
Digital Transformation
“Identifying those faculty members and getting them to be champions is critical.”Champion with governance focus fact2026-02-11, EdTech interview · EdTech
Spencer Magleby, Dean
Smith Engineering & Technology
“Make sure that it’s not just cool, but that it locks in with people.”Cautious champion; human-centered fact2026-06-19, AI health hackathon · TechBuzz
Sue Jackson, Dean
Health & Public Service
no attributable public statement foundPublicly silent unknownChecked 2026-09-03 · Role; targeted name/AI searches found no attributable statement
Steven Clark, Dean
Humanities & Social Sciences
no attributable public statement foundPublicly silent; college itself is active unknownChecked 2026-09-03 · Role; no attributable statement found
Daniel Horns, Dean
Science
no attributable public statement foundPublicly silent unknownChecked 2026-09-03 · Role; a public AI-project repost added no attributable words
Krista Ruggles, Interim Dean
Education
“Artificial intelligence is transforming elementary STEM education, yet evidence remains fragmented.”Champion with evidence and privacy guardrails fact2025-10-30, coauthored review · Paper; role
Courtney Davis, Dean
Arts
“We are proud to provide education to our students that will put them at the forefront of arts technology.”Champion fact2026-02-20, AI Design Awards announcement · UVU story; syndicated full quote
Bob Allen, Dean
Woodbury Business
no attributable public statement foundPublicly silent; college faculty are highly active unknownChecked 2026-09-03 · Role; no substantive attributable statement found

Continuity is strong: the institute was launched under the previous president with the stance that everyone, regardless of major, should engage with the technology; the current president championed AI integration at his prior institution. Four deans are publicly silent while their colleges are visibly active — an invitation to bring them in, not evidence against.

Faculty champions and skeptics, by college

No named faculty member was found publicly advocating a blanket AI ban. The common design choice among the people who have built something — a revision coach that refuses to draft, a marketing "boss" that challenges strategy, virtual patients, a course bot with disclosure and verification controls — is guardrails, verification, and human judgment. That is also the design of this plan, which is why these people are its natural first cohort.

NameCollegePublic stance and evidenceSource
Anne Arendt
Technology Management; associate dean at statement
SmithChampion; urged preparation across fields and focus on opportunities factDate UNKNOWN, circa 2024, UVU Review
Jenny Nehring; Troy Taysom
Information Systems & Technology
SmithChampions; presented “ChatGPT–AI: Embracing Change” to educators fact2023-06-13, Nehring profile, Taysom profile
Armen Ilikchyan
Technology Management/Mechatronics; Applied AI program leader
SmithChampion with guardrails; uses AI-simulated student feedback for course improvement fact2025-02-27, profile
George Rudolph
Computer Science chair
SmithChampion with guardrails; moved from treating AI work as near-cheating to teaching multi-agent orchestration after employer feedback fact2026-08-10, TechBuzz
Majid Memari
Computer Science
SmithChampion/cautious; built AI-fluency modules and a grounded course chatbot with disclosure, verification, privacy, and autonomy controls fact2026-08-10, TechBuzz
Xi Chen
Computer Science
SmithCautious integrator; “Coding Twice” has students code independently before using AI fact2026-02-19, SIGCSE
Zac Taylor
Applied Engineering & Transportation
SmithChampion with caution; supports responsible use and broader access to technical knowledge fact2024-09-11, UVU Review
Tyson Riskas
Information Systems/Applied AI
SmithChampion/cautious; presented on assessing IS learning in the GenAI era and teacher preparation fact est2024-07-25, profile
Noah Myers
Accounting
WoodburyChampion with guardrails; uses agents for forensic accounting after students learn manual foundations fact2024-08-29 and 2026-08-10, KSL, TechBuzz
Diego Alvarado-Karste
Marketing
WoodburyChampion with guardrails; built an AI “boss” that challenges strategy rather than writing the finished work fact2026-08-10, TechBuzz
Yang Huo
Strategic Management/Operations
WoodburyChampion; completed GenAI and ChatGPT training to update curricula and workforce skills fact2025-04-11, profile
Kari Olsen
Accounting
WoodburyCautious/evaluative; coauthored empirical work on ChatGPT performance on accounting assessments fact2023, profile
Qianwen “Rachel” Bi
Finance/Personal Financial Planning chair
WoodburyChampion/cautious; teaches practical AI while raising credit, privacy, ethics, and general-AI questions fact2023-06 and 2024-05, UVU Summit, graduation program
Jacob Burdis
Product Management/Marketing adjunct
WoodburyChampion in professional work; co-founded an AI-supported education company factDate UNKNOWN, profile
Maritza Sotomayor
Economics; UNESCO Chair associate director
WoodburyCollective ethical-AI role is public; individual classroom position was not found unknown2025, UNESCO report
Krista Ruggles
Elementary Education/STEM
EducationChampion with guardrails; co-developed a multi-agent classroom simulation with teacher feedback fact2026, i-ETC program
Uzeyir “Adam” Ogurlu
Elementary Education
EducationCautious/champion; research finds openness to training alongside plagiarism, overreliance, reasoning, and authenticity concerns fact2023-12-08, RESSAT
Bing Han
Elementary Education
EducationNamed on the UNESCO AI and Education Committee; personal stance not separately attributable unknown2025, UNESCO report
Angie McKinnon Carter
English
CHSSCautious adopter; built a revision coach that refuses to draft, retains human meetings, and permits opt-outs fact2026-08-10, TechBuzz
Yi Yin
Sociology/Behavioral Science
CHSSChampion with opt-outs; uses AI for course materials and comparative research while supporting informed refusal fact2024–26, profile, TechBuzz
Christa Albrecht-Crane
English; Writing Program chair
CHSSStrongest public skeptic; argues NotebookLM’s frictionless compression can conflict with writing pedagogy and educational values fact2025-10-24, SIGDOC paper
Tom Henry
English
CHSSCautious/mixed; separates useful writing support from harmful substitution fact2025, Utah English Journal
Shane Smith
Philosophy adjunct
CHSSPragmatic champion; opposes leaving students unprepared but warns that dependence can remove learning fact2024-09-11, UVU Review
Brian Whaley
English and Literature chair
CHSSCautious/champion; supports experimentation and says policing-only responses waste effort factCirca 2024, UVU Review
Kelsey Hixson-Bowles
English
CHSSChampion/researcher; studies faculty and student perceptions to guide classroom practice fact2025–26, profile
Devin Gilbert
Languages & Cultures
CHSSCautious integrator; teaches bounded use of machine translation and GenAI fact2024-11-20, profile
Acacia Overono
Psychology
CHSSCautious/mixed; promotes disclosure and critical review of AI’s classroom and social effects fact2024–26, profile
Chris Weigel
Philosophy; Center for the Study of Ethics
CHSSCautious; teaches AI ethics and publicly led “Wrestling With AI Policy at UVU” fact2025-10-02, Ethics Week
Eunmi Joung
Mathematics
ScienceChampion with guardrails; found that critiquing ChatGPT answers can deepen mathematical reasoning fact2025-09-19, study
Roxanne Brinkerhoff, Inyoung Lee, Ka Lun Wong, Ofa Ioane, Eunmi Joung
Mathematics
ScienceStudy reports moderate openness, low confidence, changed teaching, and demand for responsible-use support fact est2026-03-15, EJMSTE
Britt Wyatt
Biology
ScienceEngaged/champion; presented on adapting practical science-learning tools for a GenAI environment fact est2025-02-28, profile
Jill Johnson
Nursing
Health & Public ServiceChampion after initial caution; used AI Academy learning to create five interactive virtual-patient chatbots fact2024–25, profile
Hsiu-Chin Chen
Nursing
Health & Public ServiceChampion; joined the 2026 AI Fluency Summer Institute to improve teaching and research fact2026-05-05, profile
John Fisher
Emergency Services/Public Administration
Health & Public ServiceMixed; publicly covers AI’s benefits, competence, access, cost, ethics, and risks fact est2026-04-09, profile
Merilee Larsen
Health Sciences chair
Health & Public ServiceCautious/champion; participated in health/public-safety ethics and UVU policy panels fact2025-10-02, Ethics Week
Brandon Truscott
Entertainment Design
ArtsMixed; presented “The Quick and the Dread: Threats and Benefits of AI in Teaching” fact est2024-08-14, profile
Shirin Abedinirad
Art & Design, Sculpture/Ceramics
ArtsPublicly engaged through an “AI and Creativity” panel; exact position not recoverable unknown2025-09-30, Ethics Week
Cheung Chau
Music
ArtsUNESCO executive-board participation is public; no individual AI teaching statement was found unknown2025–26, UNESCO

About 37 faculty joined the May 2026 AI Summer Institute; six are publicly profiled. The behavioral program in the guide (§07) starts from peer nomination, not this list — but the list shows every college already has someone to nominate.

The friction people describe in public — and what the plan does about each

Friction (public evidence)What it calls for — and where the plan answers itGrade / date
Copyleaks false positives reached Faculty Senate
“CopyLeaks has been reporting false positives regarding AI generated content.” Minutes
Do not use detector output as sole proof; require human review and process evidencefact
2025-04-29
An ENGL 2010 syllabus scanned every paper while warning students to keep version history against false flags
Syllabus
Consistent due process and clear evidence standardsfact
Spring 2025
UVU says assignment, course, department, and university AI rules may conflict
Student guidance
A simple baseline policy, with course-specific additionsfact
Accessed 2026-09-03
Gateway is employee-only; free usage resets weekly at a $5 allowance; heavier use needs a paid tenant
UVU employee page, TechBuzz
Student access plan; published quota behavior; department path for heavier usefact
2026-07-14
Wilson is in selected courses and carries a hallucination warning; no audited public error rate
Wilson, EdTech
Publish accuracy, outcome, accessibility, and rollout testsfact unknown
2025–26
Employee AI use rose from 61% to 76%, yet 75% expressed caution; “very satisfied” fell from 23% to 15%
Board packet
Trust-building and measured outcomes, not just activation countsfact
2026-01-29
Faculty’s largest training request was ethics/responsible use/academic integrity, 34%; staff asked for ethics/privacy/policy, 26%
Board packet
Discipline-specific training plus usable policy examplesfact
2026-01-29
Diego Alvarado-Karste abandoned a standalone build because institutional approval was unrealistic inside the two-week institute
TechBuzz
A preapproved sandbox, reusable components, and fast reviewfact
2026-08-10
Policy 441 AI entered Stage 1 with a June 2026 Board target, but it is absent from the current approved manual
Executive summary, current manual
Publish current stage, operative rules, and relation to student academic useunknown
2025-09-25 to 2026-09-03
Optional adjunct pedagogy/technology workshops are unpaid; appointments end each semester
Policy 639
Paid AI training and account continuity for active/returning adjunctsfact est
Current
No UVU-specific public accessibility complaint or AI-product accessibility evaluation was located
UVU AI guidelines
Publish accessibility testing and feedback results; absence of complaints is not proof of accessibilityunknown
Checked 2026-09-03

The most telling pair: employee AI use rose from 61% to 76% in a year while "very satisfied" fell from 23% to 15% and 75% expressed caution. Access is not the bottleneck; trust and fit are. That is the case for a service designed around guardrails, predictable quotas, a real non-AI path, and no detector-as-proof — the rights charter in §07.

Student voice

What students saidThe ask it impliesGrade
Computer Science student Jacob Barrus found AI useful for explanations but warned, “if you rely too heavily on AI, you don’t learn.” English student Salem Kimball worried about lost investment in learning.
2024-02-20, UVU Review
Permit help without replacing the student’s thinkingfact
Poster reported 27% Copyleaks, 4% Grammarly, and 10% Turnitin on the same claimed self-written paper.
2024-11-10, r/UVU
Do not treat one detector score as prooffact
Commenters described supplying edit history, waiting for responses, and changing writing style to avoid flags.
2024-11-11, r/UVU
Prompt evidence review and a clear appeal pathfact
Anonymous respondents reported false flags and preferred human grading and transparent rules. More than one-third reportedly avoided AI for ethical or environmental reasons; sample details were not published.
2025-11-20, UVU Review
Human review, disclosure rules, and a valid non-AI pathfact unknown
Named student panels addressed AI and creativity and the future of AI. Public schedule proves participation; individual positions are UNKNOWN without transcripts.
2025-09-30 to 10-03, Ethics Week
Continue structured student forums and publish their recommendationsfact
Student preference survey: ChatGPT 71%, other tools 13%, Gemini 9%, Copilot 7%. Common uses included research/questions, concept explanation, ideas, and homework help.
2026-01-29, Board packet
A student service must support the tools they already use or give a clearly better equivalentfact
Student-reporting synthesis raised employment, environment, mental-health, training-data, and artists’ rights concerns alongside educational value.
2026-02-18, UVU Review
Teach social and rights impacts, not only promptingfact
CHSS and the Institute held “Student Voices on AI” and faculty/student-perception sessions.
2026-03-04 and 03-27, Yi Yin profile
Turn these forums into published design requirements and outcome measuresfact
No public UVUSA AI resolution, Copyleaks discussion, Policy 441 position, or recorded AI debate was found.
Minutes checked through 2026-03-19, UVUSA minutes
Obtain and publish a formal student-government position before broad student rolloutunknown

Adjunct facts

Most UVU instructors work part-time. The public record on their AI access is thin, and thin in a specific way:

EvidenceWhat it meansGrade
Shane Smith, adjunct Philosophy professor, publicly discussed AI proof limits and estimated that an AI pre-grader could reduce about eight grading hours to three. UVU Review, 2024-09-11Adjunct workload is a direct adoption use case, but automated grading also raises due-process risk.fact
UVU’s AI Task Force charge expressly includes training and support for full- and part-time faculty. Task Force chargeAdjunct inclusion exists at the intent level.fact
The Board reports 120 AI Academy completers and 22 advanced participants, with no employment-class breakdown. Board packetWhether any adjunct completed the Academy is UNKNOWN.fact unknown
Public AI Academy pages do not state adjunct eligibility, compensation, selection rules, or account continuity.The phrase “for faculty” is not enough to prove practical adjunct access.unknown
The 2026 Faculty Summer Institute offered selected faculty a $1,500 completion stipend, but did not name or exclude adjuncts. Institute, AI briefEligibility by faculty type remains UNKNOWN.fact unknown
Gateway is available to “all employees.” Policy 639 ends adjunct employment each semester. Employee page, Policy 639Active adjunct access is a reasonable EST; access before, after, or between appointments is UNKNOWN.fact est
Required adjunct pedagogy orientations are paid hourly; optional pedagogy and technology workshops are unpaid. Policy 639Optional AI training competes directly with unpaid time.fact
Published Faculty Senate eligibility is limited to salaried, benefits-eligible faculty; adjunct welfare is a Senate purpose, but adjuncts cannot serve as senators. ConstitutionAI governance lacks a clearly published direct adjunct seat.fact
No “UVU Adjunct Association” or AI-specific statement from the public UVU AAUP/AFT chapter was found. UVU AAUP/AFTA public adjunct AI voice or advocacy channel was not located.unknown
UVU’s own Common Data Set counts 1,310 part-time and 802 full-time instructional faculty in fall 2025. That works out to 62.0% part-time. UVU People & Culture publishes the same two counts. Common Data Set 2025–26, §I-1, People & CultureThe 62% share is confirmed at a UVU source. It counts people, not full-time jobs, and UVU does not say here whether every part-time teacher is an adjunct or whether concurrent-enrollment teachers are inside the count.fact counts · est share
The narrower count. IPEDS reports 937 part-time and 809 full-time UVU instructional staff for fall 2024, which is 53.7% part-time. IPEDS counts only staff whose main job is teaching, and leaves out graduate assistants, so its list is shorter than the Common Data Set’s. IPEDS, UVU staff tableA reader who finds 53.7% somewhere else is looking at a different year and a narrower list of people, not a different answer.fact counts · est share

Service implication: paid, short, asynchronous training; accounts and course agents that survive the semester boundary; eligibility stated in plain language; an adjunct advisory seat before any policy or detector rule changes. The plan's artifact-gated stipend and adjunct-first design (§07) already assume this — the public record confirms why.

Does the public voice change the sequencing?

CollegeAppendix A ratingVerdict after the public-voice sweepReason
Smith Engineering & TechnologyHIGHStay HIGHWidest technical roster, several named classroom uses, department-chair leadership, coding-tool evidence, applied programs, and explicit employer pull.
Woodbury BusinessHIGHStay HIGHMyers, Alvarado-Karste, Huo, Bi, Olsen, and others show practical use across accounting, marketing, finance, management, and assessment.
School of EducationHIGHStay HIGHInterim dean Ruggles is a named AI researcher; Ogurlu supplies cautious evidence; AI Academy completions and a multi-department panel show breadth. Named classroom implementations remain fewer than Smith or Woodbury.
Humanities & Social SciencesHIGHStay HIGHThis is the broadest public conversation: writing tools, source verification, sociology research, translation, disclosure, ethics, opt-outs, and the strongest skeptical voice. Caution is evidence of active engagement, not low receptivity.
College of ScienceMEDStay MED, strengtheningMathematics research, biology teaching work, the science panel, and DIFFRAX show activity, but named public classroom deployment remains concentrated.
Health & Public ServiceMEDMove to HIGH, controlledJill Johnson’s five virtual-patient bots, Chen’s Summer Institute work, Fisher’s adoption talk, and a panel spanning health, forensics, PA, and emergency services fill x24’s local-evidence gap. Start with synthetic cases and human review.
School of the ArtsMEDStay MED, low confidenceDean Davis supports arts technology; Truscott and Abedinirad show public engagement. However, there is still little public evidence of a repeatable classroom program, policy, or broadly used tool.

One change carried into Appendix A: Health & Public Service moves to HIGH (controlled) — named nursing champions with built virtual-patient bots filled the local-evidence gap. The controls do not move: synthetic cases first, human review always.

K

What we plug into — and the five questions that remain

Integration surfaces from the public record, the house's compute reality, and the ask list shrunk to what only UVU can answer

Appendix K in five lines

  1. QuestionWhat campus systems must the service join, and what must Utah Valley University tell us?
  2. AnswerPublic records confirm the main campus systems, but five questions that only the university can answer still block the connection.
  3. Deciding numbers360-person Gateway pilotfactabout $5 of tokens per employee per weekfact44 courses and 2,000 studentsfact
  4. What the plan doesBring its own computers, user page, and support, then use only approved sign-in, course, data, network, and support links.
  5. Still unknownThe five questions cover sign-in, Gateway and Wilson design, data rules, the network handoff, and operating owners; no login was attempted.

The service stands on its own; this appendix maps the house it stands in. Every surface below was established from public sources — UVU's own pages, job postings, policy documents, network registries, and a handful of passive checks any browser performs on first visit — so that the questions we finally ask UVU are few, precise, and cannot be answered from outside. Sweep date: September 3, 2026.

Integration surfaces

SurfaceWhat a self-contained service integrates withEvidence
Microsoft Entra / Azure ADInstitutional SSO, user identity, group or role claims, MFA, lifecycle controlsPasswordless setup expressly requires Azure AD synchronization. UVU’s Microsoft job owns Entra, tenant management, Conditional Access, Defender, and DLP. Passwordless guide, job posting
MFAUVU’s current page uses Microsoft MFA for myUVU/Microsoft apps and also documents Duo for some flows.MFA page, Atlassian support flow. Exact endpoint-by-endpoint chain is UNKNOWN.
CanvasLTI 1.3 placement, deep links, context/role claims, approved APIs, test courses, accessibility reviewCanvas is UVU’s official LMS. UVU requires legal, stewardship, purchasing, accessibility, and systems review for integrations. Canvas, technology review
BannerAuthoritative enrollments and course roles; preferably consumed through the approved Canvas/Banner or API pathUVU says Banner automatically adds students and governs teacher/assistant roles in Canvas. Canvas roles
Banner versionBanner 9 Self-Service is the stated replacement for retired SSB8. Core Banner/Ethos/API versions are not public.SSB8 retirement
Existing Canvas LTIsAvoid duplicate tool placement and account for content/data flowsUEN’s April 2024–April 2025 report gives UVU 60,735 unique Canvas users and lists Proctorio, Kaltura, Watermark Insights, McGraw Hill Campus, and Cengage among its top LTIs. UEN Canvas report
Teams LTIEvidence that UVU supports roster-aware Microsoft/Canvas integration patternsInstructure reports Teams/Canvas schedule and roster synchronization at UVU. Case study, Microsoft 365/Canvas article
Pathify/myUVUPortal card/deep link, notifications, service discovery, possibly SSOMobile moved to Pathify in April 2026. Desktop myUVU is scheduled for November 16, 2026. UVU mobile, FAQ
API layerPrivate UVU API gateway, data lake and governed data productsThe Cloud Data Engineer role says “Own API Management space at UVU.” UVU’s public tool list shows Banner AppNav, IDM Person Sync, and AD tools, many gated to campus/VPN. No public developer portal was found. Job, tools
IT service managementIncident, request, change, and service catalog integrationUVU Status confirms Atlassian/Jira Service Management. The SRE job confirms PagerDuty on-call. ServiceNow appears only as a possible skill. Status, SRE job
Monitoring/securitySplunk route, UVU monitoring, endpoint protection, security reviewInternet2 lists UVU as a Splunk subscriber. CrowdStrike Falcon with OverWatch is the stated default on faculty/staff computers. Splunk, CrowdStrike
Data governanceData classification, steward approval, minimum-access personas, architecture review, data catalogPolicy 445, effective 2024-03-28
PrivacyApproved purpose, authorized storage, collection notice, third-party disclosurePolicy 446, effective 2023-05-09
SecurityVendor review, MFA, least privilege, encryption, logging, inventory, tested changes, network controlsPolicy 447, effective/reviewed 2026-08-10
RecordsPrompt, output, support, audit, roster-cache, and analytics schedules; GRAMA/legal-hold handlingUVU says chat may be logged and records may be subject to GRAMA. Retention follows record content, schedules, and holds—not a universal deletion period. Privacy statement, Policy 451
AI policyNo approved Policy 441 appeared in the current manual at cutoff. A 2025 AI-policy proposal existed, but its PDF now returns 404.Policy manual, pipeline, Library AI guide
Identity tenant (probed)Microsoft Entra ID, cloud-managed ("Managed" namespace, no on-premises federation) — a standard app registration is all our sign-on needsPublic tenant-discovery endpoints, 2026-09-03
AI Gateway hosting (probed)Both Gateway hostnames are published through Microsoft Entra Application Proxy — a privately hosted app, pre-authenticated by the tenant, two separate app registrations. Our hook to it is an API endpoint its backend can call; we never sit inside itPublic DNS + one unauthenticated request per hostname, 2026-09-03
Wilson hosting (probed)Public Wilson answers from a Cloudflare-fronted host whose page title names a server node — self-hosted, node-named, consistent with the vendor-platform finding abovePublic DNS + page title, 2026-09-03
Campus address space (probed)UVU owns its own public /16 block; core portals and VPN answer from it — an on-campus data center with public routing exists, so a rack for the fleet is a facilities conversation, not a hosting-model onePublic DNS + ARIN registry, 2026-09-03
Advertised model endpoints (probed)None — no public "llm", "ollama", "copilot", or similar hostnames exist. Nothing to collide withPublic DNS, 2026-09-03

Reading it together: a Microsoft-shaped house (the CIO's own words: "We're kind of a Microsoft shop. We have an A5 license"), Canvas with Banner-fed roles, a portal moving to Pathify this fall, Jira Service Management with PagerDuty on-call, CrowdStrike on managed devices, Splunk available, and Policies 445/446/447 as the written rulebook. Everything our design already assumed — now sourced.

The house's compute and network, as publicly known

AreaPublic factWhat it means for usGrade
UVU internet uplinkUEN increased UVU internet capacity by a factor of five within ten days, completing the change March 18, 2025. Absolute pre/post capacity was not given. UENSize external model traffic and updates against a private bandwidth figure; do not interpret “5×” as a line rate.fact
Research networkInternet2 lists UVU as a primary participant through UETN since 2018-01-01. Internet2UETN/Internet2 is an available institutional network path, but private routing and peering details are still needed.fact
Cloud/service procurementInternet2 currently lists UVU as a subscriber to AWS, Canvas, Pathify, and Splunk.These are real institutional relationships and possible contract routes. They do not prove workload location, editions, spend, or Direct Connect.fact
UVU-owned HPCThree UVU-owned Notchpeak nodes at University of Utah CHPC: one 96-core/380 GB node and two 64-core/150 GB nodes. Total: 224 cores and 680 GB RAM. UVU accounts receive priority. No GPUs are listed. UVU HPCUseful research CPU capacity exists, but it is not a proved production hosting surface.fact
One-U H200 nodesCHPC documents ten nodes, each with eight H200 GPUs, 96 CPU cores, and 2 TB RAM. It says all CHPC users can use guest access and AI researchers may request priority. CHPCEST a UVU CHPC account may submit guest work. UVU entitlement, priority, protected-data approval, and service hosting rights remain unproved.fact unknown
RedtailRedtail is a statewide, University of Utah-managed AI supercomputer. No public UVU allocation, proposal, account, workload, or Gateway/Wilson connection was found. CHPC, funding accountDo not count Redtail as UVU capacity.fact unknown
Smith Engineering BuildingOpened January 22, 2026; nearly 200,000 square feet. Public descriptions name an AI lab, VR functions, and a machine-learning/autonomous-systems drone lab. No server or GPU specifications are published. UVUBuilding and labs are real; production rack, power, cooling, network, and GPU capacity are private.fact
SCaFLIndexed/archived material describes a donated 10 Gb UTOPIA connection and local servers for a public/business lab. The current SCaFL pages and old impact report returned 404.It is not evidence of UVU’s campus backbone or current service capacity.HISTORICAL · C
Campus data centerUVU’s design guide follows data-center and server-room standards, including TIA-942 considerations, power, cooling, security, and growth space. It does not reveal an operating data-center location or capacity. Telecom guideRack location, available power/cooling, segmentation, and remote-hands service remain private.fact unknown
JamfUVU’s page says Jamf handles Apple inventory, deployment, updates, encryption, and security. It says university-owned Apple devices were to be enrolled, but gives no count. JamfA managed Mac client may need Jamf packaging; this does not provide usable hardware inventory.fact
Citrix/MyAppsUVU publicly lists Citrix-delivered MyApps and many student applications. No GPU-backed VDI profile, vGPU type, host count, or capacity was found. Software inventoryBrowser-based delivery is safer than assuming a GPU-capable VDI client.fact unknown

Gateway and Wilson: what is actually known

Treat the employee Gateway, the public-website Wilson, and the in-course Wilson as three separate services until UVU shares an as-built diagram. The only safe shared assumption is Microsoft identity and governed UVU content — not a shared model layer. Notably, the public-website Wilson runs on a vendor platform inside UVU's own cloud tenant with a named UVU engineer operating it; whether the Canvas Wilson shares that backend is not public.

ComponentPublic evidenceGrade
Gateway front doorUVU says every employee has a free, secure platform for approved AI tools. The public endpoint invokes Microsoft sign-in, but no login was attempted. UVU employee pagefact
Gateway launch and populationReported employee launch July 1, 2026, after a 360-person pilot. Week-one registered users rose 24%; active users rose 31%. TechBuzzfact
Gateway model familiesReported access to ChatGPT, Claude, Gemini, and locally hosted open-source models. Exact providers, versions, endpoints, and local models are UNKNOWN. TechBuzzfact
Gateway functionsMulti-model comparison, projects/folders, personal knowledge bases, prompt chains, and agent creation/export are publicly reported.fact
Gateway implementationUVU’s Tyler Small said it used “FERPA-secure APIs and off-the-shelf templates.” TechBuzz describes the result as assembled in-house.fact
Gateway allowanceAbout $5 of tokens per employee per week, resetting Monday; a paid tenant was reported for heavier use. Actual model rates and spend are private.fact
Gateway student rolloutJuly reporting said students were next. No public student launch date, eligibility rule, or budget was found.unknown
Gateway stackSearches found no reliable tie to LibreChat, Open WebUI, LiteLLM, Portkey, Azure OpenAI, Bedrock, or any named router/database. “Locally hosted” does not identify where.unknown
Gateway data termsUVU’s employee page describes protective Microsoft terms, but those terms are not expressly the Gateway’s retention, logging, or provider-training contract.unknown
Wilson product surfaceUVU calls Wilson a UVU-developed assistant spanning its public website and selected Canvas courses. Course access excludes grades, tests, and private files; UVU says it does not store personal information. Wilson AIfact
Wilson rollout14 classes in 2024; 26 courses and 1,800+ students in fall 2025; 44 courses and 2,000 students by February 11, 2026. No later count was found. PBA, board minutes, EdTechfact
Wilson retrievalCourse Wilson is grounded in permitted course content. Baum specifically named Kaltura recordings and links to relevant points in a recording. The vector store, embeddings, chunking, reranking, and indexing cadence are private.fact unknown
Public Wilson platformCloudforce names the public Wilson as a “nebulaONE-powered conversational agent” and says Daniel Razo “built it, supports it, manages it, and continually improves it.” Cloudforcefact
Canvas Wilson relationshipNo source proves whether public Wilson and the 44-course Canvas Wilson share a tenant, agent, retrieval store, codebase, or model route.unknown
nebulaONE defaultsCurrent Cloudforce terms say nebulaONE ordinarily processes inside the client’s Azure tenant and does not train models on client data. Its privacy page describes anonymized telemetry retention. These are product defaults—not UVU’s executed terms or configuration. Terms, privacyest
Wilson model/providernebulaONE documentation names Microsoft Foundry/Cognitive Services and gives Azure OpenAI and Claude as examples. No public source identifies UVU’s selected models.unknown
Legacy WilsonOcelot powered the 2020 Wilson. UVU announced a replacement/transformation effective May 1, 2023. Ocelot is not evidence of the current stack. Ocelot, UVU transformation pagefact
Wilson roadmapFebruary 2026 plans included connecting course agents and bringing universitywide answers into course Wilson. Public evidence does not establish completion.fact unknown

The twelve-item data request, after the public sweep

None of the twelve could be fully closed from outside — each bundles public facts with private contracts or configuration — but most no longer need answering before we build. The service brings its own compute, front door, and operations; only the interface details block integration.

#ItemStateClosed by public evidenceWhat stays private — and whether it blocks us
1Software and spendunknownNamed public platforms now include M365/Entra, Canvas, Pathify, AWS, Splunk, Jira, CrowdStrike, Kaltura, and others.Full SKU/seat/spend/renewal/reseller ledger. Does not block a self-contained service except where a contract governs an integration.
2MicrosoftunknownA5 environment, M365 tenant operations, Entra, Conditional Access, DLP, Defender, Power Platform, Graph direction.Exact population mix, Copilot/ChatGPT SKUs, Azure OpenAI/Copilot Studio rights, $600,000 request outcome. Only identity/application registration blocks integration.
3GatewayunknownModel families, employee rollout, pilot size, allowance, features, in-house/template build, FERPA-secure API claim, students-next direction.Exact software, models, providers, router, hosting, APIs, connectors, spend, retention, logs, and training terms. Blocks coexistence/routing decisions, not independent hardware.
4WilsonunknownPublic Wilson is nebulaONE; Daniel Razo operates it. Canvas/Kaltura grounding and 44-course/2,000-student latest count are public.Whether web/course products share a backend; model/provider, RAG, LTI/API, hosting, retention, cost, current count, assessment, oversight. Blocks overlap and Canvas design.
5CanvasunknownOfficial LMS, Banner-fed roles, approval process, major LTIs, Teams sync pattern, Internet2 subscription.Core/Plus/Next edition, IgniteAI flags, current Copyleaks state, contract/data terms, exact LTI/API configuration. Interface fields block integration; commercial tier mostly does not.
6ComputeunknownThree Notchpeak nodes: 224 CPU cores/680 GB; statewide One-U and Redtail resources described.Current status/use, GPU entitlement, allocations, protected-data approval. Does not block because the planned service brings its own compute.
7Campus hardwareunknownJamf function, MyApps/Citrix, Smith labs, SCaFL historical 10 Gb link, and room purposes.Counts, GPU/server models, Citrix profiles, rack/power/cooling, operational state. Only the proposed service’s facility/network envelope blocks us.
8NVIDIAunknownThree-year, no-funds collaboration; training, workshops, tools, cloud workstations, ambassador pathway.Executed MOU, active deliverables, credits, equipment, usage. Does not block the self-contained service.
9People and fundingunknownKahlert $5.2 million; $2 million one-time recommendation; AI/Dx reinvestment figures; public program counts from x29.Full rosters, outcomes, grant ledger, placements, gift restrictions. Useful diligence, not a technical blocker.
10Data and agreementsunknownUVU Data Lake, Azure/Synapse/API Management, hybrid environment, governance/security/privacy policies.As-built schemas, system-of-record map, current tools, API contracts, executed MOUs/DPAs/DUAs, retention and subprocessors. Data/interface parts block integration.
11State and UNESCOunknownState task force/credential intent, Redtail facts, UNESCO chair and committee are public from x29.Credential launch/platform/enrollment/RFP; chair funding, executed agreements, live projects. Not a core infrastructure blocker.
12Payment reconciliationSTILL PRIVATENo reliable FY2024–FY2026 provider/reseller payment export was completed.Vendor payments, reseller routes, contracts and ledger reconciliation. Useful diligence; not needed to size our own service.

The five questions that genuinely require UVU

Asked once, together, of the technical side of the house:

  1. Identity and application interface pack: What Entra registration, protocol, issuer/tenant, claims, groups, MFA/Conditional Access rules, test accounts, Canvas LTI 1.3 deployment details, Canvas API scopes, Banner-fed role source, Pathify placement, and UVU API Management route must the service use?
  2. Gateway/Wilson as-built boundary: Please provide one current data-flow diagram for Gateway, public Wilson, and course Wilson showing platform/version, hosting, models/providers, routing, retrieval sources, Canvas/Kaltura interfaces, and whether any components or stores are shared. Must our service hand off through any of them?
  3. Approved data and control envelope: Which Policy 445 data classes may enter the service; what are the required retention, deletion, legal-hold, logging, no-training, subprocessor, region, encryption, accessibility, privacy-notice, security-review, and final Policy 441 requirements?
  4. Network and facilities handoff: For owner-supplied hardware, what rack location, power/cooling, physical access, VLANs, firewall/egress rules, DNS/TLS, load-balancing, UETN/Internet2 path, private-cloud links, remote administration, monitoring, backup, and disaster-recovery rules apply?
  5. Operating contract and rollout: Who owns product approval, data stewardship, Canvas administration, security, networking, support, change control, incident response/JSM/PagerDuty, and service-level decisions; what pilot users/courses and current Wilson/Gateway rollout constraints must we plan around?

Everything else in the twelve-item ledger is commercial diligence that can be requested later as an export, or never. The four decisions that belong to Dr. Burns rather than to IT stay in Appendix H: whether a Moonshot or research-fund notice of intent was filed, whether a budget number exists or the dial stays the instrument, which college sponsors the pilot and how Dx co-sponsors, and whether a seat quote under $30 per user per year should flip the plan automatically.

L

The free path, honestly

Free model access for a whole university — what is really on offer, where it breaks, and what it is good for

Appendix L in five lines

  1. QuestionCan free online model access serve all 36,600 people at Utah Valley University?
  2. AnswerNo; three measurable offers cover only 1.38% of a normal day when run from university accounts.
  3. Deciding numbers78,690 turns per day neededest1,082 turns per day suppliedest1.38% coveredest
  4. What the plan doesUse these offers only for public or made-up data, tests, and a separate sandbox; keep the owned local service as the floor.
  5. Still unknownMany rate limits, retention terms, and end dates are UNKNOWN; no account was made and no request was sent.

A fair challenge was put to this plan: before buying anything, have we exhausted the free options? Free tiers exist from a dozen providers, and as of this week Meta's newest model is free through a coding gateway. This appendix takes that path seriously, on the providers' own published terms, and shows exactly where it holds and where it breaks for 36,000 people. Verified directly on September 3, 2026, then extended by a full provider survey the same day.

Meta Muse Spark 1.3 — the fact sheet

QuestionAnswerGrade
What is it?Meta's newest frontier model, released September 2, 2026 — trained for long-horizon, agentic work; 1M-token context; text, image, and video input.fact
Is it open-weight — can a campus run it on its own machines?No. It is API-only today. Meta's own announcement lists an open-weights release on its roadmap with no date. Until weights ship, it cannot be part of an owned floor.fact
How good is it?Artificial Analysis Intelligence Index: 61 (xhigh) / 62 (max) — level with GPT-5.6 Sol (61) and Grok 4.6 (61), just under Claude Opus 5 (63), under Claude Fable 5.1 (66), and above the best open-weight model, GLM-5.3 (60). Meta's coding claims hold on two benchmarks: it ties Sol on Terminal-Bench 2.1 (88.8) and leads both Opus 5 and Sol on DeepSWE (75.4 vs 74.0 / 73.0); it trails Opus 5 on knowledge-work and computer-use benchmarks.fact
What does it cost?Standard tier: $1.25 per million input tokens, $4.25 output (cache hits $0.15). Roughly 40% cheaper per benchmark task than GPT-5.6 Sol at the same score.fact
Where is it "free"?Through the contributor tier: $0.10 / $0.20 per million tokens on Meta's platform, and $0 on OpenCode Zen "for a limited time" — in both cases in exchange for permission to use prompts and completions to train future Meta models ("used to improve our products").fact
What is the catch for a university?The contributor tier is a data-for-discount trade. Every prompt becomes training data, which rules it out for anything touching student records, unpublished research, licensed course material, or personal data — the Sensitive and Restricted classes in UVU's own policy. The free window has no stated end date, and the offer is aimed at individual developers; there is no institutional, education, or admin-controlled program published.fact terms · est fit
Where does it fit this plan?As a candidate for the metered frontier pool under the standard (no-training) tier — a strong, cheaper coding and agent route beside Gemini Flash, Sol, and Fable — and as a legitimate free sandbox for CS students working on their own code with no UVU data. Not as a floor, and never on the contributor tier for real work.est

OpenRouter's free models — the limits, from the source

TermWhat OpenRouter statesGrade
Free-model rate limit20 requests per minute and 50 requests per day on an account that has never bought credits; 1,000 requests per day once at least $10 of credits has ever been purchased. Paid models carry no platform-level request cap.fact
Can a campus multiply accounts?No — by design and by contract. Docs: "Making additional accounts or API keys will not affect your rate limits, as we govern capacity globally." Terms (updated August 31, 2026) prohibit creating "multiple accounts as a single user, for purposes of bypassing or circumventing use limits" and reselling API access.fact
What happens to the data?"Some Models may store or train on your Inputs for improving their own large language models" per each model's terms; OpenRouter has opted out "where possible." On free models the upstream provider's terms govern — often training-permitted.fact
StabilityFree models rotate in and out without notice and are throttled hardest at peak hours (community reports across 2025–26); the free-tier limits themselves were tightened in 2025.est

The arithmetic at university scale

ScenarioMathVerdict
Phase-1 daily demand36,629 people × ~2.15 prompts per person per day (whole-population campus telemetry, §06) ≈ 78,800 prompts per day; exam-week stress ≈ 9× the base concurrency.The bar any "free" plan must clear.
One institutional OpenRouter account, funded1,000 requests per day ÷ 78,800 ≈ 1.3% of daily demand; 20 per minute ≈ 0.33 concurrent streams against a base need of ~53.A sandbox, not a service.
One free account per person50 per day each would technically exceed 2.15 per person — but the terms forbid organized multiplication of accounts, capacity is "governed globally," data terms are per-user consumer terms, and there is no admin control, audit, budget, or SSO.Not permitted, not governable, not private.
Muse Spark contributor tier at scaleZero dollars — and every prompt trains Meta's models. At 78,800 prompts a day the campus would be donating the largest free labeled dataset in Utah higher education.Only for code and questions with no UVU data in them.
Muse Spark standard tier in the frontier pool$1.25 / $4.25 per million tokens; a typical 1,500-token research exchange ≈ $0.006. For the 200-researcher pool, adding Spark to the routed mix changes yearly cost by tens of dollars either way.Belongs in the pool as a route; changes nothing about the floor.

What free tiers are genuinely good for

  • Sandboxes with no UVU data — a CS student trying a new model on their own code; a faculty member evaluating a release before the fleet qualifies it.
  • Model evaluation — free tiers are the cheapest way to benchmark a candidate against the fleet's current model before pulling its weights.
  • Teaching the routing lesson — the plan's data-class rules (§08) are exactly the discipline that makes free tiers safe: Public-class work may burst to labeled outside routes; nothing Sensitive ever does.
  • Never the floor. No published free tier offers institutional terms, an admin plane, an SLA, or a promise to exist next semester. The owned floor is what makes the free options safe to use at all.

Sources verified directly: Meta's model page and September 2 announcement; Artificial Analysis, September 2; OpenCode Zen documentation; OpenRouter's rate-limit documentation and Terms of Service (updated August 31, 2026).

Every route to Muse Spark 1.3 — and its terms

Prices are per million tokens: input / cache read / output. The pattern to notice: the same model is sold at two prices, and the cheap one is paid for with your data.

RoutePrice and limitsData termsEligibility, duration, institutional readingGrade
OpenCode Zen Contributor Free; raw ID muse-spark-1.3-contributor-free, OpenCode config ID opencode/muse-spark-1.3-contributor-free$0 / $0 / $0. Exact RPM, TPM, and request caps are UNKNOWN.Prompts and completions may train future Meta models. OpenCode says prompts pass upstream without being stored by OpenCode, but its general unpaid-account term also permits service-improvement use. Meta retention is UNKNOWN.“Limited time,” with no end date. General terms permit organizational accounts; no education program was found. Setup requests billing details. Post-free removal, conversion, and fallback are UNKNOWN. Zen, privacy, termsfact unknown
OpenCode Zen workspacePay-as-you-go credits; default auto-reload adds $20 when balance drops below $5 unless disabled. Card fee: 4.4% + $0.30.Admins can disable data-collecting models.Team workspace administration is free during beta; future price is unknown. Admin/member roles and per-member/monthly limits exist.fact
OpenCode Go; opencode-go/muse-spark-1.3-contributor$10/month; accounting rate $0.10 / $0.002 / $0.20. Usage-value caps: $12 per five hours, $30/week, $60/month. One subscribed member per workspace.Training: Yes. Zero-data-retention: No. Exact retention time unknown.Region limited by Meta policy. After limits, users can use free models or optionally fall back to paid Zen balance. Not a campus-wide entitlement. Go documentationfact
Meta Model API standard; muse-spark-1.3Meta-upstream public routes list $1.25 / $0.15 / $4.25. The unauthenticated Meta pricing page could not be read, so direct-account rate limits and max pricing remain UNKNOWN.Public gateway disclosures describe the standard route as non-contributor/no-training. Direct Meta retention and ZDR wording could not be verified anonymously.Public preview. Exact tier-based RPM/TPM and education programs are UNKNOWN. Meta release, OpenRouter, Artificial Analysisfact unknown
Meta Model API Contributor; muse-spark-1.3-contributorPublic routes list $0.10 / $0.002 / $0.20. Conflicting secondary rate-limit figures were found; none is safe to publish as exact.Meta may use prompts and completions to improve/train future products; not ZDR. Retention duration unknown.Geography page was login-gated. No education exception was found.fact unknown
OpenRouter standard and ContributorStandard: $1.25 / $0.15 / $4.25. Contributor: $0.10 / $0.002 / $0.20.Contributor route explicitly allows training. OpenRouter can enforce no-data-collection/ZDR routing only when a compliant provider is available; Contributor is not such a route.Both currently relay Meta as the upstream, not separately hosted weights. Standard, Contributor, provider privacy controlsfact
Vercel AI GatewayStandard public rate matches $1.25 input / $4.25 output; Contributor is also listed.Inherits the Meta route’s data distinction.Gateway, not an independent host. Vercel provider pagefact
Muse Code subscriptionsCurrent reporting lists $5 Everyday, $15 High, and $50 Power monthly plans. First-party subscription material was rate-limited during research.Plan-specific institutional data terms were not verified.Consumer/developer product, not a proven UVU entitlement.est
AWS Bedrock, Vertex AI, Azure AINo exact Spark 1.3 listing found.Search miss, not proof of global absence.unknown

Two more cautions from the survey: independent Terminal-Bench runs score Spark at 85–86 rather than Meta's 88.8, and most of Meta's launch comparisons used a "max" reasoning mode that was not generally available. And an instructive earlier incident: a misconfigured third-party test of Spark 1.1 gave the model internet access to a real target, and it changed real data — not a jailbreak, but proof that isolation and target validation are the harness's job, not the model's.

Longevity — the dated record behind the "free" offer

DateEventWhat it means for durabilityGrade
2023-02-24Original LLaMA used research-only, case-by-case access. MetaMeta’s meaning of “open” has changed across generations.fact
2023-07-18Llama 2 added broad commercial use under a custom license, a 700M-user gate, and restrictions on training other LLMs. LicenseOpen-weight did not mean OSI open source.fact
2024-04-18 to 2024-07-23Llama 3 retained restrictions; 3.1 loosened output use but added derivative naming duties. Llama 3, 3.1New generations can arrive with materially different terms. This is not evidence of retroactive relicensing.fact
2024-09-25 to 2025-04-05Llama 3.2 and Llama 4 added/retained EU multimodal restrictions. 3.2 policy, Llama 4Geography and use rights require version-specific review.fact
2025-04-29Meta launched hosted Llama API in limited free preview. MetaHosted Meta model access has precedent—but not permanence.fact
2026-07-06Reporting said Meta’s hosted Llama API closed after roughly 14 months. The official deprecation page existed but was unreadable anonymously; shutdown was not live-tested. Official page, contemporaneous reportSupports high endpoint-lifecycle risk, with incomplete primary proof.fact unknown
2026-03-01Older Llama repositories were consolidated; the Llama 3 repository was archived. RepositoryTooling/repository retirement did not withdraw downloaded weights.fact
2026-04-08 to 2026-09-02Four Spark releases in under five months: original, 1.1, 1.2, 1.3. Version intervals were 92, 27, and 28 days.Fast improvement, but poor evidence for long pinned-version life.fact
2026-08-10Meta released separate Muse Glimmer weights under Apache-2.0. MetaMeta can release Muse weights; Glimmer does not guarantee Spark weights or compatibility.fact
2026-09-02Spark 1.3 shipped while the Model API remained a public preview; future Spark weights remained only a roadmap item.No support window, overlap period, EOL notice, or SLA was found.fact
2025-04 to 2025-06Current OpenCode code line had a 0.0.1 tag on April 22; a founder interview says public launch was June 19. TagExact “first release” is unresolved; the project is about 17 months old.fact unknown
2025-09Zen was publicly visible during September; sources disagree on the exact launch day.Hosted gateway is about one year old.unknown
2026-02-06 to 2026-08-05Zen documents 18 model retirements across GPT/Codex, Claude, Gemini, GLM, Kimi, Qwen, and MiniMax. ZenDirect evidence of catalog churn.fact
2026-09-03OpenCode: about 203.4k stars, 26.5k forks, 15,650 commits, approximately 950–1,000 contributors, and a release on September 2. MIT client; company/core-team governance. Repository, licenseStrong project momentum and an easy fork/exit path.fact
2026-09-03Operator is Anomaly Innovations, Inc.; YC lists an active 24-person W2021 company. Predecessor SST raised $1M in 2021. Terms, YC, funding reportReal backing, but audited current finances are unavailable.fact
If the plan depended on…RiskWhy
Pinned Muse Spark 1.3 endpointHIGHPublic preview, four versions in five months, no lifecycle policy, closed weights, and prior hosted-API sunset evidence.
Future Spark weightsHIGH/UNKNOWNOnly a roadmap statement; no checkpoint, date, license, size, or hardware requirement.
OpenCode open clientLOW–MEDIUMMIT, unusually active, broad provider support, large contributor base, and readily forkable. Company-led governance remains a concentration risk.
OpenCode Zen/Go serviceMEDIUMYoung, upstream-dependent, terms allow limits/discontinuation, and the catalog changes quickly. Paid products, backing, and client portability reduce exit risk.

The harness landscape — what a university could actually stand on

A harness is the software people touch: the chat front door, the coding agent, the research tool. A powerful harness does not make a model smarter; it gives the model context, memory, tools, and authority — which is exactly why the choice matters for a campus. Snapshot September 3, 2026. Controls legend: I identity/SSO, B enforceable per-user budget, A audit/admin policy; Y verified, N absent, U not publicly proved. No product published a current accessibility conformance report — "a11y U" means unproven, not unusable.

Chat and assistant front doors

Harness · license · maturityModels and local fitControlsData default · accessibilityUniversity fitGrade
LibreChat; MIT; ≥2023; ~42.8k★/~5.4k commits; active 0.8.8 RC lineOpenAI-compatible, Ollama/local, major clouds; works cleanly behind LiteLLMI:Y OIDC/SAML/OAuth/LDAP; B:Y token-credit balances; A:Y, though admin UI is still marked PreviewApp RUM off and GTM only when configured; usage/cost transactions stored. Bundled Meilisearch analytics needs explicit disabling. A11y U, but current keyboard, focus, contrast, and screen-reader work is documented.Best broad front door. Keep the current plan; validate RC admin controls and telemetry settings.fact unknown
Open WebUI; custom Open WebUI License, BSD-3 base plus branding requirement above the threshold; not ordinary OSI open source; created 2023-10-06; ~150.8k★/~18.4k commitsOllama and any OpenAI-compatible endpointI:Y OIDC/LDAP/groups/RBAC; B:U; A:Y, audit export off by defaultNo outbound app telemetry by default; messages and usage remain in its DB. Community functions can execute unaudited Python. A11y U; high-contrast/text-scale controls exist.Strong alternative pilot if branding terms are accepted and LiteLLM enforces spend.fact unknown
AnythingLLM; MIT; ≥2023; ~65k★/~2.4k commits; Mintplex Labs, regular releasesGeneric OpenAI, LiteLLM, llama.cpp, Ollama, LM Studio, LocalAIBuilt-in multi-user accounts; I:U campus SSO; B:U; A:UAnonymous PostHog usage is on unless DISABLE_TELEMETRY=true; vendor says documents/chats are excluded. No-auth installs and admin-enabled unrestricted agents are supported. A11y U.Department/lab pilot, not first central choice.fact unknown
Jan; Apache-2.0; ≥2023; ~44.3k★/~8.6k commits; active Jan/Menlo projectOpenAI/Anthropic-compatible, Ollama, LocalAI, TGI, llama.cpp, LiteLLMDesktop I:N/B:N/A:N; separate Jan Server has Keycloak/OIDCLocal chats/settings/logs; no telemetry without consent. Cloud prompts go to the selected provider. A11y U.Good personal faculty/research desktop, not a campus front door.fact
Cherry Studio; AGPL-3.0 Community Edition; ≥2024; ~51.4k★; rapid 1.x→2.x cadence50+ providers, Ollama/LM Studio, MCP and RAGCE desktop I:N/B:N/A:N; enterprise detail insufficientAnonymous usage, errors, and crashes on by default; can be disabled, but upgrades reset the settings on. A11y U.Individual or department use until enterprise controls are proved in writing.fact unknown
Msty; proprietary per-user license; first changelog 2024-01-03; active 2.9.xOpenAI-compatible, MLX/llama.cpp/Ollama, cloud providersI:Y enterprise SSO; B:U; A:Y logs, allowlists, access controlsMsty says Studio has no analytics/telemetry. Enterprise retains team and audit metadata in a single-tenant service. A11y U.Plausible managed pilot, subject to procurement and security evidence.fact unknown
LobeHub, formerly LobeChat; LobeHub Community License, Apache-derived with commercial/distribution restrictions; ≥2023; ~82.2k★/~13.4k commitsMajor clouds, OpenAI-compatible, Ollama; now an agent workspaceI:Y generic OIDC options; B:U; A:UAnalytics packages exist, but no complete official default/payload statement was found. A11y U.Legal/security pilot only; license and control gaps outweigh feature breadth.fact unknown
Onyx, formerly Danswer; MIT Community Edition, separate enterprise license; ≥2023; ~31.9k★/~8.1k commits; company-backed, active v4Major clouds, Ollama, LiteLLM, vLLM; deep connector catalogI:Y, but CE/enterprise docs conflict; B:U; A:Y group/model/agent controlsSelf-host sends opt-out aggregate telemetry; other data stays in UVU infrastructure except selected providers. A11y U.Strongest governed knowledge-search candidate, after obtaining a versioned entitlement map.fact unknown
Khoj; AGPL-3.0; ≥2023; ~37k★/~5.2k commits; 2.0 beta, cloud deprecation announcedOpenAI-compatible, Ollama/vLLM/llama.cpp, major providersGoogle OAuth/magic link; generic OIDC/SAML U; B:U/A:UAnonymous telemetry on by default; disable with KHOJ_TELEMETRY_DISABLE=True; vendor says prompts and PII are excluded. A11y U.Personal/research-lab self-host only while service direction contracts.fact unknown

Coding and agent harnesses

Harness · license · maturityProvider postureControls and data postureUniversity fitGrade
OpenCode; MIT; ≥2025; ~203k★/~15.7k commits; Anomaly; extremely activeProvider-neutral: OpenAI-compatible, llama.cpp, LM Studio, Ollama, broad clouds; direct LiteLLM Meta supportEnterprise I:Y and central configuration; generic client B:U/A:U. Direct/BYOK prompts reportedly not stored; /share uploads a full session. Hosted unpaid terms allow improvement use. A11y U.Preferred CLI builder pilot. Disable sharing, use UVU keys/gateway, and ship a deny-first permission profile.fact unknown
Cline; Apache-2.0; ≥2024; ~67.4k★/~7.2k commits; company-backed, rapid releasesProvider-neutral; custom OpenAI-compatible endpoints, Ollama/LM Studio and major cloudsI:Y SAML/OIDC/RBAC; B:U; A:Y tool/model policy, cost telemetry and selective audit. Anonymous telemetry opt-in; prompt archive is explicit. A11y U.Top managed IDE pilot. Centrally disable automatic unrestricted execution and use sandboxes.fact unknown
Roo Code; Apache-2.0; ≥2024; ~24.3k★Historically provider-neutralAll products and maintenance ended 2026-05-15. Sunset noticeReject: discontinued.fact
Continue; Apache-2.0; ≥2023; ~35.7k★Historically provider-neutralRepository is read-only and not actively maintained; final 2.0 removed auth and telemetry.Reject: retired/archive only.fact
aider; Apache-2.0; ≥2023; ~48.7k★/~13.1k commits; slowing — last release Aug 2025, last code push May 2026 (GitHub API, Sept 3, 2026)Almost any cloud or local model, including Ollama/OpenAI-compatible routesI:N/B:N/A:N; enforce all controls at LiteLLM. Analytics asks consent; the selected prompt default is Yes, excluding code/chats/keys/PII. A11y U.Good expert tool, not a managed campus platform by itself — and watch its maintenance cadence before standardizing on it.fact unknown
Goose; Apache-2.0; ≥2025; ~53.9k★; 500+ contributors; Linux Foundation Agentic AI FoundationProvider-neutral; Ollama, LiteLLM, 15+ providers, MCP/ACPCentral I/B/A:U. Telemetry off by default—but prompt-injection checks and desktop sandbox are also off by default. A11y U.Strong sandboxed IT/CS/research pilot.fact unknown
OpenHands; MIT; ≥2024; ~86k★/~8.1k commits; company-backedAny LLM/BYOK; can supervise other coding agents; LiteLLM supportedEnterprise I:Y, SAML/SSO/RBAC; B:U; A:Y with event coverage U. Self-host can keep code/conversations at UVU; traces may contain tool and LLM I/O when enabled. A11y U.Enterprise automation pilot only, with containers and narrow repository credentials.fact unknown
Kilo Code; MIT core; ≥2025; ~27.2k★/~30k commits, some inherited fork history; acquired by AnacondaProvider-neutral; 500+ hosted, BYOK and local modelsI:Y SSO/SCIM/RBAC; B:Y user/sub-org limits; A:Y logs and model allowlists. Operational telemetry may contain linked diagnostic data. A11y U.Strong managed candidate after residency and telemetry review.fact
Codex CLI; Apache-2.0; public 2025-04; ~121k★/~10.2k commits; OpenAI, very activeOpenAI-first; --oss supports Ollama/LM Studio and custom Responses-compatible providersWorkspace I:Y/A:Y; client B:U. Business/Enterprise/Edu traffic is not trained on by default; Plus/Pro needs opt-out. A11y U.Strong controlled engineering/CS pilot, but operationally vendor-centered.fact unknown
Gemini CLI; Apache-2.0; 2025-06-25; ~106.8k★/~6.4k commits; Google, weekly stable cadenceVendor-tied: Gemini/Google/Vertex modelsGoogle identity; B:U in client; A:Y immutable policy, strict mode, MCP allowlist, OTEL. OTEL is off, but if enabled prompt logging defaults true. A11y U.Good Google-centered pilot, not a neutral campus standard.fact unknown
Claude Code; proprietary core; preview 2025-02-24; distribution/plugin repo ~144k★Vendor-tied: Claude through Anthropic, Bedrock, Vertex, Foundry, or a LiteLLM gatewayI:Y/A:Y; budget through provider/gateway. Retention is channel-specific; enterprise local-session transcripts can default to six years unless changed. ZDR may be available. A11y U.Controlled vendor pilot only after retention and route are contractually fixed.fact unknown

Research, retrieval, and workflow harnesses

Harness · license · maturityProvider supportControls, data, accessibilityUniversity fitGrade
Onyx/Danswer; MIT CE plus separately licensed enterprise; active v4Clouds, Ollama, LiteLLM/vLLM; 40+ enterprise connectorsIdentity available but CE/enterprise docs conflict; hard user budget U; model/agent/document controls available; aggregate telemetry opt-out; a11y UBest central library, policy, and department knowledge pilot. Require a written feature/entitlement map.fact unknown
Kotaemon; Apache-2.0; ≥2024; ~25.7k★ but only ~327 commitsOpenAI/Azure/Cohere, Ollama and llama.cppBasic multi-user login; OIDC/SAML, budgets, audit and telemetry defaults U; a11y UResearch prototype, not a central service.fact unknown
RAGFlow; Apache-2.0; ≥2024; ~90k★/~9k commits; InfiniFlow, active 0.27.xOpenAI-compatible, Ollama, LocalAI, vLLM/Xinference and cloudsOIDC present; budgets/audit/telemetry U; a11y U. Heavy deployment; official docs require gVisor when code execution is enabled.Advanced research/RAG pilot with dedicated platform engineering.fact unknown
Dify; modified Apache-2.0, with multi-tenant permission and branding restrictions; ≥2023; ~154k★/~24k forks; activeBroad provider/plugin catalog; OpenAI-compatible and self-hosted modelsEnterprise SSO and operation logs; hard user budget U. OTEL off by default; marketplace connections exist; complete analytics default U; a11y U.Strong app/workflow builder, but likely needs commercial licensing for campus multi-tenancy.fact unknown
n8n; Sustainable Use License, source-available rather than OSI open source; ≥2019; ~203k★/~60.5k forks/~23.8k commits1,500+ integrations, major model providers and OllamaEnterprise SAML/OIDC/LDAP, RBAC and log streaming; per-user AI budget U. Anonymous diagnostics on by default. A 2026 issue reports an external banner request despite isolation flags; a11y U.IT integration tier only, with least-privilege service accounts, approval gates, and outbound monitoring.fact unknown

Lifecycle matters more than popularity: Roo Code was discontinued in May 2026, Continue's repository is read-only, and Khoj's hosted direction is shrinking. Open WebUI, LobeHub, Dify, and n8n carry licenses that are not ordinary permissive open source and should not be presented as such.

What an agentic harness buys — and risks — in a classroom

Beyond a chat boxConcrete valueMain riskLikely users
Repository discovery and planningSearches files, follows imports, reads tests and plans multi-file changesPrivate code and secrets enter contextCS, IS, central IT, engineering
File creation and editingApplies patches, refactors modules, writes tests and documentationWrong or destructive changes; licensing/IP contaminationSoftware and data teams
Shell and test executionInstalls dependencies, runs builds, linters, tests, notebooks and CLIsHost compromise, malicious packages, resource exhaustionCS labs, engineering, analytics
Browser/web toolsReads documentation, investigates errors and interacts with web applicationsPrompt injection from pages; accidental action on real systemsIT, cybersecurity, research
MCP and business toolsCalls databases, ticketing systems, cloud APIs and internal servicesCredential misuse and cross-system data movementIS, business analytics, integration teams
Iterative agent loopObserves failures, revises its plan and continues without one prompt per stepRunaway token spend and compounding mistakesAdvanced builders and researchers
Git integrationProduces reviewable diffs, branches and commitsUnapproved pushes, releases or deploymentManaged development teams
RAG and permission-aware searchAnswers across policy, library, research and department repositoriesStale ACLs, poisoned documents and cross-unit leakageLibrary, research administration, service desks

The controls that make it safe — each one is already a design commitment of this plan:

  • Treat repository text, web pages, documents, issues, and MCP output as untrusted input. Prompt injection can become shell execution.
  • Run every student or researcher agent in an ephemeral container/VM. Do not mount the home directory or credential stores.
  • Deny outbound network access by default; allow only named package registries and approved endpoints.
  • Use short-lived, task-scoped credentials. Default repository access to read-only; deny deploy, publish, push, billing, and production tools.
  • Require a person to approve file writes outside the workspace, shell/network actions, package installation, and consequential tool calls.
  • Put all model traffic through LiteLLM for identity, model allowlists, per-user budgets, cost alarms, and emergency shutoff.
  • Keep FERPA, HR, finance, security, unpublished research, controlled data, credentials, and proprietary partner material out of contributor/free-training routes.
  • Set course-level rules for disclosure, allowed tools, retained work logs, individual contribution, and assessment design.
  • OpenCode’s current permission system supports allow/ask/deny rules, but most ordinary actions are permissive by default. UVU must ship its own deny-first profile. Permissions
  • The UnderSpecBench study found that agents guess action boundaries when operational tasks are underspecified. Assignment and run instructions must explicitly define allowed targets and actions. Paper

Three tiers, one router

TierRecommended harnessesLocal floor and frontier routeFree-tier policy and controls
EveryoneLibreChat as the supported front door; Onyx search exposed only for approved knowledge collectionsLibreChat → LiteLLM → Qwen3.8-27B local default. Use GLM-5.3 only on infrastructure proven to fit it. Metered frontier routes selected by task/data class.Do not expose Contributor Free in the general model menu. Use only institutionally accepted free entitlements with suitable no-training terms. Campus SSO, per-user budgets and logs are mandatory.
BuildersOpenCode default CLI; Cline managed IDE option; Goose or OpenHands for advanced sandboxed work; Kilo worth a procurement pilotHarness → LiteLLM → local model first; route hard coding tasks to paid non-contributor Spark, Sol, Opus, or Gemini according to evaluation and budgetContributor Free may appear only in a separate public-data sandbox with separate keys, a warning banner, no persistent connector credentials, disabled sharing, and no production access.
ResearchersOnyx for governed knowledge search; RAGFlow for advanced ingestion; Kotaemon for prototypes; Dify/n8n only as controlled workflow platformsLocal embeddings and local Qwen for protected corpora; paid no-training frontier route only when the data agreement allows itNo contributor/free-training route for source documents, unpublished work, grant material, participant data or licensed collections. Connector ACL tests, retention rules and source citations are required.

The layers stay separate on purpose: a broad chat portal, a shell-capable coding agent, and a credentialed workflow engine do not belong inside one permission boundary.

Verdict on Spark

Muse Spark 1.3 changes the frontier-pool candidate list, not the architecture.

OptionResultBenefitMain riskEase of undo
(a) Add paid, non-contributor Spark plus (b) a separate Contributor Free sandbox
est
Pilot muse-spark-1.3 behind LiteLLM for coding/agent tasks. Offer free Contributor only in an isolated public-data environment.Captures the strong model and the legitimate $0 experiment without trading UVU data for compute.Meta endpoint churn; first-party retention/rate terms still need procurement confirmation; sandbox users may paste restricted data.Easy if model names are abstracted behind LiteLLM and no workflow depends on Spark-only behavior.
(a) Paid non-contributor only
est
Cleanest institutional route.Simpler policy and lower data risk; still much cheaper than some frontier models.Gives up free experimentation; direct Meta terms still need confirmation.Easy.
(b) Contributor Free broadly
est
Free campus access while promotion lasts.Near-zero token spend.Training use, non-ZDR status, unknown limits/retention/end date, behavioral leakage, and weak durability.Technically easy, but data already disclosed cannot be recalled. Do not choose.
(c) Ignore Spark
est
Keep the present model pool unchanged.Least procurement and policy work.Misses a capable, relatively inexpensive coding route and useful competitive pressure.Easy.

The paid pilot should not start until Meta's or the gateway's current terms confirm, in writing: no training, retention period, geography, incident handling, rate limits, and permitted educational use. Survey: working paper 34, September 3, 2026.

Every provider with free access — the inventory

Twenty-eight providers checked on their own pricing, limit, and terms pages on September 3, 2026. "Free" turns out to mean six different things: a renewable allowance, a one-time signup credit, a thirty-day trial, a limited-time promotion, a data-for-discount trade, or nothing at all. The survey's headline: OpenAI, Anthropic, xAI, Together, DeepInfra, and Meta's own API have no recurring free generation tier; GitHub Models — the most generous free gateway of 2025 — was fully retired on July 30, 2026, after about 23 months.

ProviderFree access and capability anchorExact public limits and scopeReliability / practice signal
OpenRouterfact 18 literal free IDs above; random openrouter/free router.20 RPM; 50 RPD. After buying at least $10: 1,000 RPD, but no longer a zero-cash path. TPM/TPD/concurrency and exact org aggregation unknown. Limits, 2026-06-12Free model availability varies; current free endpoint showed 96.8% availability. Users report 429/500 errors and model switching.
Google AI Studio / Gemini APIfact Zero-price rows include Gemini 3.8/3.7/3.6/3.5 Flash; 3.5/3.1 Flash-Lite; Gemini 3 Flash Preview; Gemini 2.5 Pro, Flash and Flash-Lite; selected live/audio/TTS/transcription, embeddings, robotics, and Gemma models. Gemini 3.8 Flash: AA Intelligence 59 on 2026-09-02.Limits apply per project, not key. Exact free RPM/RPD/TPM are now visible only in AI Studio and are not guaranteed: unknown. Representative Flash context: 1,048,576 input/65,536 output. Rate limits, updated 2026-09-02Google staff confirmed a 20-RPD reduction in Dec. 2025 and said free access was not designed as a long-term application base.
Groqfact GPT‑OSS‑120B/20B, Qwen 3.6/3.8 27B, Compound/Mini, Prompt Guard, Whisper and Orpheus. GPT‑OSS endpoint: about 471 output tokens/s and 86% provider-test accuracy on 2026-09-03.GPT‑OSS/Qwen: 30 RPM, 1,000 RPD, 8K TPM, 200K TPD. Compound: 30 RPM, 250 RPD, 70K TPM. Limits are per organization. Context generally 131K. Current tableOfficial status operational. Llama 3.1 8B and Llama 3.3 70B free/developer access ended 2026-08-16.
Cerebrasfact GPT‑OSS‑120B and Gemma 4 31B. GPT‑OSS endpoint: about 1,643 output tokens/s and 87% test accuracy.Trial only: $5 credit, verified payment method, expires after 30 days. Each model: 5 RPM, 30K TPM, 1M tokens/hour and 1M TPD, per organization. Free/trial context 65K. Current limits100% trailing-90-day status on 2026-09-03. Former always-free service is no longer renewable.
SambaNovaconflicting sources Rate-limit docs list free DeepSeek V3.1, Llama 3.3 70B, GPT‑OSS‑120B, DeepSeek V3.2 and Gemma 4 31B. Its plans page says users must buy credits before the first request.Docs say 20 RPM, 20 RPD, 200K TPD per user and model; most contexts 128K. Current practical zero-payment access was not login-tested. Limits · conflicting plan pageStatus showed 99.93%–100% over 90 days. Treat free availability as unknown until a new account succeeds.
Together AIfact No ordinary free trial. A separate $150-credit page exists, but eligibility and current availability are unknown.Minimum $5 purchase; limits dynamic per organization/model. No-trial notice, updated 2026-06-01Serverless model availability ranged about 96.85%–100% over 30 days.
Fireworks AIfact $1 one-time signup credit across eligible serverless models. Capability depends on chosen model.Expiration, default RPM/TPM and concurrency are not public: unknown. PricingServerless status operational. Self-service changed to prepaid on 2026-07-01.
Mistral Studio / La Plateformefact Free plan provides $10/month in API credits; model access is dynamic. labs-* models are free, experimental, and may silently change. Medium 3.5 AA Intelligence 30; Small 4, 20.RPS/concurrency, TPM and monthly token caps are organization/model-specific and visible only after login: unknown. Current model contexts generally 128K–256K. Pricing · limitsAPI trailing-90-day uptime 99.395%; embeddings 94.337%. Labs are expressly not for production.
Hugging Face Inference Providersfact $0.10 monthly credit for Free users; model and upstream-provider choice is broad. Capability and context depend on the selected model.Dollar credit is exact; public RPM/TPM/concurrency are provider-specific and unknown. PricingInference Endpoints API showed 100% over 30 days, but upstream availability differs. Free credit was reduced from a larger launch allowance.
Cloudflare Workers AIfact Dynamic catalog within neuron allowance. Current free alternatives include GLM 4.7 Flash, Gemma 4 26B, Nemotron 3 120B and Llama 3.3 70B.10,000 neurons/day/account; general text 300 RPM; selected frontier models 20 RPM. Llama 3.3 70B example: about 125 standard turns/day. Pricing, updated 2026-08-28Workers AI operational; two incident-days in 30. Several frontier models became paid-only 2026-07-28.
GitHub Modelsfact None. Fully retired 2026-07-30.N/A. RetirementThe free gateway lasted about 23 months. Strongest direct warning against using previews as a campus floor.
NVIDIA build.nvidia.com / NIMfact Developer Program members get dynamic hosted “Free Endpoint” models for prototyping. Recent examples include Kimi K3, DeepSeek V4 Pro and Nemotron 3.5 Lightning.Public RPM/RPD/TPM/credit amount and context inventory: unknown. Community reports commonly see 1,000 credits and 40 RPM, but these are not contractual limits. Current programNo public hosted-NIM status page found. NVIDIA warns trial services and endpoints can change or end.
OpenCode Zenfact Big Pickle, MiMo‑V2.5, Ling 3.0 Flash Fin, Nemotron 3 Ultra, Nemotron 3.5 Lightning, Muse Spark 1.3 Contributor and Muse 1.2 Contributor are listed free. Live API also exposed two other free IDs. Muse 1.3: AA Intelligence 62 on 2026-09-02.RPM/RPD/TPM/TPD/concurrency/context: unknown. Every listed free model is limited-time. Zen, updated 2026-09-03Official-repository issue reported intermittent free endpoint unavailability on 2026-08-25.
Meta Model API / Musefact No verified direct recurring free tier. Muse 1.3 Standard is paid; Contributor is deeply discounted and trades data rights for price. Zen temporarily supplies Contributor free.Direct current public limits and retention pages required login: unknown. OpenRouter shows 1,048,576 context. Muse 1.3 AA Intelligence 62. Meta release, 2026-09-02Product moved from private preview in April to public preview in July 2026. Durable paid availability is more likely than durable free availability.
Ollama Cloudfact Recurring monthly Starter allowance and a smaller Starter model set, both undisclosed.One concurrent request; additional requests queue. Monthly allowance, RPM/RPD/TPM and model roster: unknown. PricingCloud started 2025-09-19; limits were redesigned 2026-08-31. Users reported large free models unexpectedly moving behind payment.
DeepInfrafact No current free or trial allowance found. PAYG only.Paid default is 200 concurrent requests/model; capacity can still return 429. LimitsAPI status 100% over 74 observed days; individual models about 98.78%–99.96%.
Novita AIfact New-user voucher of unspecified value; some zero-price development models such as Llama 3.3 3B and Qwen 2.5 7B. Free-specific capability anchor unknown.Amount, expiry, RPM/RPD/TPM and permanence: unknown. QuickstartModel status about 98.4%–100%; zero-price model inventory changes.
Hyperbolicfact $1 one-time credit after phone verification; broad paid model catalog.60 RPM on basic accounts; Pro 600 RPM after depositing $5. Token allowance depends on model price. LimitsNo incident in prior seven days; no usable long-horizon uptime figure. Older $5/$10 signup pages are stale.
Chutesfact No current free API. Its 200-RPD Early Access plan ended 2026-03-15.N/A; current service is PAYG or paid subscription. Retirement notice, 2026-02-27API showed 100% over 30 days, but the free path is gone.
Klusterunknown No current public self-service general inference offer or current free-price table found.Limits, model roster and present signup credits: unknown. Historical $25 credits expired in Feb. 2025.No public current status evidence found. Do not use cached 2025 offers.
Cohere trial/evaluationfact Trial keys provide access to evaluation models; Command-family availability is account-dependent.1,000 API calls/month; Chat 20 RPM; Rerank 10 RPM. Non-production only. LimitsOfficial status operational. Trial allowance fell from 5,000 calls/month in 2023 to 1,000.
xAI APIfact No current free tier. The $25/month launch promotion ended in 2024.Prepaid; $25 minimum top-up. Tier 0 means zero prior spend, not free inference. BillingOfficial status includes API and billing incidents.
Anthropic APIfact No renewable tier. New users receive a small, unspecified one-time test credit.Evaluation limits and credit value are unknown; paid limits are organization/model based. PricingOfficial 90-day API uptime was 99.52%; selective research grants exist.
OpenAI APIfact No general-generation free tier. One test request and free moderation only.omni-moderation-latest: 250 RPM, 5,000 RPD, 10K TPM. GPT generation Free tier is not supported. Model pageDurable paid API; free generation cannot be planned as capacity.
Vercel AI Gatewayfact Every free team receives $5 every 30 days, usable across the catalog. GPT‑OSS‑120B is 131K context and starts at $0.10/M input, $0.50/M output.No public free RPM/concurrency cap found. At a 700-input/300-output turn: about 22,727 turns/30 days. Pricing · model pageGateway adds retries/fallbacks, but upstream status and terms remain relevant.
Z.AIfact GLM‑4.7‑Flash and GLM‑4.5‑Flash are free; both advertise 200K context.Numeric rate limits appear only in the account console: unknown. Model overviewUsers report tight dynamic throttling. Free availability is a conversion path to paid coding plans.
Alibaba Model Studiofact New users receive model-specific quotas, typically 1M tokens/model. OAuth separately advertises 2,000 calls/day. Qwen Max/Plus/Flash models are included depending on region.Account and RAM users share quotas. Signup quota is time-limited; a 90-day policy takes effect 2026-09-08. Re-registering does not issue new quota. Free quota, updated 2026-09-01Trial, not durable floor. Regional restrictions and rapid model sunset notices apply.
IBM watsonx.aifact Lite provides up to 300,000 foundation-model tokens/month and open-model playground/API access. Current catalog includes Granite and selected third-party models.300K combined tokens/month; one Lite instance and one authorized user. RPM/concurrency/context depend on model: unknown. Runtime plan, updated 2026-07-30Evaluation plan; idle Lite instances may be deleted after 30 days. The token allowance increased in July 2025.

Data terms and account rules — in the providers' own words

ProviderFree-tier data treatmentAccount / university-use rule
OpenRouterRouter logging is off by default; upstream providers still differ. ZDR routing is available.Organizational authorized users are supported, but terms prohibit creating multiple accounts to bypass limits: “create multiple accounts…for purposes of bypassing or circumventing use limits.Terms, updated 2026-08-31
GoogleUnpaid prompts and responses may improve Google products and may be human-reviewed. Retention for ordinary unpaid calls is unknown.18+; no under-18-directed API client. “Do not submit sensitive, confidential, or personal information to the Unpaid Services.Additional Terms, effective 2026-03-23
GroqNo training without permission; inference is not retained by default. Reliability/abuse logs can last 30 days; ZDR is available.Organization-wide quota. “beyond published…limitations, including by registering multiple accounts or orchestrating usage between multiple organizations.AUP, effective 2025-10-15
CerebrasDoes not train on or retain ordinary inference I/O; operational retention duration is not stated.Per organization; verified payment required. “Buy, sell or transfer API keys without our prior written consent.Terms, effective 2024-08-27
SambaNovaEULA limits customer-content processing to service delivery or law; staff says prompts are not stored or trained on. Output/log retention unknown.Free docs say per user. Access may be given only to authorized users; resale or third-party availability is prohibited.
TogetherInputs/outputs are not stored by default; temporary caching can occur; training is opt-in.Project keys may be shared only with trusted collaborators and have full project rights. Multi-account-evasion clause unknown.
FireworksOpen-model I/O is not logged or stored by default unless the customer opts in.Buying, selling, or transferring API keys is prohibited without prior written consent.Terms
MistralStandard API I/O is not used for training and normally has 30-day abuse-monitoring retention. Labs/Preview can train and offer no opt-out.Limits are per organization. Buying, selling, or transferring keys/accounts is prohibited.
Hugging FaceHF does not store bodies/responses; content-free debugging logs remain up to 30 days. Upstream terms also apply.Organization tokens exist. Explicit multi-account evasion rule unknown.
CloudflareNo model training or service improvement without explicit consent; persistence occurs only if the customer invokes storage products.Terms prohibit quota circumvention and “automated agents or scripts to create multiple accounts.Terms, updated 2025-09-12
GitHub ModelsRetired. Historical service said model providers did not receive training rights.One free account per person/entity; users could not share tokens to exceed limits.
NVIDIATrial terms can permit deidentified content to improve services; fine-tune input can persist 30 days and generated content 90 days.Keys are tied to one user. Trial access may not be transferred or made available to third parties.
OpenCode ZenOpenCode says it does not store prompt content, but free upstreams have exceptions. Muse Contributor explicitly permits Meta training.Terms prohibit bulk or multiple accounts that evade usage or promotions.
Meta MuseContributor permits training on prompts/completions. Standard is described as no-training, but current primary retention terms were login-blocked.Direct account-sharing, limit scope, and education terms: unknown.
Ollama CloudPrompts/responses are transient, not retained after fulfillment, and not used for training.Ollama is one account per person.” Teams require a separate shared account. Pricing
DeepInfraI/O is not stored or trained on without consent; support/security exceptions may last 30 days.Scoped keys and spending caps exist. Explicit multi-account rule unknown.
NovitaTerms state default ZDR, transient processing, and no training; account metadata has longer retention.Multiple accounts for extra quotas are prohibited.
HyperbolicInference input is transient, discarded, and not used for training.Account/password sharing and multiple-account limit evasion are prohibited.
ChutesNo current free tier. Public-model I/O is not stored, logged, or trained on; third-party chutes may differ.Multiple subscriptions used to evade caps are prohibited.
KlusterCurrent public self-service data terms: unknown. Enterprise page claims zero prompt logging.Current consumer/inference account rules: unknown.
CohereTrial I/O may support R&D, performance, and safety work. Trial users are told not to submit personal data.One account; user IDs cannot be shared or used to bypass limits.
xAIAPI data is not used for training without permission; default retention is 30 days; eligible teams can use ZDR.Team-scoped limits and confidential credentials; explicit multi-account clause unknown.
AnthropicCommercial/API I/O is not used for training by default. Public pages conflict between default non-retention and deletion within 30 days.Each user/service should have its own key; automated account creation and ban evasion are prohibited.
OpenAIAPI data is not used for training by default; abuse-monitoring logs can contain content for up to 30 days. Approved organizations can request ZDR.Keys must remain secret; institutional projects/service accounts are supported.
VercelVercel says Gateway I/O is not stored or trained on; upstream policies still apply. ZDR/no-training routing controls exist.AI terms allow authorized users but prohibit sensitive personal information in input. Mass free-team farming is not expressly authorized.
Z.AIA portion of request content may be cached; cache retention is undisclosed. Free-model training treatment: unknown.Paid coding plans are individual and cannot be shared. General free-API account rules are unknown.
AlibabaFree-tier training and detailed retention terms were not found: unknown.One Alibaba account and its RAM users share quota; new accounts cannot be re-registered for another grant.
IBM watsonx.aiLite-specific prompt-training and retention terms were not clearly stated on the pricing/plan pages: unknown.Lite is limited to one instance and one authorized user; it is an evaluation plan, not a campus entitlement.

The capacity math, done properly

One "turn" below is one prompt-and-response of 700 input and 300 output tokens; a four-turn chat is four turns. Phase-1 demand: 36,600 people × 2.15 turns per day ≈ 78,690 turns per day; exam-week stress ≈ 708,000. The three most generous recurring, publicly measurable free offers, combined:

Provider / modelBinding free limitOne institution-controlled account36,600 genuine personal accounts (raw maximum)
Vercel / GPT‑OSS‑120B$5 per 30 days; $0.00022/standard turn22,727/30 days = 758/day average27,742,800/day
Groq / GPT‑OSS‑120B200K TPD; 1K tokens/turn200/day7,320,000/day
Cloudflare / Llama 3.3 70B10K neurons/day; 80.109 neurons/turn124/day4,538,400/day
Combined1,082/day39,601,200/day
ScenarioBase demand coveredExam-stress demand covered
One UVU-controlled account/org at all three1.38%0.153%
36,600-account raw multiplication50,326% raw capacity5,592% raw capacity
Genuine, self-managed personal accounts100%, subject to eligibility and each person’s terms100%, because 19.35 turns/person/day is below each allowance
Centrally created/poolable account farm100% only by violating or evading termsMathematically sufficient, contractually unusable

The answer in one sentence: free tiers cover about 1.4% of the campus's base day as a centrally operated service; the only way to 100% is to centralize thousands of personal accounts, which the providers' terms forbid — and the daily numbers are generous, because the same providers cap requests per minute at 20–30 with no promise of burst capacity against a base of ~53 simultaneous conversations and an exam-week peak near 474. Separately, a real person's own free account can cover that person's own demand — a genuine personal benefit, but not campus infrastructure UVU can promise, audit, or revoke.

What breaks at university scale

  1. Account farming is not a lawful architecture. Groq expressly prohibits orchestrating usage across multiple organizations. Cloudflare prohibits scripted multi-account creation and quota circumvention. OpenRouter and OpenCode prohibit multiple accounts used to bypass limits.
  2. A real person’s account is not UVU capacity. The person controls consent, settings, deletion, model selection, and account continuity. UVU cannot promise access, audit activity, revoke all access centrally, or move unused personal quota into a campus pool.
  3. 36,600 credentials become a security system of their own. Providers say keys must remain confidential and must not be embedded client-side. Central collection would create a high-value credential store; sharing keys often violates the applicable terms.
  4. Sensitive data leaves campus under inconsistent terms. Google explicitly warns against confidential, sensitive, or personal information in its unpaid API. Muse Contributor permits training. Cohere trial data may support R&D and safety. NVIDIA trial terms allow some service-improvement use.
  5. FERPA needs control, not merely a privacy promise. The FERPA school-official exception requires the institution to exercise direct control over the vendor’s use and maintenance of education records. Personal consumer accounts normally do not create that relationship. U.S. Department of Education FERPA guidance
  6. UVU’s own rules follow the data wherever it is stored. UVU classifies grades, schedules, disciplinary, financial, payroll, research and similar data as Sensitive/Confidential, including when a contractor or cloud stores it. UVU data classification · Policy 445, effective 2024-03-28 · Policy 447, effective 2026-08-10
  7. Free services lack institutional controls. Some have team dashboards, but personal free accounts generally lack UVU SSO, SCIM, domain claims, legal hold, central retention, eDiscovery, audit export, data-region choice, and contract remedies.
  8. Models rotate silently or with short notice. OpenRouter’s router chooses among available models; Mistral Labs permits silent updates; NVIDIA trial endpoints can change; OpenCode says free models are limited-time. GitHub Models demonstrates that an entire service can disappear.
  9. No SLA means no floor. Status histories show good averages for several vendors, but free/trial terms do not promise capacity during registration, class-change bursts, finals, or provider-wide shortages.
  10. Age and consent complicate rollout. Google’s developer API requires users to be 18+. Groq’s current service agreement is also adult/business focused. UVU still has some under-18 students.

Governance conclusion, not legal advice: Public-class prompts may use personal or free tiers after clear notice; anything Sensitive or Confidential needs an approved institutional service with a data-processing agreement and a reviewed processing chain — which is what the owned floor is.

Free, metered, or owned — the honest three-way comparison

ChoicePrivacy and termsPredictability and controlThree-year cost at modeled base demand
Free cloud tiersMixed: from no-training/ZDR to explicit training and human review. Mostly self-service terms.Lowest. Quotas, models and access can change; little SLA or central administration.$0 inference, but the three measurable providers cover only 1.38% of base demand centrally. Integration, support and governance still cost staff time.
Metered hosted open APIsPaid terms usually offer no-training defaults, better ZDR choices, DPAs and institution-owned keys. Requests still leave campus and may traverse a gateway plus upstream provider.Stronger: central gateway, budgets, rate controls, audit, fallbacks and negotiated capacity. Vendor/model churn remains.est 86.17M standard turns over three years. GPT‑OSS‑120B raw inference is about $6.6K at a low OpenRouter route, $19.0K through Vercel’s displayed route, or $24.6K at Groq’s rate. With a 30% reserve: roughly $8.6K–$31.9K, excluding gateway and labor.
Current owned Apple fleetPrompts remain on campus if all routing, storage and telemetry are local. Best control over retention and logs.Strongest governance and fixed local capacity, but UVU owns uptime, upgrades, security, model qualification and staffing.est 26 M5 Pro/48GB minis plus kit: $63.9K upfront / $79.6K three-year machine TCO, excluding labor. Rated around 104 streams: enough for the 53-stream base, not the 474-stream whole-campus exam case.
Owned fleet scaled to whole-campus exam peakSame local benefits.Fixed 9× peak capacity, much of it idle on ordinary days.est ceil(474÷4)+1 spare = 120 minis; roughly $367K three-year machine TCO by linear scaling, before larger network/rack/operations costs. This is a planning bound, not a quote.

The finding worth sitting with: at the observed prompt rate, a cheap hosted open model can cost less than the machines — roughly $9k–$32k over three years for inference alone against about $80k of owned hardware. The owned floor earns its place through privacy, continuity, and control, not because local inference is always cheaper. The guide has said this since its first version (§12); this survey puts numbers under it from the free-provider side.

Longevity — twenty dated events

DateEventWhat it proves
2023-09-27Cloudflare launched Workers AI for all plans and warned access/limits could change.A recurring allowance can last, but model eligibility is movable.
2024-07-29NVIDIA launched free hosted NIM credits for Developer Program members.Durable prototyping funnel; no production promise.
2024-08-01GitHub Models public preview launched.Start of a free gateway later retired.
2024-08-27Cerebras launched an always-free inference tier.Historical free promise did not survive unchanged.
2024-09-10SambaNova Cloud launched with free Llama 3.1 405B access.Model later disappeared; catalog churn is real.
2024-09-17Mistral launched its free experimentation tier.Now framed as monthly credits plus experimental Labs.
2024-11-04xAI offered $25/month only through the end of 2024.Promotion had an explicit end and is not current.
Feb. 2025Hugging Face launched Inference Providers with a larger small allowance.Current Free allowance is only $0.10/month.
Apr. 2025Users reported OpenRouter free limits falling from 200 to 50 RPD.Anecdotal exact cut; current 50-RPD primary limit is confirmed.
2025-07-10OpenRouter said two providers withdrew free capacity and that it was subsidizing replacements.Free availability depends on upstream marketing budgets.
July 2025Together ended ordinary free trials.Mature inference shifted to prepaid use.
2025-09-19Ollama Cloud preview began.Free cloud history is only about one year.
2025-12-07Google staff confirmed a 20-RPD free reduction.Google free API quotas can shrink without supporting production use.
2026-03-15Chutes retired its 200-RPD Early Access plan.Loss-making free inference was removed.
2026-07-30GitHub Models fully retired.Whole free gateway can disappear, not just one model.
2026-07-28Cloudflare moved Kimi/GLM/DeepSeek frontier models to paid access.Free allowance remained, but strongest models moved out.
2026-08-16Groq ended free/developer access to Llama 3.1 8B and Llama 3.3 70B.Model availability is not durable even when the provider survives.
2026-08-31Ollama changed usage accounting to monthly token pools.Capacity mechanics can change quickly.
2026-09-03OpenCode lists Muse Spark 1.3 Contributor Free as limited-time.Today’s attractive offer carries no longevity promise.
2026-09-30OpenRouter’s Dots 3 Note free model is scheduled to expire.Current catalog contains explicit near-term expiry.

Reading the business models: small renewable credits that lead naturally to pay-as-you-go (Vercel, Mistral, Cloudflare, IBM) are the more durable kind; subsidized GPU giveaways to win developers (OpenRouter free routes, NVIDIA's hosted trial, OpenCode promotions) are the less durable kind; one-time signup credits are not capacity at all. GitHub, Chutes, and xAI ended their free offers rather than maintaining them.

Verdict, provider by provider

ProviderBest legitimate UVU roleHard limitData-term riskLongevity riskRecommendation
OpenRouterEvaluation, Public-data fallback, model discovery50 RPD/20 RPMMED: upstream variesHIGH: subsidized rotation and 2025 withdrawalsUse experimentally; pay for controlled routing if adopted.
Google Gemini API FreeCS labs and voluntary experimentsExact quota private/per projectHIGH: unpaid data may train; human reviewHIGH: staff-confirmed quota cutsNever send Sensitive data; do not size infrastructure from it.
Groq FreeFast open-model labs and public prototypes200K TPD/org for strong text modelsLOW: no training, ZDR availableMED: free models rotateBest direct free API, but not a campus floor.
CerebrasThirty-day performance evaluation$5/30 daysLOWHIGH: renewable tier already endedTrial only.
SambaNovaConditional model-speed lab20 RPD/model in docs, but signup conflictMEDHIGH: current access contradictedVerify manually before any course promise.
TogetherMetered open-model providerNo ordinary free tierLOW–MEDHIGH for freeConsider paid, not free.
FireworksOne-time benchmark/evaluation$1 creditLOW–MEDHIGH: finite creditTrial only; paid candidate.
MistralEU-hosted model evaluation and course labs$10/month; private rate capsMED; Labs HIGHHIGH: hidden caps/silent Labs updatesUseful sandbox; contract/pay for production.
Hugging FaceSmall interoperability tests$0.10/monthMED: upstream termsHIGH: allowance was cutCatalog/research hub, not capacity.
CloudflarePublic-data edge inference and burst valve10K neurons/day/accountLOWMED: recurring since 2023, but frontier gatingStrong free complement; paid plan for reliability.
GitHub ModelsNoneRetiredN/AHIGH: already goneRemove from all plans.
NVIDIA hosted NIMPrototype models before self-hostingNumeric allowance unknownHIGH for trialHIGH: pre-release/changeableEvaluation only.
OpenCode ZenDeveloper evaluation, especially MuseLimits unknown; limited-timeHIGH: contributor trainingHIGHPublic/synthetic prompts only.
Meta Model APIPaid Muse evaluationNo direct verified free tierContributor HIGHHIGH for freeTreat Contributor as paid data-for-discount, not free.
Ollama CloudPersonal experimentation1 concurrent; allowance undisclosedLOWHIGH: recent limit redesignGood personal benefit; cannot capacity-plan.
DeepInfraCheap metered open APINo free tierLOW–MEDHIGH for freeSerious paid candidate.
NovitaSmall-model developmentUnknown voucher/free limitsLOWHIGH: dynamic zero-price modelsSandbox only.
HyperbolicShort evaluation$1, 60 RPMLOWHIGH: one-timeTrial only.
ChutesPaid specialist inferenceFree plan retiredLOW–MEDHIGH: already endedDo not count as free.
KlusterNone pending evidenceCurrent service unknownUNKNOWNHIGHExclude until vendor supplies current terms.
Cohere TrialCourse demonstration/evaluation1,000 calls/month; non-productionHIGH: trial may support R&DHIGH: allowance cutSynthetic/Public data only.
xAIPaid Grok testingNo free tierLOW–MEDHIGH for freePaid-only consideration.
AnthropicResearch grants or paid ClaudeNo recurring free APILOW commercialHIGH for freePursue grants; price Claude for Education separately.
OpenAIResearch grants, Codex student benefit, paid APINo free generation tierLOW commercialHIGH for freeUse education/research programs, not imagined API quota.
Vercel AI Gateway$5/month model evaluation and fallback testingDollar cap; throughput unknownMED: upstream chainMED: renewable conversion funnelStrong missed free option; institutional paid use needs review.
Z.AIPublic/synthetic GLM experimentsQuotas privateUNKNOWN–HIGHHIGH: free conversion pathExperimental only.
Alibaba Model Studio90-day Qwen course/research trialUsually 1M tokens/modelUNKNOWNHIGH: new-user grantTime-boxed course or benchmark.
IBM watsonx.ai LiteGranite/open-model teaching labs300K tokens/month; one userUNKNOWNMED–HIGH: evaluation planWorth a controlled faculty pilot, not campus capacity.

Survey: working paper 33, September 3, 2026 — every limit read from the provider's own page, every quoted clause from its terms; "unknown" marks where a number lives only behind a login. No accounts were created and no requests were sent; the numbers are what the providers publish, not what they deliver under load.

M

The harness market, extensively

171 tools across eight categories — what is adopted, what is rising, what is funded, what is dying — and how each fits this plan

Appendix M in five lines

  1. QuestionWhich chat, coding, research, and work tools are used, rising, or dying, and which fit the university?
  2. AnswerThe review supports the current core tools while new tools are tested around them.
  3. Deciding numbers171 toolsest86 code projects checkedest4 archived projects and 9 with no code update in over 90 daysest
  4. What the plan doesKeep the current core, test new tools for 90 days, and reject shut-down or hype-only tools.
  5. Still unknownFast risers still lack proven users, owners, safety, and campus controls; the social scan found no campus-wide university use.

Appendix L compared twenty-five harnesses in depth. This appendix widens the lens to the whole market as of September 3, 2026: 171 products and projects, each with its form, license and first release, owner or backer, public scale, growth and release cadence, disclosed capital, stability signals, any rating or benchmark placement, provider and local-model support, identity and budget controls, data and telemetry defaults, and a university verdict. Three feeds were merged: a dedicated survey lane over GitHub trending, awesome-lists, marketplaces, package registries, funding announcements, and press; a live pull from the GitHub API for 86 repositories (below); and an X trend scan whose findings appear as their own section below. Every cell carries its grade: fact a source states it, est derived, unknown not independently verified. Stars are an interest signal, never proof of adoption, quality, or safe operation.

The short version

  • Reach is concentrated in the big consumer and enterprise assistants — ChatGPT, Gemini, Meta AI, GitHub Copilot, Microsoft 365 Copilot, Replit, Gemini Notebook, Claude. Open-source attention concentrates in Ollama, AutoGPT, Dify, Langflow, Open WebUI, LangChain, llama.cpp, OpenCode, and n8n.
  • Funding is concentrating around foundation-model vendors and coding agents: OpenAI, Anthropic, Cognition, Replit, Lovable, Cursor, CodeRabbit, Ollama, Braintrust, Qodo.
  • The sharpest new GitHub spikes — DeepSeek Harness, Ponytail, Graphify, Caveman, ECC, Paperclip, OpenMAIC, openclaude — need provenance and real-usage validation before any procurement conversation; several show star patterns no organic project produces.
  • Clear retirement or migration risk: Flowise, Roo Code, Continue, OpenAI Swarm, the old AutoGen and Semantic Kernel agent paths, Sourcegraph Cody's retired tiers, Amazon Q CLI, and the consumer Gemini CLI path.
  • The plan's base still maps well: LiteLLM as the routing boundary, LibreChat as the front door, llama.cpp as the floor, OpenCode/Cline-class coding, Onyx-class research. The market evidence argues for testing more harnesses around that boundary, not replacing it.

Four rankings — how the market is moving

All four are market-motion scores on a 0–100 scale est, not quality, security, or procurement scores. Method:

  • Adoption: 55% active/paid-user evidence, 20% installs or package use, 15% repository reach, 10% organization/revenue evidence. Missing dimensions were renormalized. Vendor claims received a confidence discount.
  • Growth: 70% percentile rank of the best recent user, install, download, or star-growth measure; 30% recency and evidence quality. Shorter intervals were annualized only for ranking and are shown explicitly.
  • Funding momentum: 65% log-scaled disclosed amount, 20% recency, 15% strategic or revenue evidence. Acquisitions without a disclosed price received no invented amount.
  • New and rising: limited to products first released or materially relaunched within nine months; 60% recent velocity, 20% age, 20% release activity.
  • These are market-motion scores, not security, quality, accessibility, or procurement scores.
#Absolute adoption90-day growth (or closest public interval)Funding momentumNew and rising, ≤9 months
1est 100 ChatGPT — >1B WAUest 100 DeepSeek Harness — +210,708 stars/22d; anomalyest 100 OpenAI — $122B, 2026-03-31est 100 DeepSeek Harness — 210,708 stars/22d; validation required
2est 98 Gemini — >1B MAUest 96 Ponytail — 122,907 stars/84dest 98 Anthropic — $65B, 2026-05-28est 96 Ponytail — 122,907 stars since 2026-06-12
3est 92 Meta AI — 700M MAU, stale 2025 evidenceest 91 Dify — ≈16.8K stars/month over 97dest 88 Cognition — >$1B, 2026-05-27est 93 Graphify — 114,218 stars since 2026-04-03
4est 90 GitHub Copilot — 50M usersest 88 OpenCode — ≈12K stars/month over 118dest 86 Lovable — $400M, 2026-08-12est 91 Caveman — 102,927 stars since 2026-04-04
5est 88 Replit — >50M platform usersest 86 Codex — +30,401 stars/86dest 84 Replit — $400M, 2026-03-11est 89 ECC — 246,773 stars since 2026-01-18
6est 86 Microsoft 365 Copilot — >30M paid seatsest 84 OpenMAIC — +9,426 stars/7dest 82 Cursor/SpaceX — acquisition 2026-08-14; prior $900M roundest 87 Paperclip — 79,934 stars since 2026-03-02
7est 85 Gemini Notebook — >30M usersest 82 qm — +6,712 stars/7dest 80 CodeRabbit — $143M, 2026-08-12est 84 OpenMAIC — +9,426 stars/week
8est 84 Claude/Claude Code — 39% survey share; use doubledest 80 Cloudflare OS — +9,549/monthest 78 Ollama — $88M, 2026-07-09est 82 openclaude — +1,035/week; 32,228 stars
9est 82 LangChain — 229.9M PyPI downloads/monthest 79 Cloudflare Computer — +8,868/monthest 76 Braintrust — $80M, 2026-02-17est 80 nanobot — 47,685 stars since 2026-02-01
10est 81 Lovable — 60M projects; 900M visits/monthest 76 Paperclip — ≈4,962 stars/monthest 74 Qodo — $70M, 2026-03-30est 78 codebase-memory-mcp — 42,025 stars since 2026-02-24
11est 80 Codex — 1.6M weekly usersest 75 Browser Use — ≈4,860 stars/monthest 72 Arize — $70M, 2025-02est 76 QwenPaw — 34,846 stars since 2026-02-24
12est 78 Cursor — 12% survey; hundreds of millions weekly requestsest 74 llama.cpp — ≈4,260 stars/monthest 70 Gumloop — $50M, 2026-03-12est 74 Prime Agent — 19,733 stars since 2026-05-08
13est 76 Vercel AI SDK — 92.1M npm downloads in Augest 72 Browser Harness — ≈3,730 stars/monthest 68 Consensus — $30M, 2026-05-11est 72 Cloudflare OS — 9,539 stars since 2026-04-15
14est 74 OpenCode — 7% survey; 1.85M npm/weekest 70 OpenHands — ≈3,330 stars/monthest 67 Dify — $30M, 2026-03-10est 71 Cloudflare Computer — 8,954 stars since 2026-06-05
15est 73 Ollama — 8.9M developers claimedest 69 Open WebUI — ≈3,100 stars/monthest 66 Mastra — $22M, 2026-04-09est 70 qm — 14,523 stars since 2026-07-29
16est 71 CrewAI — 26.5M PyPI/monthest 64 RAGFlow — ≈2,520 stars/monthest 65 LangChain — $125M, 2025-10-20est 66 OpenHarness — 15,633 stars but already stale
17est 70 Cline — >5M installsest 63 Ollama — ≈2,430 stars/monthest 63 OpenHands — $18.8M, 2025-11-18est 64 MiMo Code — ≈4,580 stars/month
18est 68 Strands — 14M downloads claimedest 59 CrewAI — ≈1,470 stars/monthest 61 Cline — $32M, 2025-07-31est 62 Kimi Code — 7,238 stars since 2026-05-22
19est 67 Dify — 1.4M machines claimed; 154K starsest 58 Cline — ≈1,450 stars/monthest 60 Glean — $150M prior round plus $300M ARRest 60 open-connector — 5,524 stars since 2026-06-29
20est 66 Open WebUI — 150.8K stars and strong container useest 57 Langflow — ≈1,360 stars/monthest 58 Browserbase — $40M, 2025-06est 58 fx — new 2026-08-11
21est 65 n8n — 203K stars; strong commercial growthest 54 Crush — ≈952 stars/monthest 56 Relevance AI — $24M, 2025-05est 57 Browser Harness — 17.3K stars since 2026-04-17
22est 64 LangGraph — 13.15M npm downloads in Augest 53 Aider — ≈872 stars/monthest 54 LlamaIndex — $19M, 2025-03est 56 Antigravity CLI — 6% survey within four months
23est 63 Glean — $300M ARR and 45% wDAU/wMAUest 51 Qwen Code — ≈804 stars/monthest 52 Onyx — $10M seed, 2025est 54 IBM Bob — GA 2026-04; 80K internal IBM users
24est 62 Elicit — >400K monthly researchersest 49 Mastra — ≈663 stars/month sampledest 50 Elicit — $22M, 2025-02est 52 Bedrock AgentCore — GA 2026-06-18
25est 60 AnythingLLM — >5M Docker pullsest 47 LobeHub — ≈663 stars/month sampledest 48 CrewAI — >$18M, 2024-10est 50 Meta Muse Code — beta 2026-08-05

The growth and new-entrant columns deserve special caution: the DeepSeek Harness, Ponytail, ECC, Graphify, Caveman, OpenClaw, and qm star patterns are large enough that stars should not be accepted as adoption without package use, identifiable users, repository provenance, and contribution-distribution checks.

Declining or dead — with dates

Product / pathStateEffective dateEvidenceUVU action
Flowisefact frozen, archived and ended2026-08-31Official EOL discussionDo not start; export existing flows
Roo Codefact repository archived/shutdown2026-05-15Archived repositoryDo not start; migrate maintained forks only after review
Continuefact no longer actively maintainedcurrent 2026-09-03Repository noticeDo not make a campus standard
OpenAI Agent Builderfact scheduled unavailable2026-11-30AgentKit noticeDo not build new durable workflows
OpenAI Swarmfact replaced by Agents SDK2025Swarm repositoryUse Agents SDK or another framework
AutoGenfact maintenance modecurrent 2026-09-03AutoGen repositoryNew Microsoft work goes to Agent Framework
Semantic Kernel agent pathfact new agent work moving to Agent Framework2026Microsoft Agent FrameworkKeep only supported existing workloads
Sourcegraph Cody Free/Pro/Enterprise Starterfact unavailable2025-07-23Cody FAQDo not plan around retired tiers
Amazon Q CLIfact superseded by Kiro CLIcurrent 2026Migration guideEvaluate Kiro separately; do not assume compatibility
Gemini CLI consumer service pathfact ended/transitioned to Antigravity CLI2026-06-18Google transition noticeFreeze before adopting either path campus-wide
SWE-agent full harnessfact directs new users to mini-SWE-agentcurrent 2026RepositoryUse for research replication, not a campus default
HuggingChat original servicefact shut down, later replaced by Omni2025-07-01 / 2025-10-16Hugging Face ChatTreat Omni as a new service review
Quivrfact latest release 2025-02-042025-02-04RepositoryAvoid new standardization
MetaGPTfact last release 2025-03; last push 2026-01-212026-01-21RepositoryResearch reference only
GPT4Allfact latest release 2025-02-252025-02-25RepositoryPrefer Ollama, llama.cpp, Jan or LM Studio
Ragasfact no release in last 90 days; last push 2026-02-242026-02-24RepositoryDo not rely on it as the sole evaluation layer
AgentOpsfact latest GitHub release 2025-082025-08RepositoryPrefer Langfuse, Phoenix, Promptfoo or Opik
OpenHarnessfact last push 2026-06-04, despite April launch2026-06-04GitHub snapshotTreat as stalled until maintenance resumes

Open-source status also needs care:

  • Open WebUI’s custom license restricts rebranding for deployments with 50 or more users. License explanation
  • Dify’s modified Apache license restricts certain multi-tenant and branding uses. Dify license
  • n8n moved to its Sustainable Use License. License announcement
  • Crush uses FSL-1.1-MIT, which is source-available rather than immediately permissive.
  • Phoenix’s server uses the Elastic License 2.0.
  • These are not necessarily blockers for internal university use, but they are procurement and redistribution constraints.

Stability and longevity — the top 40, scored on six axes

L, M, H mean low, medium, or high risk, not product quality. Axes: VD vendor/platform dependence · LIC license or future-use risk · RUN funding/runway · BUS maintainer concentration · GOV governance concentration · BRK breaking-change churn. A dated desk assessment, not a security audit.

#HarnessOverallVDLICRUNBUSGOVBRKDated basis
1ChatGPTMHHLLHMfact >1B WAU and $122B committed by 2026-03-31; closed service
2ClaudeMHHLLHMfact $65B Series H, 2026-05-28; closed Anthropic service
3GeminiMHHLLHMfact >1B MAU, 2026-08-11; Google-controlled
4Microsoft 365 CopilotMHHLLHMfact >30M paid seats, 2026-07-29; Microsoft-controlled
5GitHub CopilotMHHLLHMfact 50M users, 2026-07-29; service and policy controlled by Microsoft/GitHub
6OpenAI CodexMMLLLHHfact OSS CLI plus closed service; version 0.153 by 2026-09-03
7Claude CodeMHHLLHHfact fast releases through 2.1.259; Anthropic routes only
8CursorHHHLMHHfact joined SpaceX 2026-08-14; survey share declined from 18% to 12%
9Replit AgentHHHLMHHfact $400M round 2026-03-11; vertically integrated platform
10LovableHHHLMHHfact $400M Series C, 2026-08-12; fast-moving hosted platform
11Open WebUIHLHMHHHfact custom license; multiple advisories patched June–Aug 2026
12LibreChatMLLHMMMfact MIT and active; unknown institutional funding; R90 only 4
13OllamaMMLLMHHfact MIT, $88M raised 2026-07-09, R90 29
14llama.cppMLLLMMHfact MIT/community; R90 ≥100; HF organization since 2026-02-20
15ClineMLLMMMHfact Apache-2.0; $32M; npm incident and later advisories patched in 2026
16OpenCodeMLLHMHHfact MIT, R90 55; unknown funding and institutional governance
17OpenHandsMLMMMMHfact MIT core, $18.8M Series A, R90 33
18GooseMLLMLLMfact Apache-2.0 and foundation home; safety switches off by default
19DifyHLHMMHMfact modified license; $30M round 2026-03-10
20LangflowMLLLMMHfact MIT/active; August 2026 advisories patched
21n8nHLHLMHHfact source-available license, $240M raised, R90 ≥100
22LangChainMLLLMHHfact MIT, $125M Series B, R90 86
23LangGraphMLLLMHMfact MIT; parent well-funded; active evolution
24LlamaIndexMLLMMHMfact MIT; $27.5M total as of 2025-03
25CrewAIMLLMMHHfact MIT; >$18M; R90 37; telemetry on
26Microsoft Agent FrameworkMLLLLHMfact MIT, Microsoft-backed, 1.0 on 2026-04-03
27PydanticAIMLLMMMHfact MIT; parent funded; R90 56
28Google ADKMMLLLHMfact Apache-2.0, Google-backed, R90 20
29Vercel AI SDKMLULMHHfact 23.6M npm/week; R90 ≥100; license not frozen in evidence set
30MastraMLMMMHHfact core/EE split; $35M total; R90 22
31OnyxMLMMMMMfact MIT core/EE split; $10M seed; active university reference
32RAGFlowMLLHMHMfact Apache-2.0; 1,581 open issue/PR count; funding unknown
33ElicitMHHMMHMfact $22M Series A; closed research service
34GleanMHHLLHMfact $300M ARR on 2026-05-28; closed enterprise service
35Gemini NotebookMHHLLHMfact >30M users; Google-controlled/rebranded 2026-07-16
36LangfuseMLMLMHHfact acquired by ClickHouse 2026-01; R90 ≥100
37PhoenixHLHLMHHfact ELv2 server, auth off and analytics on by default; R90 97
38BraintrustMMHLMHMfact $80M Series B 2026-02-17; managed control plane
39PromptfooMLLLMHMfact MIT; OpenAI acquisition agreement 2026-03-09
40PaperclipHLLHHHHfact created 2026-03-02; 5,370 open issue/PR count; funding/governance unknown

Category maps

CategoryCurrent leaderFastest riserSafest university choiceWildcard
Chat/front doorest ChatGPT by weekly reach; Gemini close by monthly reachest Gemini/Claude by current user and business growthest LibreChat behind LiteLLM for provider control; ChatGPT Edu for managed SaaSest Duck.ai for low-risk, privacy-oriented public use
Codingest GitHub Copilot by users and installsest Codex and Claude Code by current use growth; OpenCode in OSSest Cline Enterprise or governed OpenCode through LiteLLMest Goose because it is provider-neutral and foundation-backed
Browser/computer useest ChatGPT agent by parent-product reachest Browser Harness and Cloudflare Computer by new-project velocityest Playwright MCP inside a locked-down managed runtimeest Cloudflare Computer
Frameworksest LangChain by downloads and ecosystemest Microsoft Agent Framework among governed entrants; Paperclip by raw starsest Haystack or Microsoft Agent Framework, depending language/cloudest Paperclip, after provenance and maintenance checks
RAG/researchest Glean for enterprise knowledge; Gemini Notebook for personal research reachest Onyx/Dify among deployable platforms; Elicit among research toolsest Onyx for campus search; Gemini Notebook for governed individual studyest Elicit Research Agent
Workflowest n8n by OSS reach and commercial growthest Gumloop by recent funding; Activepieces by package growthest n8n self-hosted with license review, or Activepiecesest Gumloop
Local runtime/desktopest Ollama by developer reachest llama.cpp by repository velocityest llama.cpp as the technical floor; Ollama as the managed developer layerest Jan for a telemetry-off desktop
Evaluation/observabilityest Langfuse by OSS/package reachest Opik by releases; Promptfoo by use/acquisitionest Promptfoo offline plus Phoenix behind authenticationest MLflow where UVU already operates ML infrastructure

Fit to this plan — pilot, watch, avoid

The market evidence does not justify replacing the architecture. The smaller, safer move is to keep the router as the boundary and test additional harnesses around it.

CategoryPilot (2–3)WatchAvoid nowWhy
CodingCline Enterprise; OpenCode through LiteLLM; Goose or OpenHandsAntigravity; IBM Bob; PaperclipRoo Code; Continue; Grok BuildThe pilots cover IDE, terminal and autonomous modes while retaining provider choice. The avoided set is retired or too new to govern.
Browser/computer usePlaywright MCP in a sandbox; ChatGPT agent Enterprise with narrow permissions; Browserbase with contract limitsCloudflare Computer; Browser HarnessLogged-in Browser Use deployments without telemetry and egress controlsBrowser actions create a larger data and authority boundary than ordinary chat.
FrameworksMicrosoft Agent Framework; LangGraph; HaystackPaperclip; Antigravity SDK; Bedrock AgentCoreNew AutoGen, Semantic Kernel agent work, OpenAI SwarmThe pilots cover .NET/Python, graphs and mature RAG while avoiding announced migration paths.
RAG/researchOnyx; Gemini Notebook under Workspace; ElicitRAGFlow; ConsensusFlowise; QuivrOnyx matches UVU’s current research tier. Gemini Notebook and Elicit cover individual source-grounded research.
Workflown8n self-hosted; Activepieces; Dify self-hostedGumloop; Relevance AIOpenAI Agent Builder; unreviewed acquired platformsThese pilots offer useful visual workflows while preserving deployment choice. License terms must be checked before shared-service use.
Localllama.cpp; Ollama; Jan or LM StudioLocalAIGPT4All and Text generation web UI as campus standardsllama.cpp remains the strongest Apple-silicon floor. Ollama improves developer experience. Jan offers a telemetry-off desktop path.
EvaluationPromptfoo offline; Phoenix with authentication enabled; Langfuse with telemetry disabled or contractually approvedBraintrust; OpikRagas, AgentOps and Portkey as the sole standardThis gives UVU regression testing, traces, red-team checks and a self-host route without tying all proof to one vendor.

Pilot gates that apply to every harness:

  1. Route provider calls through LiteLLM unless the pilot requires a documented exception.
  2. Use university identity, role separation, audit logs, budget caps and short retention.
  3. Test FERPA exposure, accessibility, source attribution, prompt-injection handling, export and deletion.
  4. Deny browser/file/shell/network permissions by default; enable only the minimum needed.
  5. Record model, harness, version, policy and evaluation set together. A model benchmark is not a harness benchmark.
  6. Require an exit path: configuration export, data export, provider substitution and rollback.
  7. Run pilots for 90 days before changing the campus standard.

Live GitHub snapshot — 86 repositories, pulled directly from the API

Our own pull on September 3, 2026, independent of the survey: stars, forks, creation and last-push dates, latest release, archive state. Four projects are archived and nine had no code pushed in over 90 days — including two the earlier survey had described as active. This is the check that keeps "popular" and "maintained" from being confused.

RepositoryCategoryStarsForksCreatedLast pushLatest releaseFlagLicense
anomalyco/opencodeCoding / agent203,44426,5412025-04-302026-09-032026-09-02MIT
n8n-io/n8nWorkflow builder203,22360,5352019-06-222026-09-032026-09-03NOASSERTION
ollama/ollamaLocal runner / serving180,04317,6682023-06-262026-09-032026-08-27MIT
langgenius/difyWorkflow builder154,32624,3972023-04-122026-09-032026-08-25NOASSERTION
langflow-ai/langflowWorkflow builder154,18410,0122023-02-082026-09-032026-09-01MIT
open-webui/open-webuiChat front door150,80822,0362023-10-062026-09-022026-08-31NOASSERTION
anthropics/claude-codeCoding / agent143,89723,0032025-02-222026-09-022026-09-02
ggml-org/llama.cppLocal runner / serving126,89722,7032023-03-102026-09-032026-08-25MIT
openai/codexCoding / agent121,17418,5672025-04-132026-09-032026-09-03Apache-2.0
browser-use/browser-useBrowser agent112,15612,3342024-10-312026-09-032026-08-16MIT
google-gemini/gemini-cliCoding / agent106,80014,5252025-04-172026-09-032026-09-01Apache-2.0
vllm-project/vllmLocal runner / serving90,88121,6532023-02-092026-09-032026-08-26Apache-2.0
modelcontextprotocol/serversOther90,04911,5482024-11-192026-09-032026-08-31NOASSERTION
infiniflow/ragflowResearch / RAG89,98210,6102023-12-122026-09-032026-08-28Apache-2.0
zed-industries/zedCoding / agent89,70010,4202021-02-202026-09-032026-09-02NOASSERTION
OpenHands/OpenHandsCoding / agent86,06911,2862024-03-132026-09-032026-08-27MIT
lobehub/lobehubChat front door82,19815,8572023-05-212026-09-032026-08-28NOASSERTION
daytonaio/daytonaAgent framework / infra71,8295,6502024-02-062026-07-242026-06-23
FoundationAgents/MetaGPTAgent framework / infra70,1978,9212023-06-302026-01-212024-04-22stale >90dMIT
openinterpreter/openinterpreterOther68,2305,8732023-07-142026-08-202026-08-20Apache-2.0
cline/clineCoding / agent67,4027,2802024-07-062026-09-032026-09-02Apache-2.0
Mintplex-Labs/anything-llmChat front door65,5657,2472023-06-042026-09-032026-08-27MIT
microsoft/autogenAgent framework / infra60,7919,1782023-08-182026-04-152025-09-30stale >90dCC-BY-4.0
crewAIInc/crewAIAgent framework / infra58,0458,3262023-10-272026-09-032026-08-27MIT
BerriAI/litellmObservability / gateway57,93411,1312023-07-272026-09-032026-09-02NOASSERTION
zylon-ai/private-gptResearch / RAG57,4897,6142023-05-022026-09-022026-06-18Apache-2.0
FlowiseAI/FlowiseWorkflow builder55,40424,9712023-03-312026-08-132026-07-29ARCHIVEDNOASSERTION
AntonOsika/gpt-engineerCoding / agent55,1117,2932023-04-292025-05-142024-06-06ARCHIVEDMIT
aaif-goose/gooseCoding / agent53,8796,1612024-08-232026-09-032026-08-27Apache-2.0
CherryHQ/cherry-studioChat front door51,4014,9092024-05-242026-09-032026-09-03AGPL-3.0
mudler/LocalAILocal runner / serving48,8454,4122023-03-182026-09-032026-08-20MIT
Aider-AI/aiderCoding / agent48,6984,9202023-05-092026-05-222025-08-09stale >90dApache-2.0
oobabooga/textgenLocal runner / serving47,6145,9832022-12-212026-08-172026-05-20AGPL-3.0
exo-explore/exoLocal runner / serving47,2263,4922024-06-242026-08-252026-04-23Apache-2.0
janhq/janChat front door44,3133,0072023-08-172026-09-032026-07-23NOASSERTION
danny-avila/LibreChatChat front door42,7708,8492023-02-122026-09-03MIT
agno-agi/agnoAgent framework / infra42,0315,8652022-05-042026-09-032026-09-01Apache-2.0
chatboxai/chatboxChat front door41,6334,2242023-03-062026-09-022026-09-02GPL-3.0
langchain-ai/langgraphAgent framework / infra40,9926,9162023-08-092026-09-032026-08-27MIT
The-Vibe-Company/quivrResearch / RAG39,4903,7322023-05-122026-08-312025-02-04NOASSERTION
CopilotKit/CopilotKitCoding / agent37,1814,6012023-06-192026-09-032026-09-01MIT
khoj-ai/khojResearch / RAG37,0362,4522021-08-162026-08-022026-03-26AGPL-3.0
continuedev/continueCoding / agent35,7395,3222023-05-242026-09-032026-06-19Apache-2.0
langfuse/langfuseObservability / gateway34,1553,6882023-05-182026-09-032026-09-03NOASSERTION
TabbyML/tabbyCoding / agent33,8601,7882023-03-162026-06-302026-01-25NOASSERTION
sgl-project/sglangLocal runner / serving33,7798,5132024-01-082026-09-032026-08-22Apache-2.0
Pythagora-io/gpt-pilotCoding / agent33,6803,4732023-08-162026-06-18NOASSERTION
mckaywrigley/chatbot-uiChat front door33,3429,4212023-03-112024-08-03stale >90dMIT
onyx-dot-app/onyxResearch / RAG31,9104,3992023-04-272026-09-032026-09-02NOASSERTION
openai/openai-agents-pythonAgent framework / infra29,1694,6582025-03-112026-09-022026-08-19MIT
huggingface/smolagentsAgent framework / infra29,1402,9182024-12-052026-08-252026-05-29Apache-2.0
voideditor/voidCoding / agent28,8142,6442024-09-112026-06-02ARCHIVEDApache-2.0
charmbracelet/crushCoding / agent27,8762,2182025-05-212026-09-032026-08-31NOASSERTION
mastra-ai/mastraAgent framework / infra27,6682,7252024-08-062026-09-032026-08-28NOASSERTION
Kilo-Org/kilocodeCoding / agent27,1563,1142025-03-102026-09-032026-09-02MIT
vercel/aiAgent framework / infra26,5645,0692023-05-232026-09-032026-09-02NOASSERTION
huggingface/open-r1Local runner / serving26,4492,4482025-01-242026-04-02stale >90dApache-2.0
Cinnamon/kotaemonResearch / RAG25,7292,1552024-03-252026-07-142026-05-31Apache-2.0
letta-ai/lettaAgent framework / infra24,6012,6122023-10-112026-08-232026-05-14Apache-2.0
browserbase/stagehandBrowser agent24,1351,6652024-03-242026-09-032026-08-28MIT
Skyvern-AI/skyvernBrowser agent22,9252,1532024-02-282026-09-032026-08-31AGPL-3.0
PromtEngineer/localGPTResearch / RAG22,2062,4622023-05-242026-08-26MIT
stackblitz-labs/bolt.diyCoding / agent19,83910,5612024-10-132026-02-072025-05-12stale >90dMIT
pydantic/pydantic-aiAgent framework / infra19,6992,6342024-06-212026-09-032026-09-03MIT
agent0ai/agent-zeroCoding / agent19,0773,7792024-06-102026-09-022026-08-27NOASSERTION
avante-corp/avante.nvimCoding / agent18,1478482024-08-142026-08-282026-08-27Apache-2.0
camel-ai/camelAgent framework / infra17,6722,0662023-03-172026-09-032026-03-22Apache-2.0
plandex-ai/plandexCoding / agent15,6191,1772023-10-242025-10-032025-07-16stale >90dMIT
nanobrowser/nanobrowserBrowser agent13,7201,4512024-12-312026-08-182025-11-22Apache-2.0
e2b-dev/E2BAgent framework / infra13,6641,0182023-03-042026-09-022026-09-02Apache-2.0
Portkey-AI/gatewayObservability / gateway12,8901,2822023-08-232026-05-252026-01-12stale >90dMIT
bentoml/OpenLLMLocal runner / serving12,5258372023-04-192026-08-312025-04-21Apache-2.0
simonw/llmChat front door12,4599722023-04-012026-09-022026-09-02Apache-2.0
LostRuins/koboldcppLocal runner / serving11,6037582023-03-162026-09-032026-08-29AGPL-3.0
Arize-ai/phoenixObservability / gateway11,3071,0942022-11-092026-09-032026-09-03NOASSERTION
github/copilot-cliCoding / agent11,1361,9102023-01-062026-09-022026-08-29NOASSERTION
sigoden/aichatChat front door10,4257432023-03-032026-02-232025-07-06stale >90dApache-2.0
microsoft/magentic-uiCoding / agent10,0831,0162025-05-052026-09-032026-05-21MIT
microsoft/vscode-copilot-chatCoding / agent9,9712,0112025-06-102026-05-202026-04-07ARCHIVEDMIT
ml-explore/mlx-lmLocal runner / serving6,8821,0182025-03-112026-09-032026-04-22MIT
olimorris/codecompanion.nvimCoding / agent6,8334492023-12-272026-08-312026-08-24Apache-2.0
AgentOps-AI/agentopsObservability / gateway5,8106182023-08-152026-06-252025-08-29MIT
lmstudio-ai/lmsLocal runner / serving5,2584502024-04-152026-09-01MIT
ag2ai/ag2Agent framework / infra4,9007142024-11-112026-09-022026-08-28Apache-2.0
Fannovel16/comfyui_controlnet_auxOther4,1713732023-08-172026-08-27Apache-2.0
OpenHands/OpenHands-CloudCoding / agent76442025-02-172026-09-032026-09-02NOASSERTION

The master table — all 171

Chat and assistant front doors — 16

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact ChatGPTfact consumer/enterprise assistantfact closed; 2022-11-30fact OpenAIfact >1B weekly users, 2026-08-31fact active continuous servicefact $122B committed, 2026-03-31; unknown teamest high vendor dependence; strong runwayunknown no comparable marketplace scorefact OpenAI-hosted; no localfact Enterprise/Edu identity, admin, retention and usage controlsfact plan-dependent; enterprise data excluded from training by defaultest everyone, governed tier
fact Claudefact consumer/enterprise assistantfact closed; 2023-03-14fact Anthropicfact Claude Code weekly use doubled since 2026-01; business subscriptions quadrupledfact activefact $30B Series G, 2026-02-12; $65B Series H, 2026-05-28est low runway risk; high vendor dependenceunknown no comparable ratingfact Anthropic/Bedrock/Vertex; no localfact Enterprise SSO, SCIM, audit and spend controlsfact commercial/API data policy differs from consumer planest everyone, governed tier
fact Geminifact consumer/Workspace assistantfact closed; Bard 2023-03-21; Gemini rename 2024-02-08fact Googlefact >1B monthly users, 2026-08-11fact activefact Alphabet-backed; unknown standalone funding/teamest strong longevity; high ecosystem dependenceunknown no comparable ratingfact Google models; no localfact Workspace identity, audit, DLP and admin controlsfact Workspace protections differ from consumer useest everyone, governed tier
fact Microsoft 365 Copilotfact enterprise assistantfact closed; 2023-11-01 GAfact Microsoftfact >30M paid seats, 2026-07-29fact activefact Microsoft-backedest strong longevity; Microsoft lock-inunknown no comparable ratingfact Microsoft-hosted; multiple partner models in some featuresfact Entra, Purview, audit, DLP and budget administrationfact commercial data protectionsest everyone where M365 is standard
fact Perplexityfact search/research assistantfact closed; 2022fact Perplexity AIunknown current active-user countfact activeunknown complete current funding/team verificationest medium vendor riskunknown no comparable ratingfact hosted multi-model; no localfact Enterprise SSO/SCIM, roles, credit and audit controlsfact enterprise no-training commitments; consumer terms differest researchers
fact Poefact multi-bot chat front doorfact closed; 2022-12fact Quoraunknown current active-user countfact activefact Quora-backed; unknown Poe-specific fundingest medium/high third-party-bot riskunknown no comparable ratingfact many hosted providers; no localunknown complete campus-grade control setfact third-party bot operators can have separate data practicesest not suitable for regulated use
fact Le Chatfact chat/enterprise assistantfact closed; 2024-02fact Mistral AIunknown active-user countfact active; Vibe coding brand changed 2026-05-28unknown current round total in this sweepest medium vendor riskunknown no comparable ratingfact Mistral hosted; enterprise/on-prem optionsfact enterprise identity and deployment optionsunknown consumer telemetry default not fully verifiedest everyone, controlled pilot
fact Meta AIfact consumer assistantfact closed; 2023-09-27fact Metafact 700M MAU reported 2025-03; staleunknown current standalone growthfact Meta-backedest strong runway; high account/data couplingunknown no comparable ratingfact Meta/Llama hosted; no local front doorunknown university control planefact consumer service tied to Meta privacy termsest not suitable as campus standard
fact Grokfact consumer/enterprise assistantfact closed; 2023-11fact xAIunknown standalone active usersfact activefact $20B Series E announced 2026-01-06; unknown teamest strong runway; high vendor/platform riskunknown no comparable ratingfact xAI-hosted; no localunknown full university control setunknown default telemetry and training treatment by planest not suitable pending controls review
fact DeepSeek Chatfact consumer assistantfact closed service; 2025-01fact DeepSeek/High-Flyerunknown active-user countfact activeunknown outside funding; privately backedest jurisdiction and data-governance riskunknown no comparable ratingfact DeepSeek-hosted; models can be run separately locallyunknown campus-grade identity/auditfact policy covers prompts, files and history collectionest not suitable for protected data
fact Qwen Chatfact consumer assistantfact closed service; 2023fact Alibaba Cloudunknown active-user countfact activefact Alibaba-backedest high jurisdiction/vendor dependenceunknown no comparable ratingfact Qwen hosted; separate open-weight local modelsunknown campus-grade controlsunknown default consumer telemetry not fully verifiedest not suitable for protected data
fact Kimifact consumer assistantfact closed; 2023-10fact Moonshot AIunknown active-user countfact activeunknown current cumulative capital in this sweepest high jurisdiction/vendor dependenceunknown no comparable ratingfact hosted; no supported campus-local front doorunknown campus-grade controlsunknown default data use not fully verifiedest not suitable for protected data
fact Z.aifact consumer/developer assistantfact closed service; unknown first datefact Zhipu AIunknown active-user countfact activeunknown current funding/teamest high jurisdiction/vendor dependenceunknown no comparable ratingfact hosted; some GLM weights available separatelyunknown campus controlsunknown data defaultest not suitable pending review
fact You.comfact search/assistant front doorfact closed; 2021fact You.comunknown current active usersfact activeunknown current capital/teamest medium vendor riskunknown no comparable ratingfact hosted multi-modelfact enterprise controls advertised; depth not fully testedunknown plan-specific data defaultsest researchers
fact HuggingChat / Omnifact open-model chat front doorfact service; original 2023-04; Omni 2025-10-16fact Hugging Faceunknown active usersfact original service ended 2025-07-01 and returned as Omnifact Hugging Face-backedest medium product-continuity riskunknown no comparable ratingfact multiple open models; separate self-host pathsunknown complete campus controlsunknown current Omni telemetry defaultest researchers
fact Duck.aifact privacy-oriented assistantfact closed front door; unknown first datefact DuckDuckGounknown active-user countfact activefact DuckDuckGo-backedest medium vendor riskunknown no comparable ratingfact hosted third-party models; no localunknown SSO/budget/auditfact chats are not stored or used for training by defaultest everyone for low-risk work

Coding agents, editors and review systems — 50

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact Claude Codefact CLI/cloud coding agentfact proprietary client; 2025fact Anthropicfact 143,875/23,001/≈14,668; npm 20,032,224/week; 39% JetBrains survey usefact latest 2.1.259, 2026-09-02; fact use up from 18% in 2026-01fact Anthropic $65B Series Hest rapid change; strong runwayfact Terminal-Bench scores are agent+model, not harness-onlyfact Anthropic/Bedrock/Vertex; no localfact enterprise identity, policy and audit optionsfact route/plan dependentest builders
fact OpenAI Codexfact CLI/cloud coding agentfact Apache-2.0 CLI; 2025fact OpenAIfact 121,101/18,554/≈15,011; npm 22,794,236/week; 1.6M weekly usersfact latest 0.153, 2026-09-03; >3× users since 2026-01fact OpenAI $122B committedest rapid breaking-change riskfact Terminal-Bench entry is agent+modelfact OpenAI-first; compatible/custom and local endpoints in CLIfact enterprise admin; CLI policy surfacesfact OpenTelemetry off and prompt logging false by defaultest builders
fact GitHub Copilotfact IDE/CLI/cloud agentfact closed; 2021 previewfact Microsoft/GitHubfact 50M users; VS Code 74,501,253 installs; 4.6/5 from 1,054 ratingsfact users up from 26M in 2025-10fact Microsoft-backedest low runway; high vendor dependencefact marketplace rating; benchmarks depend on model/modefact hosted multi-model; no campus-local inferencefact enterprise policy, audit, content exclusions and spendfact business data protections; telemetry remains service dependentest builders
fact Cursorfact AI editor/cloud agentfact closed; 2023fact Cursor; joined SpaceX 2026-08-14fact 12% JetBrains survey; hundreds of millions weekly requests reportedfact survey down from 18% in 2026-01fact $900M Series C, 2025-06-06; >$500M ARR thenest ownership transition and vendor concentration riskunknown comparable marketplace ratingfact hosted multi-model; no supported local control planefact enterprise SSO, privacy and admin featuresfact privacy mode available; default depends on planest builders
fact Windsurffact editor/coding agentfact closed; 2024fact Cognition since 2025-07-14fact vendor reported hundreds of thousands of DAU and ≈$82M ARR at acquisitionunknown current 90-day use growthfact acquired by Cognition; price not disclosedest integration/ownership riskunknown comparable ratingfact hosted multi-modelfact enterprise identity/adminunknown verified defaults after acquisitionest builders
fact Google Antigravityfact IDE/agent environmentfact closed preview; 2026-05-19fact Googlefact 6% JetBrains surveyfact new Q2 2026fact Google-backedest preview/breaking-change riskunknown comparable ratingfact Gemini/Google hostedunknown full production control setfact usage telemetry enabled in previewest builders, pilot only
fact Gemini CLIfact coding CLIfact Apache-2.0; 2025fact Googlefact 106,800/14,525/850fact R90 96 including nightly; latest 2026-09-02fact Google-backedfact consumer service path ended 2026-06-18; Antigravity transitionunknown harness-only ratingfact Gemini/Vertex; no first-class localfact Google Cloud/Workspace controls by routeunknown CLI telemetry default in current transitionest builders
fact Julesfact cloud coding agentfact closed; 2025fact Googleunknown active usersfact activefact Google-backedest medium vendor/preview riskunknown public ratingfact Gemini-hostedfact Google-account controls; unknown full campus auditunknown data defaultest builders
fact Devinfact cloud software agentfact closed; 2023-12fact Cognitionunknown public active usersfact activefact Cognition raised >$1B at $26B valuation, 2026-05-27est strong runway; proprietary lock-inunknown independent current harness scorefact hosted; no localfact enterprise controls advertisedunknown default task-data retentionest builders
fact Replit Agentfact browser IDE/application agentfact closed; 2024fact Replitfact >50M platform users; 85% of Fortune 500 representedunknown agent-active countfact $400M round, 2026-03-11, $9B valuationest strong runway; platform lock-inunknown independent harness ratingfact Replit-hosted models/toolsfact teams, roles and enterprise controlsunknown precise default telemetryest builders
fact Amazon Q Developerfact IDE/cloud coding assistantfact closed; 2023fact AWSunknown active usersfact Q CLI superseded by Kiro CLIfact Amazon-backedest product-line migration riskunknown current independent scorefact AWS-hostedfact IAM, organization and logging integrationfact service telemetry under AWS termsest builders
fact Juniefact JetBrains coding agentfact closed; 2025fact JetBrainsfact 9% JetBrains surveyfact activefact JetBrains-backedest medium IDE lock-inunknown separate marketplace scorefact hosted providers; no local standardfact JetBrains enterprise controls vary by IDEunknown data defaultest builders
fact Kirofact IDE/CLI coding agentfact closed; 2025fact AWSunknown active usersfact current successor path for Amazon Q CLIfact Amazon-backedest new/product-transition riskunknown ratingfact AWS-hostedfact AWS identity/policy integrationunknown detailed defaultest builders
fact Ampfact coding agentfact closed; 2025fact Sourcegraphunknown active usersfact activefact Sourcegraph-backedest medium vendor riskunknown ratingfact hosted modelsunknown complete control matrixunknown defaultest builders
fact Factory Droidsfact cloud coding agentsfact closed; unknown first datefact Factoryunknown active usersfact activefact $50M Series B, 2025-09est medium vendor riskunknown independent ratingfact hostedfact enterprise controls advertisedunknown defaultest builders
fact Augment Codefact IDE coding agentfact closed; 2024fact Augmentfact VS Code 776,022 installs; 3.5/5 from 340 ratingsfact activefact $227M Series B, 2024-04; $252M total thenest strong funding; proprietary lock-infact marketplace ratingfact hostedfact enterprise security/adminunknown default telemetryest builders
fact Warpfact agentic terminalfact closed client/service; 2020fact Warpunknown active usersfact activeunknown current capital in this sweepest medium vendor riskunknown comparable ratingfact hosted agent; shell remains localfact team/admin optionsfact terminal telemetry configurable; exact current default not recheckedest builders
fact TRAEfact AI editorfact closed; 2025fact ByteDanceunknown active usersfact activefact ByteDance-backedest jurisdiction/vendor riskunknown ratingfact hosted modelsunknown university controlsunknown default data useest not suitable pending review
fact IBM Bobfact enterprise coding agentfact closed; GA 2026-04-28fact IBMfact used by about 80,000 IBM employeesfact new Q2 2026fact IBM-backedest low runway; new-product riskunknown independent ratingfact hosted/enterprise IBM stackfact enterprise identity, governance and auditunknown exact telemetry defaultest builders
fact Meta Muse Codefact coding agentfact closed beta; 2026-08-05fact Metaunknown beta usersfact new Q3 2026fact Meta-backedest high preview riskunknown ratingfact Meta-hostedunknown production controlsunknown telemetryest builders, watch
fact Grok Buildfact coding agent/clientfact source published 2026-07-15; license unknown herefact xAIfact ≈26.4K stars; issues disabledfact periodic upstream sync; new Q3 2026fact xAI-backedest high governance and issue-transparency riskunknown ratingfact Grok-hostedunknown campus controlsfact egress reporting described; unknown full defaultest not suitable yet
fact OpenCodefact CLI/TUI coding agentfact MIT; unknown first datefact Anomalyfact 203,446/26,541/5,643; npm 1,848,183/week; 7% surveyfact R90 55; est ≈12K stars/month over 118 daysunknown funding/teamest rapid-change and maintainer-concentration riskunknown independent harness-only ratingfact provider-agnostic; local modelsunknown native enterprise control planeunknown current default telemetryest builders
fact Aiderfact terminal coding agentfact Apache-2.0; 2023fact independentfact 48,698/4,920/1,849fact no GitHub release since 2025-08-09; development/tags continueunknown funding/teamest medium maintainer/bus-factor riskfact public model benchmark suite; mostly model+prompt resultsfact provider-agnostic; local modelsunknown SSO/audit/budgetfact consent telemetryest builders
fact Clinefact IDE coding agentfact Apache-2.0; 2024fact Clinefact 67,402/7,280/1,194; VS Code 5,198,510; 4.0/5, 312 ratingsfact R90 ≥100; est ≈1.45K stars/monthfact $32M seed+Series A, 2025-07-31fact npm compromise 2026-02-17; patched; later advisories patchedfact marketplace rating; vendor benchmark not independentfact provider-agnostic; local modelsfact Enterprise SSO, RBAC, audit and cost controlsfact code remains local in extension; optional telemetry, default unknownest builders
fact Roo Codefact IDE coding agentfact Apache-2.0; 2024fact Roo Code Inc.fact 24,313/3,415/1,034; 1,978,032 installs; 4.5/5, 348fact archived 2026-05-15unknown funding/teamfact shutdown/archivedfact marketplace ratingfact provider-agnostic/local historicallyunknown continuing controlsunknown post-shutdown data stateest not suitable
fact Kilo Codefact IDE coding agentfact MIT; unknown first datefact Kilo; acquired by Anaconda 2026-07-15fact 27,156/3,114/552; 1,494,211 installs; 4.3/5, 205fact activefact acquisition; price unknownest medium integration riskfact marketplace ratingfact multi-provider/localunknown complete campus control matrixunknown defaultest builders
fact Continuefact IDE coding assistantfact Apache-2.0; 2023fact Continuefact 35,739/5,322/941; 4,061,691 installs; 3.5/5, 181fact repository says no longer actively maintainedunknown current funding/teamfact maintenance retreat/read-only directionfact marketplace ratingfact provider-agnostic/local historicallyunknown continuing enterprise supportunknown defaultest not suitable
fact OpenHandsfact autonomous coding platformfact MIT core; 2024fact All Hands AIfact 86,069/11,286/640; PyPI 623,690/monthfact R90 33; latest 1.16, 2026-08-27fact $18.8M Series A, 2025-11-18est active but fast-movingfact benchmark results are agent+modelfact multi-provider/localfact self-host and enterprise controlsunknown telemetry defaultest builders
fact Goosefact local coding/general agentfact Apache-2.0; 2024fact Block; LF AI & Data/AAIFfact 53,879/6,161/267fact activefact corporate/foundation backedest favorable governance; active-change riskunknown independent ratingfact provider-agnostic/localunknown native SSO/budget; self-host policy possiblefact telemetry off; sandbox and prompt-injection defenses off by defaultest builders
fact Qwen Codefact coding CLIfact Apache-2.0; 2025fact Alibaba/Qwenfact 27,611/2,983/1,264est ≈804 stars/month over sampled intervalfact Alibaba-backedest medium governance/jurisdiction riskunknown ratingfact Qwen/OpenAI-compatible endpoints; local possibleunknown enterprise controlsunknown telemetryest builders
fact Mistral Vibe CLIfact coding CLIfact Apache-2.0; unknown first datefact Mistral AIfact 4,913/684/285fact activefact Mistral-backedest medium provider dependenceunknown ratingfact Mistral-firstunknown native controlsunknown telemetryest builders
fact Crushfact terminal coding agentfact FSL-1.1-MIT; 2025fact Charmbraceletfact 27,877/2,218/687est ≈952 stars/month over 97 daysunknown capital/teamest source-available license riskunknown ratingfact multi-provider/localunknown enterprise controlsfact telemetry on by defaultest builders
fact SWE-agentfact research coding agentfact MIT; 2024fact Princeton/NLP research communityunknown current count not frozen in final setfact directs new users toward mini-SWE-agentfact academic/community backedest migration riskfact SWE-bench results are agent+model/configurationfact multi-provider/local possibleunknown enterprise controlsunknown telemetryest researchers
fact mini-SWE-agentfact compact research coding agentfact MIT; 2025fact SWE-agent teamunknown exact snapshot countfact active successor pathfact academic/community backedest medium research-project riskfact benchmark results remain agent+modelfact multi-provider/local possibleunknown enterprise controlsunknown telemetryest researchers
fact CodeRabbitfact code-review agentfact closed service; unknown first datefact CodeRabbitfact >2M reviews/week; 17K customersfact activefact $143M Series C, 2026-08-12; unknown lead in this sweepest strong funding; hosted lock-inunknown independent review-quality ratingfact hostedfact enterprise SSO/policy/audit advertisedunknown code-retention default by planest builders
fact Qodofact coding/review agentfact closed plus OSS PR-Agent; unknown first datefact Qodounknown active usersfact activefact $70M Series B, 2026-03-30; $120M totalest strong funding; vendor riskunknown independent ratingfact hosted multi-modelfact enterprise controlsunknown defaultest builders
fact PR-Agentfact pull-request agentfact AGPL-3.0; 2023fact Qodounknown exact final snapshotfact activefact parent raised $120M totalest copyleft and parent-product coupling riskunknown ratingfact multi-providerunknown native complete controlsunknown telemetryest builders
fact ECCfact coding-agent toolkitfact MIT; 2026-01-18unknown owner provenance not fully verifiedfact 246,773/37,187/136fact R90 3; all growth occurred in <8 monthsunknown capital/teamest very high provenance/star-quality riskunknown ratingunknown provider/local matrixunknown controlsunknown telemetryest not suitable yet
fact DeepSeek Harnessfact coding harnessfact MIT; 2026-08-13unknown verified organizational relationshipfact 210,708/24,655/0fact R90 10; est 291,543 stars/month annualized from 22 days; anomalyunknown capital/teamest extreme provenance/manipulation risk until validatedunknown ratingunknownunknownunknownest not suitable
fact Ponytailfact coding/agent harnessfact MIT; 2026-06-12unknown owner/backerfact 122,907/6,643/202est 44,539 stars/month from 84-day lifetimeunknownest extreme new-project riskunknownunknownunknownunknownest not suitable yet
fact Cavemanfact coding harnessunknown license; 2026-04-04unknown owner/backerfact 102,927/5,986/121fact pushed 2026-09-02unknownest high license/provenance riskunknownunknownunknownunknownest not suitable yet
fact Graphifyfact coding/context harnessfact Apache-2.0; 2026-04-03unknown owner/backerfact 114,218/11,101/1,243fact pushed 2026-08-30unknownest high new-project/provenance riskunknownunknownunknownunknownest not suitable yet
fact ruflofact agent/coding orchestrationfact MIT; unknown first datefact community projectfact 70,316/8,380/900fact R90 ≥100unknownest rapid-change and maintainer riskunknownfact multi-provider/local optionsunknownunknownest builders, watch
fact oh-my-openagentfact coding-agent configurationunknown license; 2025-12unknown owner/backerfact 68,648/5,639/910fact R90 68, betaunknownest high license/beta riskunknownunknownunknownunknownest not suitable yet
fact Prime Agentfact coding agentfact MIT; 2026-05-08unknown owner/backerfact 19,733/2,154/81fact active new projectunknownest high new-project riskunknownunknownunknownunknownest builders, watch
fact OpenHarnessfact coding harnessfact MIT; 2026-04-01unknown owner/backerfact 15,633/2,545/86fact one release 2026-05-07; last push 2026-06-04unknownest high staleness risk despite ageunknownunknownunknownunknownest not suitable
fact Kimi Codefact coding agentfact MIT; 2026-05-22fact Moonshot/Kimifact 7,238/1,161/1,298fact active new projectfact parent-backed; exact project capital unknownest high new-project/jurisdiction riskunknownfact Kimi-firstunknownunknownest builders, watch
fact MiMo Codefact coding agentfact MIT; 2026-06-10unknown owner/backerfact 12,939/1,334/975est ≈4,580 stars/month over lifetimeunknownest high new-project riskunknownunknownunknownunknownest builders, watch
fact fxfact coding harnessfact Apache-2.0; 2026-08-11unknown owner/backerfact 2,710/312/184fact new Q3 2026unknownest high early-stage riskunknownunknownunknownunknownest builders, watch
fact qmfact coding harnessfact MIT; 2026-07-29unknown owner/backerfact 14,523/1,764/382est +6,712 stars in sampled seven-day intervalunknownest very high spike/provenance riskunknownunknownunknownunknownest not suitable yet

General, browser and computer-use agents — 18

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact ChatGPT agentfact browser/computer/research agentfact closed; 2025-07-17fact OpenAIunknown agent-specific active usersfact absorbed Operator/deep-research pathsfact OpenAI-backedest strong runway; high action/vendor riskunknown independent operational scorefact OpenAI-hostedfact Enterprise/Edu policy and admin controlsfact plan-dependent retention/trainingest everyone, limited pilot
fact Manusfact cloud general agentfact closed; 2025-03fact acquired by Meta 2025-12-29fact vendor reports millions of users, 147T tokens and 80M virtual machinesfact activefact Meta-backed; acquisition price unknownest ownership and hosted-action riskunknown independent ratingfact hostedunknown complete university controlsunknown task-data defaultest not suitable pending review
fact Perplexity Cometfact agentic browserfact closed; 2025fact Perplexityunknown active usersfact activeunknown capital in this sweepest browser-data/vendor riskunknown ratingfact hostedfact enterprise MDM and policy controls advertisedunknown detailed browsing-data defaultest researchers, limited pilot
fact Gemini Agent / Project Marinerfact browser/computer agentfact closed research/preview; 2024fact Google DeepMindunknown active usersfact capabilities moving into Gemini Agentfact Google-backedest preview/product-transition riskunknown independent ratingfact Google-hostedfact managed access is off by default in relevant enterprise pathsunknown detailed data defaultest everyone, watch
fact Browser Usefact browser-agent framework/front endfact MIT; 2024fact Browser Usefact 112,156/12,334/402fact R90 9; latest 0.13.8, 2026-08-16; est ≈4.86K stars/monthfact $17M seed, 2025-03-22est active; telemetry and browser-action riskunknown independent scorefact multi-provider; local possibleunknown native campus control planefact telemetry true by default; reported collection can include task/action contextest not suitable without isolation
fact Browser Harnessfact browser-agent harnessunknown license here; 2026-04-17unknown owner/backerfact ≈17.3K/1.7Kfact latest 0.1.10, 2026-08-26; est ≈3.73K stars/monthunknownest high early-stage riskunknownunknownunknownfact telemetry opt-outest builders, watch
fact Stagehandfact browser automation SDKfact MIT; 2024fact Browserbasefact ≈24.1K/1.7K; npm 1,403,303/weekfact V4, 2026-08-10fact Browserbase $40M Series B, 2025-06est medium platform dependenceunknown independent ratingfact multi-model; Browserbase or local browserfact Browserbase enterprise controls by routeunknown SDK/platform telemetry defaultest builders
fact Browserbase Agentsfact managed browser-agent platformfact closed; GA 2026-06-30fact Browserbasefact platform reports 35M browser sessions/month; not agent usersfact new Q2 2026fact $40M Series B, 2025-06est medium hosted-browser dependenceunknown ratingfact hosted browser, multi-modelfact enterprise session/admin controlsunknown exact default retentionest builders
fact Skyvernfact browser automation agentfact AGPL-3.0; 2023fact Skyvernfact ≈22.9K starsfact activeunknown current capital/teamest copyleft and browser-action riskunknownfact multi-model/self-hostunknown full native campus controlsfact telemetry on by defaultest builders
fact Playwright MCPfact browser MCP serverfact Apache-2.0; 2025fact Microsoftfact ≈36.8K starsfact activefact Microsoft-backedest low project-runway risk; browser-action risk remainsunknownfact model-agnostic; local browserunknown native SSO/budget; host supplies policyest no model telemetry intrinsic; browser data reaches chosen host/modelest builders
fact Open Interpreterfact local computer/code agentfact AGPL-3.0; 2023fact Open Interpreterfact ≈68.2K starsfact product direction rewritten around Rust/Codex; old Python community persistsunknownest high architecture/migration riskunknownfact multi-provider/local historicallyunknownunknownest builders, watch
fact UI-TARS / Agent TARSfact computer-use agentfact Apache-2.0; 2025fact ByteDancefact ≈38.8K starsfact activefact ByteDance-backedest jurisdiction and action-safety riskunknown benchmark results depend on model/environmentfact TARS models/local optionsunknown campus controlsunknown telemetryest researchers
fact Agent Sfact computer-use research agentfact Apache-2.0; 2024fact Simular/academic collaboratorsfact ≈12.2K starsfact activeunknown capital/teamest research-system riskfact OSWorld results are model+agent+environmentfact multi-modelunknown enterprise controlsunknownest researchers
fact OpenClawfact general autonomous agentunknown license; 2025-11-24unknown owner/backerfact 388,724/81,627/6,089fact activeunknownest extreme provenance, governance and open-issue riskunknownunknownunknownunknownest not suitable
fact AutoGPTfact general-agent platformfact MIT; 2023fact Significant Gravitasfact 187,099/46,040/544fact R90 13unknownest medium product-direction riskunknown current independent ratingfact multi-provider/local possibleunknown complete controlsunknown telemetryest researchers
fact nanobotfact compact agent frameworkfact MIT; 2026-02-01fact HKU Data Science groupfact 47,685/8,416/758fact active new projectfact academic backingest high early-stage/bus-factor riskunknownfact multi-provider/local possibleunknownunknownest researchers
fact Cloudflare Computerfact computer-use runtimefact MIT; 2026-06-05fact Cloudflarefact 8,954/501/21fact +8,868 stars in monthly trending windowfact Cloudflare-backedest preview and platform-dependence riskunknownfact Cloudflare runtime; model choice variesfact Cloudflare account controlsunknown telemetry/retentionest builders, watch
fact Cloudflare OSfact agent operating environmentfact Apache-2.0; 2026-04-15fact Cloudflarefact 9,539/1,123/117fact +9,549 monthly-trending signalfact Cloudflare-backedest high preview/breaking-change riskunknownfact Cloudflare platformfact account controlsunknown telemetryest builders, watch

Agent frameworks, SDKs and orchestration — 23

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact LangChainfact agent/LLM frameworkfact MIT; 2022fact LangChain Inc.fact 145,576/24,297/443; PyPI 229,927,751/monthfact R90 86; latest alpha 2026-09-02fact $125M Series B, 2025-10-20, $1.25B valuationest low runway; high change/vendor-direction riskfact benchmark claims configuration-dependentfact provider-agnostic/localfact LangSmith adds SSO/audit/spend; core does notfact core tracing off; CLI analytics onest builders
fact LangGraphfact stateful agent frameworkfact MIT; 2023fact LangChain Inc.fact 40,992/6,916/743; npm 13,152,611 in Augfact activefact parent $125M Series Best medium vendor/change riskfact deep-agent results are agent+model/configurationfact provider-agnostic/localfact managed platform adds enterprise controlsunknown core telemetry defaultest builders
fact AutoGenfact multi-agent frameworkfact MIT; 2023fact Microsoftfact 60,791/9,178/1,031fact maintenance modefact Microsoft-backedfact successor is Microsoft Agent Frameworkunknown current ratingfact provider-agnostic/localunknown core native controlsunknown telemetryest not suitable for new builds
fact Semantic Kernelfact agent/application SDKfact MIT; 2023fact Microsoftfact 28,527/4,752/271fact new work directed toward Microsoft Agent Frameworkfact Microsoft-backedest migration riskunknownfact multi-provider/local connectorsfact Azure supplies enterprise controlsunknown SDK telemetryest not suitable for new agent builds
fact Microsoft Agent Frameworkfact successor agent SDKfact MIT; 2026fact Microsoftfact 13,306/2,252/649fact 1.0 on 2026-04-03; R90 24; Python 1.17 on 2026-09-03fact Microsoft-backedest medium early-version riskunknown independent scorefact multi-provider; Azure/local connectorsfact Azure identity, audit and policy integrationunknown core default telemetryest builders
fact CrewAIfact multi-agent framework/platformfact MIT; 2023fact CrewAIfact 58,045/8,326/710; PyPI 26,457,068/month; vendor says 10M agents/monthfact R90 37 GitHub releasesfact >$18M raised by 2024-10est active; medium vendor/change riskunknown independent ratingfact provider-agnostic/localfact enterprise platform controls; core limitedfact telemetry on by defaultest builders
fact LlamaIndexfact data/agent frameworkfact MIT; 2022fact LlamaIndexfact 51,998/8,076/696; PyPI 6,405,226/monthfact activefact $19M Series A, 2025-03; $27.5M totalest medium vendor/change riskunknown comparable ratingfact provider-agnostic/localfact managed platform controls; core limitedunknown core telemetry defaultest researchers
fact Haystackfact RAG/agent frameworkfact Apache-2.0; 2019fact deepsetfact 26,403/3,071/103; PyPI 760,882/monthfact activeunknown current parent capital/teamest comparatively maturefact ADK Arena preprint reports strong build-task performance; not production prooffact provider-agnostic/localfact deepset platform provides enterprise controlsunknown core telemetryest researchers
fact PydanticAIfact typed agent frameworkfact MIT; 2024fact Pydanticfact 19,699/2,634/780; PyPI 11,040,846/monthfact R90 56fact parent $12.5M Series A, 2024-10est active; rapid-change riskunknownfact provider-agnostic/localunknown core enterprise controlsunknown telemetryest builders
fact OpenAI Agents SDKfact agent SDKfact MIT; 2025fact OpenAIfact 29,169/4,658/78; PyPI 33,424,617/monthfact R90 17fact OpenAI-backedest low runway; vendor-direction riskunknown harness-only scorefact OpenAI-first; custom model adaptersfact OpenAI org controls by routefact tracing on by default and can include prompts/resultsest builders
fact Google ADKfact agent SDKfact Apache-2.0; 2025fact Googlefact 21,390/3,937/517; PyPI 19,429,084/monthfact R90 20fact Google-backedest medium vendor/change riskfact ADK Arena results are framework+configuration, preprintfact multi-model; Google-optimized; local possiblefact Google Cloud controls by deploymentunknown SDK telemetryest builders
fact Strands Agentsfact agent SDKfact Apache-2.0; 2025fact AWS; foundation participationfact 7,141/1,089/699; vendor says 14M downloadsfact activefact AWS-backedest medium ecosystem dependenceunknown independent scorefact model-agnostic; AWS-optimizedfact IAM/CloudTrail by deploymentunknown SDK telemetryest builders
fact Mastrafact TypeScript agent frameworkfact Apache-2.0 core plus commercial EE; 2024fact Mastrafact 27,668/2,725/541; npm ≈1.58M/weekfact R90 22fact $22M Series A, 2026-04-09; $35M total; >35 staffest strong growth; license/product-boundary riskunknownfact provider-agnostic/localfact enterprise tier controlsunknown telemetryest builders
fact Agnofact agent framework/runtimefact Apache-2.0; 2023fact Agnofact 42,031/5,865/1,312; est ≈2.18M monthly downloads from secondary counterfact 31 stable/47 total releases in 90 daysunknown capital/teamest high release/breaking-change riskunknownfact provider-agnostic/localunknown complete enterprise controlsfact telemetry onest builders
fact smolagentsfact compact agent frameworkfact Apache-2.0; 2024fact Hugging Facefact 29,140/2,918/756; PyPI 580,640/monthfact no GitHub releases in R90; source activefact Hugging Face-backedest medium cadence riskunknownfact provider-agnostic/localunknown native controlsunknown telemetryest researchers
fact DSPyfact program/agent optimization frameworkfact MIT; 2023fact Stanford research/communityunknown exact final snapshot omittedfact activefact academic/community-backedest medium research-governance riskfact evaluations are task/model/program dependentfact provider-agnostic/localunknown native controlsunknown telemetryest researchers
fact BeeAIfact agent frameworkfact Apache-2.0; 2024fact LF AI & Data; originated at IBMfact ≈3,390 starsfact IBM no longer sole maintainer; foundation continuesfact foundation/corporate backingest medium governance-transition riskunknownfact provider-agnostic/localunknownunknownest researchers
fact MetaGPTfact multi-agent frameworkfact MIT; 2023fact FoundationAgents/communityfact 70,197 starsfact last push 2026-01-21; latest release 2025-03unknown capital/teamest high staleness riskunknownfact multi-providerunknownunknownest not suitable for new builds
fact Lettafact stateful/memory agent platformfact Apache-2.0 core; 2023fact Lettafact 24,601 starsfact latest release 2026-05-14unknown current capital/teamest medium cadence/vendor riskunknownfact multi-provider/localfact hosted controls vary by tierunknown telemetryest researchers
fact Vercel AI SDKfact application/agent SDKunknown license assertion not frozen; 2023fact Vercelfact 26,564/5,069/1,545; npm 23,560,904/weekfact R90 ≥100fact Vercel-backedest high change rate; ecosystem dependenceunknown harness ratingfact provider-agnostic; local endpoints possiblefact Vercel platform controls; core limitedunknown core telemetryest builders
fact Antigravity SDKfact agent SDKunknown preview license; 2026-04-29fact Googlefact 3,275/1,295/32fact no formal releases; previewfact Google-backedest high preview/API-change riskunknownfact Google-optimizedunknown production controlsunknown telemetryest builders, watch
fact Paperclipfact agent orchestration/workspacefact MIT; 2026-03-02unknown owner/backer verified only at repository levelfact 79,934/14,675/5,370fact R90 11; est ≈4,962 stars/month from sampled baselineunknown capital/teamest very high newness, issue-load and bus-factor riskunknownfact multi-provider claims; local unknownunknownunknownest builders, watch
fact Amazon Bedrock AgentCorefact managed agent runtime/harnessfact closed service; preview 2026-04; GA 2026-06-18fact AWSunknown active customersfact new Q2 2026fact Amazon-backedest low runway; high AWS dependenceunknownfact model-flexible inside AWSfact IAM, CloudTrail, budgets and policyfact CloudWatch tracing on in configured paths; memory default 30 daysest builders

RAG, research and knowledge systems — 16

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact Gemini Notebookfact source-grounded research notebookfact closed; NotebookLM 2023; renamed 2026-07-16fact Googlefact >30M people and 600K organizationsfact active/rebrandedfact Google-backedest strong longevity; Google lock-inunknown comparable research ratingfact Gemini-hosted; no localfact Workspace identity/admin by editionfact enterprise protections differ from consumerest researchers
fact Elicitfact literature-research assistantfact closed; 2017fact Ought/Elicitfact >400K monthly researchersfact Research Agent launched 2026-08-04fact $22M Series A, 2025-02est medium vendor riskunknown independent comprehensive ratingfact hosted modelsfact organization controls; complete campus matrix unknownunknown exact default by tierest researchers
fact Consensusfact scholarly search/answer systemfact closed; 2022fact Consensusfact 2.5M MAUfact activefact $30M funding announced 2026-05-11est medium vendor riskunknown independent ratingfact hostedunknown complete SSO/audit setunknown defaultest researchers
fact Gleanfact enterprise knowledge/search agentfact closed; 2019fact Gleanfact $300M ARR, 2026-05-28; nearly 2× Fortune 500 reach; 45% wDAU/wMAUfact ARR tripled from $100M in 15 monthsfact $150M Series F, 2025-06, $7.2B valuationest strong runway; high vendor dependenceunknown comparable ratingfact hosted multi-modelfact enterprise identity, ACL preservation, audit and governanceunknown exact telemetry/retention by contractest everyone where budget permits
fact Hebbiafact enterprise document researchfact closed; unknown first datefact Hebbiaunknown active seatsfact activefact $130M Series B, 2024-07est strong capital; proprietary lock-inunknown independent ratingfact hostedfact enterprise controls advertisedunknown defaultest researchers
fact SciSpacefact scholarly reading/search assistantfact closed; unknown first datefact SciSpacefact vendor reports 9.6M researchers; independent confirmation unknownunknown 90-day growthunknown current capital/teamest medium vendor/evidence riskunknown independent ratingfact hostedunknown complete campus controlsunknown data defaultest researchers
fact Onyxfact enterprise search/RAG assistantfact MIT core plus proprietary EE; 2023fact Onyxfact 31,910/4,399/427; vendor references 14K Netflix and 37K UCSD usersfact activefact $10M seed, 2025; unknown teamest medium commercial-boundary riskunknown comparable ratingfact provider-agnostic; local/air-gappedfact SSO, RBAC, audit and enterprise controlsunknown telemetry default; self-host permits stronger containmentest everyone
fact Khojfact personal/team knowledge assistantfact AGPL-3.0; 2023fact Khojfact 37,033/2,451fact activeunknown capital/teamest copyleft and small-team riskunknownfact multi-provider/localunknown complete campus controlsunknown telemetryest researchers
fact RAGFlowfact document RAG platformfact Apache-2.0; 2023fact InfiniFlowfact 89,983/10,610/1,581est ≈2.52K stars/month in sampled 11-day intervalunknown capital/teamest medium issue-load/vendor riskunknownfact multi-provider/localfact self-host; enterprise controls varyunknown telemetryest researchers
fact Difyfact RAG/agent application builderfact modified Apache; 2023-04-12fact Difyfact 154,326/24,397/1,021; vendor claims 1.4M machinesfact R90 5; latest 1.17.0, 2026-08-25; est ≈16.8K stars/month over sampled 97 daysfact $30M pre-Series A, 2026-03-10fact license restricts some multi-tenant/rebranding usesunknown independent ratingfact provider-agnostic/localfact enterprise SSO/audit/budget featuresfact telemetry present/on in default feature configurationest builders
fact Flowisefact visual RAG/agent builderfact Apache-2.0; 2023fact Flowisefact 55,404/24,971/1,046fact R90 2; frozen 2026-07-29; archived 2026-08-13; EOL 2026-08-31unknownfact endedunknownfact historically multi-provider/localunknown continuing controlsunknownest not suitable
fact Langflowfact visual RAG/agent builderfact MIT; 2023fact Langflow/Astra ecosystemfact 154,184/10,012/1,021fact R90 12; latest 1.12, 2026-09-01unknown current capital/teamfact security advisories in 2026-08; patches releasedunknownfact provider-agnostic/localfact enterprise deployment optionsunknown telemetry defaultest builders
fact Quivrfact personal/team RAGfact Apache-2.0; 2023fact Quivrfact 39,379 starsfact latest release 2025-02-04unknownest high staleness riskunknownfact multi-provider/local historicallyunknownunknownest not suitable for new standard
fact DocsGPTfact document chat/RAGfact MIT; 2023fact Arc53/communityfact 18,236 starsfact activeunknownest medium small-project riskunknownfact multi-provider/localunknown enterprise control depthunknown telemetryest researchers
fact Kotaemonfact document RAG UIfact Apache-2.0; 2024fact Cinnamon/communityfact 25,707 starsfact activeunknownest medium maintainer riskunknownfact multi-provider/localunknownunknownest researchers
fact PrivateGPTfact local document RAGfact Apache-2.0; 2023fact Zylon/communityfact 57,487 starsunknown current release cadence not frozenunknownest medium cadence/maintainer riskunknownfact local-first; multi-providerunknown enterprise controlsest local operation can avoid external telemetryest researchers

Workflow and automation builders — 15

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact n8nfact workflow/agent builderfact Sustainable Use License; 2019fact n8nfact 203,223/60,535/1,137; npm 94,686/weekfact R90 ≥100fact $180M Series C, 2025-10-09; $240M total; 6× users/10× revenue reportedfact 2026 security advisories patched; license limits commercial hostingunknown comparable ratingfact many providers/local connectorsfact SSO, RBAC, audit and external secrets in enterprisefact telemetry on unless disabledest builders
fact Activepiecesfact workflow/agent builderfact MIT core plus EE; 2022fact Activepiecesfact 24,210/4,136/519; npm 2,305,328 in Augfact R90 29unknown current capital/teamest medium commercial-boundary riskunknownfact multi-provider/self-hostfact enterprise SSO/RBAC/auditunknown telemetry defaultest builders
fact Pipedreamfact developer automation platformfact closed platform; integration repo OSS; 2019fact Workday acquisition completed by 2026-01-31fact integrations repo 11,668 starsfact activefact acquired; price unknownest integration/ownership riskunknownfact hosted multi-providerfact enterprise identity/auditunknown default telemetryest builders
fact Gumloopfact no-code agent workflow builderfact closed; unknown first datefact Gumloopunknown active usersfact activefact $50M Series B, 2026-03-12; at least $70M totalest strong funding; proprietary lock-inunknownfact hosted multi-modelfact enterprise controls advertisedunknown defaultest builders
fact Relevance AIfact agent/workforce builderfact closed; unknown first datefact Relevance AIunknown active usersfact activefact $24M Series B, 2025-05est medium vendor riskunknownfact hosted multi-modelfact enterprise controls advertisedunknownest builders
fact Lindyfact no-code assistant/agent builderfact closed; unknown first datefact Lindyunknown active usersfact activefact ≈23 staff publicly indicated; funding unknown hereest small-team/vendor riskunknownfact hostedfact SOC 2 and Safe Mode advertised; full audit matrix unknownunknown defaultest builders
fact StackAIfact enterprise workflow/agent builderfact closed; unknown first datefact acquired by Asana 2026-05-28unknown active usersfact active/acquisition integrationfact ≈$75M upfront reported; ≈62 staffest ownership/integration riskunknownfact hosted multi-modelfact enterprise controls advertisedunknownest builders
fact OpenAI Agent Builderfact hosted visual agent builderfact closed; 2025fact OpenAIunknown active buildersfact service unavailable after 2026-11-30fact OpenAI-backedfact sunset announcedunknownfact OpenAI-hostedfact OpenAI organization controlsfact tracing/data follows platform settingsest not suitable
fact Zapier Agentsfact SaaS agent builderfact closed; 2024fact Zapierfact 50K teams reported before rename; >9K app integrationsfact activefact Zapier-backedest medium vendor dependenceunknownfact hosted multi-modelfact enterprise identity, app policy and auditunknown detailed defaultest builders
fact Make AI Agentsfact SaaS workflow/agent builderfact closed; unknown first datefact Make/Celonisunknown active usersfact activefact Celonis-backedest medium platform dependenceunknownfact hostedfact enterprise controls by planunknownest builders
fact Bardeenfact browser/workflow automationfact closed; 2021fact Bardeenunknown active usersfact activeunknown current capital/teamest browser-data/vendor riskunknownfact hostedfact team controls advertisedunknownest builders
fact Relay.appfact human-in-loop workflow builderfact closed; unknown first datefact Relayunknown active usersfact activeunknown capital/teamest small-vendor riskunknownfact hosted multi-modelunknown complete campus controlsunknownest builders
fact Dustfact enterprise assistant/agent builderfact closed plus OSS components; 2023fact Dustunknown active usersfact activeunknown current capital/teamest medium vendor riskunknownfact multi-modelfact SSO, permissions and enterprise administrationunknownest everyone, controlled pilot
fact Coze Studiofact agent workflow builderfact Apache-2.0; 2025fact ByteDance/Cozeunknown exact final snapshotfact activefact ByteDance-backedest jurisdiction and platform riskunknownfact multi-model/self-host claimsunknown complete controlsunknown telemetryest builders
fact Workato Agenticfact enterprise automation/agent platformfact closedfact Workatounknown agent-specific active usersfact activeunknown current capital in this sweepest strong enterprise position; proprietary lock-inunknownfact hostedfact mature identity, governance and audit controlsunknownest builders

Local runners and desktop front doors — 18

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact Ollamafact local model runner/APIfact MIT; 2023fact Ollamafact 180,043/17,668/3,889; vendor says 8.9M developersfact R90 29; latest RC 2026-09-02fact $88M raised 2026-07-09est strong funding; company dependenceunknown harness ratingfact local-first, including Apple siliconunknown native SSO/audit; external gateway neededfact telemetry disabled by default; cloud path offers stated zero-data-retentionest everyone
fact LM Studiofact local desktop runner/chatfact closed/free; 2023fact LM Studiofact vendor reports millions of downloadsfact 0.4.17 on 2026-06-27; Bionic update 2026-07-16unknown capital/teamest medium closed-client riskunknown public comparable ratingfact local-first, Apple siliconunknown centralized campus controlsfact local privacy claims; exact analytics default unknownest everyone
fact GPT4Allfact local desktop/chat runnerfact MIT; 2023fact Nomic AIfact ≈77.4K/8.3Kfact latest release 2025-02-25unknown current capital/teamest high staleness riskunknownfact local-firstunknown enterprise controlsunknown telemetryest not suitable as new standard
fact Janfact local desktop assistantfact Apache-2.0; 2023fact Menlo Researchfact 44,313/3,007/509; vendor says 6M downloadsfact R90 2unknown capital/teamest medium cadence riskunknownfact local-first and hosted providersunknown centralized SSO/auditfact analytics opt-in/off by defaultest everyone
fact LocalAIfact local OpenAI-compatible runtimefact MIT; 2023fact communityfact 48,845/4,412fact activeunknown capital/teamest medium maintainer/bus-factor riskunknownfact 60+ local backendsfact RBAC/quota capabilities; full campus audit unknownunknown telemetryest builders
fact llama.cppfact local inference runtimefact MIT; 2023fact ggml community; Hugging Face organization since 2026-02-20fact 126,898/22,705/2,389fact R90 ≥100; est ≈4.26K stars/month in sampled 29 daysfact community/corporate ecosystem; no single funding roundest low license/runway risk; high change rateunknown harness ratingfact local-first; strong Apple silicon supportunknown native SSO/budget/auditfact no hosted telemetry intrinsicest everyone
fact llamafilefact portable local model runnerfact Apache-2.0; 2023fact Mozilla Ocho/communityfact ≈25.9K starsfact 0.10.5 on 2026-08-03fact Mozilla-backed projectest medium project-scope riskunknownfact local-firstunknown native controlsfact no hosted telemetry intrinsicest builders
fact Open WebUIfact self-hosted chat/RAG front doorfact custom license; 2023-10-06fact Open WebUIfact 150,807/22,036/243fact R90 7; latest 0.11.3, 2026-08-31unknown funding/teamfact 50+ user branding restriction; several 2026 advisories patchedunknownfact provider-agnostic/localfact SSO/roles/admin analytics; deeper controls varyfact external tracing off; admin analytics onest everyone, governed self-host
fact LibreChatfact self-hosted multi-model chatfact MIT; 2023fact community/project companyfact 42,771/8,849/732fact R90 4; latest release candidateunknown capital/teamest medium maintainer/release-cadence riskunknownfact provider-agnostic/localfact OAuth, LDAP, SAML, roles, token-spend and audit functionsunknown optional providers determine data pathest everyone
fact LobeHubfact chat/agent front endfact custom license; 2023fact LobeHubfact 82,183 stars; vendor says 6M usersest ≈663 stars/month in sampled 11-day intervalunknown capital/teamest license and vendor-claim riskunknownfact multi-provider/localfact enterprise options; exact campus matrix unknownunknown default telemetryest everyone, review license
fact AnythingLLMfact desktop/self-host RAG chatfact MIT; 2023fact Mintplex Labsfact 65,565/7,247/326; >5M Docker pullsfact R90 6unknown capital/teamest medium company dependenceunknownfact multi-provider/localfact enterprise SSO, RBAC and logsfact desktop states no data phones homeest everyone
fact Mstyfact local/multi-model desktopfact closed; unknown first datefact Mstyunknown active usersfact activefact ≈8 staff publicly indicatedest small-team/closed-client riskunknownfact local and hostedunknown complete campus controls; certifications incompletefact zero product telemetry statedest everyone, limited pilot
fact Cherry Studiofact desktop multi-model clientfact AGPL-3.0; 2024fact Cherry Studio/communityfact 51,397 starsfact activeunknownest medium governance/copyright review riskunknownfact multi-provider/localunknown centralized controlsunknown telemetryest everyone
fact PeerLLMfact local desktop assistantfact proprietary; local mode 2026-07-19unknown owner/backerunknown public scalefact new Q3 2026unknownest very high early-stage riskunknownfact local mode claimedunknownunknownest not suitable yet
fact Mantle Chatfact local desktop assistantfact closed alpha; 2026-08-06unknown owner/backerunknown public scalefact new Q3 2026fact ≈2-person teamest very high bus-factor/alpha riskunknownfact local claimsunknownunknownest not suitable yet
fact Chatboxfact desktop multi-model chatfact GPL; 2023fact community/companyfact 41,633/4,224/1,261fact R90 7unknownest medium issue-load riskunknownfact multi-provider/localunknown campus controlsunknown telemetryest everyone
fact Text generation web UIfact local model web UIfact AGPL-3.0; 2022fact communityfact ≈47.2K starsfact latest release 2026-05-20unknownest medium cadence/maintainer riskunknownfact local-firstunknown enterprise controlsunknown telemetryest researchers
fact KoboldCppfact local model runner/UIfact AGPL-3.0; 2023fact communityfact ≈11.4K starsfact release 2026-08-16unknownest medium bus-factor riskunknownfact local-firstunknown enterprise controlsunknown telemetryest researchers

Evaluation, observability and harness-adjacent systems — 15

HarnessFormLicense · first releaseOwner / backerPublic scaleGrowth / releaseCapital / teamStabilityRating / benchmarkProvider / localSSO · budget · auditData / telemetry defaultUVU verdict
fact Langfusefact tracing/evaluation platformfact MIT core plus EE; 2022fact ClickHouse since 2026-01fact 34,155/3,688/881; npm 7,933,285 in Aug; PyPI 25,881,197/monthfact R90 ≥100fact acquired by ClickHouse; ≈16+ staff before/around acquisitionest strong parent; integration and high-change riskunknown independent ratingfact provider-agnostic/self-hostfact SSO/RBAC/audit in enterprisefact product telemetry on; some enterprise self-host telemetry cannot be disabledest builders
fact Phoenixfact tracing/evaluation UIfact ELv2 server; some clients Apache-2.0; 2023fact Arize AIfact 11,307/1,094/969fact R90 97fact Arize $70M Series C, 2025-02est license and rapid-change riskunknown independent ratingfact provider-agnostic/self-hostfact authentication available but off by defaultfact analytics on by defaultest builders
fact Heliconefact LLM gateway/observabilityfact Apache-2.0; 2023fact Mintlify since 2026-03-03fact 6,131/662/156fact source active; latest GitHub release 2025-08fact acquired; price unknownest acquisition/release-process riskunknownfact provider-agnosticfact hosted enterprise controlsunknown telemetry beyond observed requestsest builders
fact Braintrustfact evaluation/observability platformfact closed platform; SDK components OSSfact Braintrustunknown public active-user countfact activefact $80M Series B, 2026-02-17; at least $121M totalest strong capital; proprietary control-plane riskunknown independent ratingfact provider-agnosticfact enterprise SSO/audit/access controlsfact managed system/billing telemetry remains onest builders
fact W&B Weavefact evaluation/tracing platformfact Apache-2.0 SDK; hosted platform closedfact Weights & Biasesfact ≈1,125 starsfact activefact W&B-backedest medium platform dependenceunknownfact provider-agnosticfact enterprise identity/auditunknown defaultest builders
fact Opikfact evaluation/observability platformfact Apache-2.0; 2024fact Cometfact 21,763/1,744/238; npm 156,598 in Augfact R90 90fact Comet-backedest high release-change riskunknownfact provider-agnostic/self-hostfact enterprise controls in commercial tierunknown telemetryest builders
fact Promptfoofact red-team/evaluation CLI and platformfact MIT; 2023fact Promptfoo; OpenAI acquisition agreement 2026-03-09fact 24,785/2,258/571; npm 2,581,681 in Aug; vendor says 350K developers/130K MAUfact R90 11fact acquisition agreement; price unknownest medium integration/vendor riskfact evaluation results are configuration-specificfact provider-agnostic/local/offlinefact enterprise controls; CLI can run without hosted servicefact telemetry on by defaultest builders
fact DeepEvalfact evaluation frameworkfact Apache-2.0; 2023fact Confident AIfact 18,081 starsfact activeunknown current capital/teamest medium vendor/change riskfact outputs depend on judge model and rubricfact provider-agnostic/localunknown core enterprise controlsunknown telemetryest builders
fact Ragasfact RAG evaluation frameworkfact Apache-2.0; 2023fact Exploding Gradientsfact 15,606 starsfact last push 2026-02-24; R90 0unknownest high staleness riskfact judge/model dependentfact provider-agnostic/localunknownunknownest not suitable as sole standard
fact OpenLLMetryfact OpenTelemetry instrumentationfact Apache-2.0; 2023fact Traceloopfact 7,414 stars; npm 653,998 in sampled monthfact activeunknown current capital/teamest medium vendor/maintainer riskunknownfact provider-agnosticunknown control plane; backend supplies controlsest sends configured traces to selected backendest builders
fact AgentOpsfact agent observability SDKfact MIT; 2023fact AgentOpsfact 5,810 starsfact latest GitHub release 2025-08unknownest high cadence riskunknownfact provider-agnosticfact hosted controls varyunknown telemetry beyond task tracesest not suitable as standard
fact LangWatchfact evaluation/observabilityfact Apache-2.0 core plus EE; 2023fact LangWatchfact 3,522 starsfact R90 70unknown capital/teamest high change/small-team riskunknownfact provider-agnostic/self-hostfact enterprise tier controlsunknown telemetryest builders
fact MLflowfact model/app evaluation and trackingfact Apache-2.0; 2018fact Linux Foundation/Databricks ecosystemfact 27,795 starsfact current 3.13 linefact foundation/corporate-backedest favorable longevity; broad-scope complexityunknown harness ratingfact provider-agnostic/self-hostfact RBAC and SSO plugin pathsfact usage tracking on by defaultest builders
fact Portkeyfact gateway/observabilityfact MIT; 2023fact Portkeyfact 12,890 starsfact last push 2026-05-25; no R90 releaseunknown current capital/teamest medium/high cadence riskunknownfact provider-agnostic/self-host gatewayfact enterprise controlsunknown telemetryest builders, watch
fact Galileofact evaluation/observability platformfact closedfact Galileounknown active customersfact activefact $45M Series B, 2024-10; $68M totalest medium proprietary-vendor riskunknown independent ratingfact provider-agnosticfact enterprise SSO/audit/governanceunknown defaultest builders

What X says — the practitioner signal, last 90 days

The third feed: a Grok Build research thread searched public X posts from June 5 to September 3, 2026, weighting the last thirty days, for what practitioners are adopting, abandoning, praising, and complaining about. Distinct-account mention counts are estimates from sampled queries, not a census; funding and incidents are recorded only when a post or linked announcement states them. It ran the morning of a simultaneous Claude Code, Codex, Cursor, and Grok Build outage — which is itself the argument for owning a floor.

  • FACT (2026-09-03): Claude Code, Codex, Cursor, and Grok Build failed together. Anthropic staff: partial outage across Claude Code, API, claude.ai (@cjav_dev).
  • FACT (JetBrains, cited on X): workplace use Claude Code 39% (from 18% Jan), Copilot 21%, Codex 16% (from 3%), Cursor 12% (from 18%), OpenCode 7%.
  • EST: Default stack is Claude Code + Codex. OpenCode Zen is the outage/free-model hatch. IDEs (Cursor, Copilot) remain, but the argument is harness, not autocomplete.
  • FACT: Muse Spark 1.3 Contributor Free on OpenCode Zen is today’s meme; a minority *do* notice Meta may train (@_yama_yu_, @CDGalpha). Go vs Zen: users pick free if training is already assumed.
  • FACT: Antigravity ToS names OpenClaw-style third-party use as Google-account-suspendable. Staff say stale; written ToS still wins the thread (@GergelyOrosz).
  • FACT as posted: Cursor = SpaceX/xAI ($60B cited via Bloomberg). Cognition/Devin ~$1B at $47B, ARR >$900M. Kilo = Anaconda (price unknown). Cline enterprise = Spec Driven + LG CNS, quiet.
  • FACT: FrontierHarness (360 runs): same model, pass 50–67%, $1.05 (Exo) to $18.34 (Claude Code). Codex if you don’t want to think (@guanlan).
  • EST star-spikes: DeepSeek Harness suspect; Graphify mixed; Ponytail/ECC thin; Caveman is a skill; qm is real YC with launch-star inflation; OpenClaw has real users and abandonment.
  • FACT: Langflow is the CVE story (CISA RCE). Goose is AAIF/LF plumbing, not the daily CLI. OpenHands is the OSS autonomous list + one faculty local stack. Aider is named more than used.
  • FACT: Open WebUI license still draws “not real OSS” drops. LibreChat vs Open WebUI relicensing war is thin this window.
  • FACT: Higher-ed on X is Stanford CS146S (agents course, 2026-09-22) and Claude Campus Ambassadors — not campus-wide harness deployments. UVU unnamed.
  • EST for UVU: keep a local/OpenCode floor; treat Spark contributor-free as public-data only; do not buy star-spike harnesses; Antigravity is a ToS risk until Google rewrites it.

Rankings as seen on X

  • Method (5 lines). (1) JetBrains workplace % when an X post cites it. (2) First-person “I use / I switched / it is down so I cannot work” over listicles. (3) Official maintainer posts count as product-alive, not as adoption. (4) GitHub star spikes on X count for *rising interest*, not *adopted now*, unless diaries exist. (5) Last-30d weighted over last-90d.
  • Most adopted now (est): Claude Code; GitHub Copilot (installed, share falling); Codex CLI; Cursor; OpenCode; n8n (automation lane); ChatGPT/Claude.ai as the chat door; Browser Use in the browser lane; Hermes in the OSS-always-on lane.
  • Fastest rising last 90d (est): Codex (3%→16% cited); Claude Code (18%→39%); OpenCode Zen + Spark 1.3 free; Grok Build (SpaceX stack); Mastra (TS shipping); DeepSeek Harness/Graphify/qm *stars*; Cognition capital.
  • Newest ≤6 months with real traction: Antigravity (conversation, negative ToS); Muse Spark 1.3 (model, not harness); Stagehand v4; Mastra computer-use; qm (YC, launch spike); AQ meta-harness; Factory public-sector. “Real” here means named workflows or eval inclusion, not stars alone.
  • Declining or dead (est): Windsurf as a brand (Cognition/Google split residue); Copilot *share*; Aider vs new CLIs; Roo Code (thin/likely archived); Continue (listicle only); personal OpenClaw at Every; Langflow *trust*; AutoGen for new builds.

Method: Native X keyword + semantic search, public posts only. Windows: last 30 (since:2026-08-04) and last 90 (since:2026-06-05). Latest plus min_faves filters. No DMs, no posting. Distinct-account mention counts are est: unique handles in returned samples (max 10 per query), scaled by how many distinct queries hit the name and whether official accounts dominate. Not a census. Funding, acquisitions, license changes, and incidents are fact only when a post or linked announcement states them. No invented numbers. Rankings prefer practitioner “I use / I switched” over listicle spam. Star counts on X are interest, not adoption. Misses are logged. A miss is not proof of absence.

Star-spike provenance — are the new "100k-star" harnesses real?

ProjectX verdictWhy
DeepSeek HarnessSuspect-until-proven100k stars in <48h claimed; FrontierHarness includes “DSH” as a real harness under test (good). Counter: Julian Goldie SOP-for-DM videos, crypto “TaiYi Super Agent” spam, almost no named production teams. Pin versions if you touch it.
PonytailThin users / mixedTrending (+1.3k/day cited). Metric claims (54/20/27%) are promo. A few “old oil” skill posts. Not a daily-driver.
ECCUnverifiedTrending explainers only. No first-person ops posts in this sample.
GraphifyMixed — demo real, scale unverifiedViral 100k-star video (208k views). One engineer: it cuts reread-token burn. Used as a clone-demo target. Not seen as a workplace standard.
CavemanNot the spike people thinkIn this window it is a token-saving *skill* (“talks like caveman”), not a 100k-star platform. Full spike check still open.
OpenClawReal users, messy opsToS-named, Steinberger “we all use it,” Every killed personal Claws for a shared Slack agent. Stars are inflated relative to maintainability, but this is not a ghost repo.
qmReal org, marketing-heavy spikeYC open-sourced Jul 2026 as the firm’s Slack+web multiplayer harness (MIT, yc-software/qm). Launch posts claim 1.6k–13k stars in days and internal accounting/legal/eng use. Last-30d first-person ops almost absent; Sep conversation moved on. Treat as real product with inflated launch stars.
PaperclipReal niche, not a spike scamOfficial release cadence, adapters, orchestrator comparisons. Name-collides with paperclip-AI-risk essays and a patent search tool.

Funding and ownership moves stated on X (dated)

DateItemAmountSourceGrade
2026-09-03NVIDIA intends to acquire Hugging Face$12,930,300,000@ClementDelanguefact (announcement post). Not a coding harness; cited because X grouped it with Cursor/Kilo.
2026-09-03Anaconda acquired Kilo Codeprice not stated@GargEtisha; @kilocode biofact acquisition; unknown price
last 30d (cited 2026-09-03)SpaceX acquired Cursorprice not stated in sampled postsmultiple, e.g. @KevinhoMoralesfact acquisition; unknown price
2026-09-02SpaceX Cursor deal$60B@wallstengine citing Bloombergfact as that post states; primary Bloomberg not opened
2026-09-02Cognition (Devin) new round~$1B at $47B val; ARR >$900M (from $492M late May); ~$10B demand@wallstengine citing Bloombergfact as posted; primary unknown
last 30d (cited 2026-09-03)Stripe acquired OpenRouterprice not stated@GargEtishafact as stated by that post; not independently confirmed in this batch
2026-08-29 (Grok summary)Google DeepMind hired Windsurf founders + $2.4B tech license; Cognition took remainder$2.4B@grokest/secondary until primary 2025 posts are fetched

Incidents and stability signals on X (dated)

DateItemSourceGrade
2026-09-03Partial outage: Claude Code + API + claude.ai@cjav_devfact
2026-09-03Codex/ChatGPT down same morning@ProductHunt, manyfact (widespread user + PH)
2026-09-03Cursor + Grok Build also failing in user screenshots@sora_bizfact as user report
2026-09-03OpenCode Zen Spark 1.3 free live; some 500s on chat-completions@hppzyyfact user
2026-09-03Antigravity ToS / Google-account ban debate@theo, @GergelyOroszfact posts; enforcement unknown
2026-09-02n8n agent silently removed 92% of AI nodes (HN via X)@echo_vicfact as cited; original HN not opened this batch
2026-09-02GitSpawn: goose 1.44.0 patched; Hermes/Qwen Code/Grok Build still open on retest@mkeremturhanfact researcher claim
2026-09-01“One flaw affects nearly every major AI coding agent” — 8 findings, 4 unpatched@Ax_Sharmafact as journalist thread; details not fully fetched
2026-06-08Cline Spec Driven enterprise with LG CNS@clinefact
2026-08-05 (cited 08-27)Langflow unauth RCE, CISA active exploitation@SnowCrashLabsfact as cited
2026-08-10Stagehand v4@Stagehanddevfact
2026-09-02Mastra sandbox computer-use@calcsamfact
2026-09-02FrontierHarness Eval published@guanlanfact
2026-09-03Browser Use × Link agent cards@browser_usefact
2026-09-03Muse Spark 1.3 contributor-free on OpenCode Zenmany; training caveat @CDGalphafact

Higher education sightings

DateSightingSourceGrade
2026-09-02Stanford CS146S *The Modern Software Developer* (Fall 2026): 85% of 2025 material thrown out; agent skills, MCP, AGENTS.md, software factories; OSS PR requirement with OpenHands, Browserbase, CrewAI, Warp, Vercel, Pi, etc.; guests from Cursor, Claude Code, Factory, Cognition, Replit. Campus start 2026-09-22; materials public.@mihail_ericfact (instructor announcement)
2026-09-02Claude Campus Ambassadors 2026–27: undergrad/grad/PhD tracks; $3,600 scholarship cited; apply by 2026-09-12@claudeaifact
2026-09-01Iowa State: SpaceXAI Grok Bot build night 2026-09-03; free Cursor Pro for attendees@jackxlaufact event post
2026-09-01Hult International Business School: 2026 stack includes Cursor Boston, Lovable, ChatGPT, Claude, Gemini, Copilot@HULT_JPfact school comms
2026-08-31University security faculty: DGX Spark ×2 + vLLM + qwen3.8-flash-next + OpenHands local stack@valdzonefact post; campus-wide unknown
2026-09-03Student moved housing to afford extra Claude Code Max@enjojoyyfact anecdote, not a deployment
UVU-named public deploymentunknown

No campus-wide harness deployment by any university surfaced on X in the window; the sightings are courses, ambassador programs, build nights, and one security faculty member's local stack. UVU is not named anywhere. That absence is consistent with the guide's first-mover finding.

The X master table — 85 tools with momentum, incidents, and what people say they are best and worst at

#ToolOne-lineCategoryOpen / closedCompany / backerFunding (stated only)X momentum, last 30 daysStability signalsScores people citeBest at / worst at
1OpenCodeModel-agnostic CLI/TUI coding agent; Zen is the hosted model router (Go paid, Zen/contributor-free routes).Coding CLIOpen (MIT claimed in promo posts)Anomaly / @thdxrunknown in this X sampleHigh. Distinct handles in Zen/Spark posts today; JetBrains 7% cited. Method: 8+ unique handles in Zen/Spark queries plus survey cite. Sentiment lean: positive-as-escape-hatch.fact: Zen Spark 1.3 free is live; some 500s unless /v1/responses (Codex-compat) is used (@hppzyy).unknown official harness score. People compare Spark-on-OpenCode vs Gemini 3.8 by feel.Best: keep working when Claude/Codex/Grok die; cheap/free models. Worst: not for non-programmers (@Anurag_barl).
2Claude CodeAnthropic’s CLI/cloud coding agent. Current default in the JetBrains workplace survey.Coding CLIClosed clientAnthropicunknown amount on X this windowVery high. Outage + survey + switch posts. 10/10 latest “switched” hits involved it or Codex. Sentiment: default, plus reliability anger.fact: partial outage 2026-09-03 covering Claude Code, API, claude.ai (@cjav_dev). Practitioner tally: “3 major outages + 6 degraded last month” (@alexgetmancom) — est, not vendor.JetBrains 39% workplace use. Terminal-Bench cited as model+harness, not harness-only.Best: default agentic work. Worst: capacity/outages; cost (“overpriced P.C. sludge” in one SaaS-stack post).
3Codex CLIOpenAI coding agent (CLI + app). Fastest-rising in the cited JetBrains numbers.Coding CLIOpen CLI (Apache claimed in other lanes; not re-stated on X here)OpenAIunknown on XVery high. Paired with Claude Code in every outage thread. Sentiment: preferred by some switchers; limit/outage complaints.fact: ChatGPT/Codex down 2026-09-03; OpenAI posted during the incident. First dual Claude+Codex outage noted (@LexnLin).JetBrains 16% (from 3%). Terminal-Bench 4.0 runs paired with gpt-5.6-luna.Best: concise, rule-following, easy switch from Claude projects (@JamesonCamp). Worst: limits; app vs CLI quality split (@scottzirkel).
4Gemini CLIGoogle’s open coding CLI; now discussed next to Antigravity and account-ban risk.Coding CLIOpen (Apache claimed in listicles)GoogleGoogle-backed; no round on XMedium-high. Less “I use Gemini CLI daily,” more ban/ToS adjacency.fact: @GergelyOrosz says Gemini CLI access was hit when Antigravity bans landed. Consumer CLI path already in transition (prior lane; not re-confirmed here).unknown harness-only. Gemini 3.8 Flash cost-per-task up on Artificial Analysis (@agentfred_ai).Best: huge Flash limits on $20 plans (@championswimmer). Worst: Google-account blast radius.
5AntigravityGoogle DeepMind agentic IDE. ToS currently the story, not the product.Coding IDEClosed previewGoogle / @\_mohansoloGoogle-backed; Windsurf-team origin is contested on XHigh (negative). Theo 793k views; Gergely 72k. Sentiment lean: avoid.fact: ToS cited as allowing Google-account suspension for third-party/OpenClaw use (@GergelyOrosz). Counter: Varun Mohan said bans were product-only (quoted). Shadow-ban support complaints exist.JetBrains 6% in prior lane; not restated in this X sample.Best: Gemini 3.8 Flash inside Google surfaces (@DjowDj0w). Worst: ToS vs Google-account risk; “ugly/hard” vs Windsurf-team expectations (@LinearUncle).
6CursorClosed AI editor, now a SpaceX/xAI asset.Coding IDEClosedCursor → SpaceX (fact on X)Acquisition price unknown on XVery high. Survey drop + SpaceX narrative + today’s outage. Sentiment: mixed; still in every stack poll.fact: outage 2026-09-03 with Claude/Codex. fact: ambassadors renamed SpaceXAI (@KevinhoMorales). est: OpenAI access cut talk (28 Aug notice / 12 Nov shutoff) is circulating, not confirmed here from OpenAI’s own account.JetBrains 12% (from 18%). CursorBench cited by @GavinSBaker.Best: Composer for simpler tasks; GUI. Worst: vendor concentration after SpaceX; model-provider cutoff risk.
7WindsurfAI editor; Cognition-owned remainder after the Google/Cognition split.Coding IDEClosedCognition (product); Google hired founders (X summaries)No new round in this sample. Historical $2.4B Google license / Cognition remainder is summarized by @grok — treat as est until primary post is fetched.Low-medium. Mostly listicle residue. Few first-person “I switched to Windsurf” hits.fact: “Anthropic pulled the plug on Windsurf” as a 2025-era event still cited (@kevinsxu). est: declining mindshare vs Claude Code/Codex/Cursor.unknown current.Best: still named as a paid option. Worst: not the conversation; Antigravity comparison is unkind.
8ClineOpen-source IDE/CLI coding agent; enterprise Spec Driven with LG CNS.Coding IDE/CLIOpenClineunknown amount on X this windowMedium. Official posts + OSS listicles + one enterprise inbound. Not in today’s outage pantheon.fact: Cline Spec Driven enterprise platform with LG CNS, 2026-06-08 (@cline). fact: inbound “enterprise license for our team” 2026-09-03 (@shettygirish75). npm-compromise history not discussed in this 30d sample.Marketplace ratings not cited on X here.Best: editor+terminal+browser agent; BYO model. Worst: quieter than Claude/Codex; enterprise story not viral.
9Kilo CodeOpen-source agent for VS Code, JetBrains, CLI; posting as acquired by Anaconda.Coding IDE/CLIOpenKilo / AnacondaAcquisition fact; price unknownMedium. Official @kilocode volume; third-party mentions are listicles. Sentiment: vendor-positive, community-thin.fact: handle/bio “Kilo (acq. by Anaconda)”; grouped with Cursor/OpenRouter deals in last 30 days (@GargEtisha). Shipping JetBrains multi-agent control room (2026-09-01).unknown independent. Demo: Fable 5.1 recreate-X for $13 (@kilocode).Best: OSS hedge vs closed SF stack (their blog). Worst: little independent “I switched to Kilo after Anaconda” testimony in this sample.
10GooseBlock-origin local agent; now under AAIF / Linux Foundation orbit.Coding/general agentOpenBlock → AAIF (@goose_oss, @AgenticAIFdn)Corporate/foundation; no roundMedium. Security + OSS-list + MCP demos. Not a daily-driver meme.fact: GitSpawn fsmonitor patched in goose 1.44.0 (@mkeremturhan). fact: AAIF ambassador shipped goose-extension-audit (@AgenticAIFdn).unknown.Best: MCP recipes, local/multi-provider, foundation governance. Worst: lower X volume than Claude/Codex; still a “list of 10” item.
11OpenHandsAutonomous coding platform (inspect, edit, test, PR). Agent Canvas as control plane.Coding agent platformOpen (MIT claimed)All Hands / @OpenHandsDevunknown on XMedium. Star-count listicles (85.7k) plus one faculty-adjacent local stack.fact: OpenHands SDK now defaults Browser Use (@mamagnus00). Databricks event with OpenHands on security agents (Sep 1).Star counts cited (78.5k–85.7k). Not harness-only benches.Best: autonomous issue→PR; self-host; local+vLLM stack (@valdzone). Worst: Docker/setup friction; promo-list noise.
12AiderTerminal pair-programmer with git-aware edits.Coding CLIOpenIndependentunknownLow-medium. Named in “what’s your harness” polls and OSS lists more than first-person 30d diaries.unknown incidents this window.Aider polyglot not cited in this 30d sample.Best: terminal+git, any model. Worst: mindshare lost to Claude Code/Codex/OpenCode.
13LibreChatSelf-hosted multi-model chat UI.Chat front doorOpenLibreChatunknownLow. Thin 90d X. One MCP success vs Open WebUI failure.unknown license-change posts this window.unknown.Best: Playwright MCP over SSE after Open WebUI failed (@mfaisal_khatri). Worst: almost invisible vs Open WebUI in this scan.
14Open WebUISelf-hosted local/cloud model UI. Branding/enterprise-license friction.Chat front doorSource-available (branding constraint cited)Open WebUIunknownMedium-low. License complaints + enterprise fork talk, not daily coding-agent chatter.fact: users drop it because “the license is stupid. It's not a real open source license” (@libertypenguin0). Branding/white-label needs enterprise license (Grok summary 2026-06-20). Forks of v0.6.5 for MIT-friendlier stacks.unknown.Best: local RAG/gateway. Worst: license/branding for white-label; MCP gaps vs LibreChat in one practitioner thread.
15n8nWorkflow automation with AI-agent nodes; still the default “no-code agent” on X.Workflow/agent automationOpen core + cloudn8nunknown on XHigh (automation lane), medium (coding-agent lane). Distinct from CLI coding; huge tutorial volume.fact: agent silently deleted 92% of AI nodes in a dataset (HN, cited 2026-09-02) — governance miss (@echo_vic).unknown coding benches.Best: Trigger→Brain→Memory→Tools without code. Worst: change-control; some builders now prefer Python/agent scripts over n8n (@uglyrobot).
16DifyVisual agentic-workflow builder. Official account still posting “build your first agent” guides.Workflow/agent platformOpen + cloudDifyunknown on XMedium-low. Official posts; little practitioner debate vs n8n/Langflow.unknown incidents this window.unknown.Best: start-from-a-goal onboarding. Worst: not in the coding-agent default stack.
17LangflowVisual agent/RAG builder (IBM-associated in security writeups).Workflow/agent builderOpenIBM / LangflowunknownMedium (security). Mentions are CVE/RCE, not adoption.fact: CISA-confirmed unauth RCE, default no login (@SnowCrashLabs). Repeat exec() sink (CVE-2025-3248 and CVE-2026-33017) (@BRuteLogic). “11 more CVEs” cited 2026-09-02.unknown.Best: visual RAG/agent graphs. Worst: internet-exposed instances; secrets next to exec().
18CrewAIMulti-agent “crews” framework.Agent frameworkOpen + platformCrewAIunknown on XMedium-low. Listicle staple; one telemetry critique.est: “78 observe-only hooks… that’s telemetry” (@_kvnloo).unknown.Best: named multi-agent starter. Worst: bloat/observe-only; skipped by some greenfield builders.
19MastraTypeScript agent framework; shipping sandbox computer-use and Render deploy.Agent frameworkOpen core + commercial EE (prior lane; not re-stated on X)Mastra / @calcsamunknown amount on XMedium. Founder launch posts + Render partnership. Sentiment: builder-positive.fact: sandbox computer-use launch 2026-09-02 (@calcsam). Suspended-run recovery across servers (@mastra).unknown harness benches.Best: TS agents, long-running workflows, computer-use in sandbox. Worst: not a coding CLI; framework not daily IDE.
20Pydantic AITyped Python agent framework from the Pydantic team.Agent frameworkOpenPydanticunknown on XMedium-low. CopilotKit channel glue + “type-safe outputs” listicles.unknown incidents.unknown.Best: schema-valid agent outputs. Worst: not a coding harness; library not product.
21Browser UseOSS + cloud “agents that use the browser”; payments and cookie-sync this week.Browser agentOpen + cloudBrowser Useunknown amount on XHigh. Official product + staff departures to SpaceXAI.fact: Link one-time cards for agent checkout 2026-09-03 (@browser_use). fact: engineer #2 left after “20k to over a million monthly agent runs” (@Alezander9). Staff → SpaceXAI (@larsencc).Odyssey #1 claimed in promo posts.Best: real-browser loop; now in OpenHands SDK default. Worst: action/payment risk; telemetry not discussed on X here.
22StagehandBrowser-agent SDK by Browserbase; v4 self-healing + WebMCP.Browser SDKOpenBrowserbaseunknown (parent Series B not restated on X)Medium. Official v4 (2026-08-10) still circulating.fact: v4 caching up to 80% faster claimed (@Stagehanddev). WebMCP first-class.unknown independent.Best: Playwright-for-agents; iframe/self-heal. Worst: Browserbase cloud coupling.
23BrowserbaseHosted browser runtime for agents.Browser infraClosedBrowserbaseunknown on XMedium. Identity/payments/Ramp, not coding-IDE talk.fact: Ramp agent-identity partnership (@browserbase).unknown.Best: authenticated web + contexts. Worst: hosted-browser data plane.
24Muse Code + Spark 1.3Meta’s coding agent (Muse Code) and model (Spark 1.3). Free contributor route via OpenCode Zen.Coding IDE + modelClosedMetaMeta-backed; no roundVery high (model), medium (Muse Code product). Spark free is today’s meme.fact: contributor-free may train (@CDGalpha). fact: Muse Code $15 plan limit complaint (@dalu_hey).AA cited; Meta DeepSWE/Terminal-Bench in Grok replies. Independent 3D-animal loop: Spark cheaper than Gemini 3.8 Flash (@thehypedotnews).Best: free/cheap coding while Claude/Codex down. Worst: training terms; Muse Code limits; not local weights.
25DeepSeek HarnessPlugin-everything coding harness; 100k GitHub stars in <48h claimed.Coding harnessOpen (MIT claimed)unknown org on XunknownHigh (hype), low (verified use). Promo videos + trending bots.fact as claimed: 100k stars <48h, faster than OpenClaw (@kubesimplify). In FrontierHarness set as “DSH.”FrontierHarness: wall-clock play, knobs. Star counts 135k-in-days in promo (@JulianGoldieSEO) — treat as est/promo.Best: swappable loop/plugins. Worst: star-spike + affiliate SOP-for-DM pattern; pin versions. Provenance: suspect-until-proven.
26PonytailSkill/harness that forces reuse-before-code. Trending GitHub.Coding skill/harnessOpenunknownunknownMedium (trending), low (use).fact: GitHub trending +1.3k stars 2026-09-02 (@soresearcher). Claimed 54% less code / 20% lower cost / 27% faster on 12 tasks (@gitgoats) — vendor/promo, not replicated here.unknown independent.Best: “do less” ladder. Worst: star velocity vs first-person production posts. Provenance: mixed/thin users.
27PaperclipOSS control panel for a team of agents (skills, adapters, operators).Agent orchestratorOpenpaperclipai / @paperclipingunknownMedium. Official releases + orchestrator roundups. Name collision with “paperclip problem.”fact: v2026.831.0, Kimi Code adapter, 175 commits / 13 contributors (@papercliping). Telemetry/observability/run-log naming split in v2026.824.0.unknown.Best: agent-as-employees dashboard. Worst: early, issue-heavy (prior lane); not a coding CLI. Users: real but niche.
28OpenClawPersonal/general autonomous agent (Claw). Named in Antigravity ToS.General agentunknown license on X@steipete / OpenClawunknownHigh. ToS villain, real users, and abandonment in the same week.fact: Every “killed the last of our Claws” — logins expired, integrations broke (@every). fact: ToS example for Google bans. Peter Steinberger: “We all use it.”unknown benches.Best: personal always-on agent. Worst: maintenance; account-ban coupling; governance. Users: real, not only stars.
29ECCRepeatable plan/build/test/review toolkit for coding agents (skills, hooks, scans).Coding toolkitOpen (MIT in prior lane)unknownunknownLow-medium. GitHub-trending explainers, almost no “I run ECC daily.”unknown incidents.unknown.Best: shared process across Claude Code/Codex/OpenCode. Worst: star-without-users pattern. Provenance: unverified.
30GraphifyLocal-folder → knowledge graph for coding agents (AST, no vectors claimed).Context harnessOpenGraphify-Labs / @safishamsiiunknownMedium-high (viral demo). One 208k-view ES video.unknown security.100k GitHub stars claimed in viral post (@SofiaSici).Best: stop agents rereading files (@i_mika_el). Worst: star-spike; few independent production writeups. Provenance: mixed — demo real, scale unverified.
31Grok BuildxAI coding agent (CLI/client). In today’s outage set and GitSpawn unpatched list.Coding CLISource published (prior lane)xAI / SpaceXxAI-backedHigh. Outage screenshots + release notes + stack posts.fact: down 2026-09-03 with Claude/Codex. fact: GitSpawn still open on retest (@mkeremturhan). v1.0.18 managed MCP policy (@XFreeze).unknown independent harness score.Best: daily coding with Grok 4.6 (@elshayib_). Worst: unpatched GitSpawn; issues-disabled governance (prior lane, not re-litigated here).
32GitHub CopilotMicrosoft/GitHub IDE+cloud coding agent. Survey share falling; Slack/Teams agent shipping.Coding IDE/cloudClosedMicrosoft/GitHubMicrosoft-backedHigh (installed base), mixed sentiment. Still #2 in JetBrains cite.fact: Copilot in Slack and Teams, Aug ship (@github). Practitioner: GitHub-hosted issue→PR agent (@Equinox_enx).JetBrains 21% (from 29% a year ago).Best: GitHub-native issue agent; student free. Worst: losing agent mindshare to Claude Code/Codex.
33AmpSourcegraph coding agent (TUI/desktop/web/mobile).Coding agentClosedSourcegraph / @ampcodeunknown on XMedium. Enthusiastic niche, not survey-scale.fact: intelligent diff-order button (@beyang).unknown.Best: cross-surface harness (@anthonywu). Worst: paid closed; quieter than Claude/Cursor.
34DevinCognition cloud software agent; Windsurf remainder lives here.Cloud coding agentClosedCognitionfact as posted: ~$1B round at $47B, ARR >$900M, Bloomberg via @wallstengine. Independent filing unknown.Medium-high (capital), medium (practitioners).Fusion pricing claims vs Fable 5.1 (@notjazii) — vendor-adj. SWE-1.7 on Cerebras 1,000 t/s.unknown independent harness-only.Best: unattended cloud agent; cheap Fusion routing. Worst: another vendor desktop; not the CLI default.
35CrushCharmbracelet terminal coding agent; FSL license.Coding TUISource-available (FSL)CharmbraceletunknownLow-medium. 27k-star listicles; sandbox mention.smolBSD sandbox for crush (@iMilnb).unknown.Best: pretty terminal agent. Worst: FSL; not in JetBrains top slice.
36Qwen CodeAlibaba/Qwen terminal coding agent.Coding CLIOpenAlibaba/QwenAlibaba-backedMedium-low. GitSpawn unpatched; Kilo collab.fact: GitSpawn still open on 0.19.6/0.22.3 (@mkeremturhan). Qwen3.8-Max #4 Arena frontend, demoed on Kilo (@Alibaba_Qwen).Arena frontend #4 (model, not harness).Best: OSS CLI in Qwen ecosystem. Worst: unpatched GitSpawn; jurisdiction.
37Hermes AgentNous-orbit general/coding agent; 24/7 worker meme.General + coding agentOpenNous / communityunknownHigh in OSS-agent lane.fact: GitSpawn still open 0.18.2/0.21.0. Consulting talk of warehouse/Slack/CTO paths (@tonysimons_).unknown.Best: local 24/7 with cheap/open models. Worst: unpatched GitSpawn; “trips over itself” in promo comparisons.
38PiMinimal coding-agent toolkit; FrontierHarness winner on cost.Coding harnessOpenIndependent (named in eval)unknownMedium. Eval-famous more than viral.fact: 90 turns / $2.50 vs Claude Code 381 / $64.36 on same DeepSWE fix (@guanlan).FrontierHarness: use if the job repeats.Best: cheap pass rate. Worst: knobs; less “just works” than Codex.
39ExoHarness in FrontierHarness; cheapest $/pass.Coding harnessunknownunknownunknownLow-medium. Eval-only in this scan.unknown ops.FrontierHarness: $1.05/pass; quit-early.Best: retries cheap. Worst: almost no practitioner diary.
40LangChainPython agent/LLM framework + LangSmith.Agent frameworkOpen + cloudLangChain Inc.unknown amount on X this windowMedium. CEO talks, OpenWiki, LangSmith deploy.fact: LangChain coding agent 52.8%→66.5% Terminal Bench 2.0 without changing model (@nykdotdev). Podium migrated to LangSmith deployments (@LangChain).Terminal Bench 2.0 harness-not-model.Best: production harness + traces. Worst: bloat; some skip it greenfield.
41LangGraphStateful graph agent runtime (LangChain).Agent frameworkOpenLangChain Inc.parentMedium. DeerFlow built on it; Stanford treats it as homework parts.unknown new incidents.unknown.Best: multi-agent graphs. Worst: “implementation part not research” (@momiji_fullmoon).
42Google ADKGoogle agent SDK; workshop circuit.Agent SDKOpenGoogleGoogle-backedMedium. Tutorial volume, not coding-CLI default.unknown.unknown.Best: GCP multi-agent patterns. Worst: Google-account adjacency (see Antigravity).
43AutoGenMicrosoft multi-agent; maintenance mode.Agent frameworkOpenMicrosoftMicrosoft-backedLow. Listicles note successor.fact as listed: maintenance; new work → Agent Framework (@kv1nsiii).unknown.Best: tutorial reference. Worst: not for new builds.
44ContinueOSS IDE coding assistant.Coding IDEOpenContinueunknownLow. Local+Ollama questions; listicles.unknown shutdown posts in this sample (prior lane said maintenance retreat — not re-confirmed on X here).DeepSWE 63.4 claimed in one promo.Best: local VS Code. Worst: not the 2026 conversation.
45Roo CodeCline-family IDE agent.Coding IDEOpenRoo Code Inc.unknownLow. History extractors still name it; almost no 30d first-person.unknown archive posts this window (prior lane archived 2026-05-15 — not re-confirmed here).unknown.Best: historical Cline fork. Worst: declining/dead on X.
46Factory / DroidModel-agnostic software agent; public-sector push.Cloud/IDE coding agentClosedFactoryunknown amount on XMedium. Readiness criteria; Carahsoft.fact: Factory + Carahsoft public sector (@FactoryAI). Long unattended runs 1–8h (@garrettwinderrr).unknown.Best: agent-readiness + autonomy. Worst: not CLI-default.
47Augment / CosmosEnterprise coding agent → “OS for agentic software development.”Coding platformClosedAugmentunknown on XLow-medium. Official Cosmos Advisor posts.unknown.unknown.Best: factory-scale prompt box. Worst: little independent X.
48LovablePrompt-to-app builder; Slack @Lovable.App builderClosedLovableunknown on XMedium (builders), low (harness nerds).fact: Slack tagging (@Lovable). Hult curriculum names it.unknown.Best: idea→app. Worst: not a repo agent; token-middleman complaints vs Claude-direct.
49Replit AgentBrowser IDE + agent + host.Cloud IDE agentClosedReplitunknown on XMedium-low in this coding-agent sample; still in vibe stacks.unknown incidents.unknown.Best: agent+host $20. Worst: not Claude Code/Codex discourse.
50qm (Quartermaster)YC-open-sourced multiplayer company harness (Slack+web).Agent orchestratorOpen (MIT claimed)YC / yc-softwareYC-backed; no roundMedium at launch (Jul–Aug), low in last 30d.fact as posted: MIT, Slack+web, harness-agnostic (Pi/OpenCode/Codex/Claude Code), used internally for accounting/legal/eng. Stars 1.6k–13k in days across posts.unknown benches.Best: company-wide rooms. Worst: launch-spike, few Sep first-person ops. Provenance: real org (YC), star-spike still marketing-heavy.
51WarpAgentic terminal. Stanford OSS partner list.Agentic terminalClosed clientWarpunknown on XLow-medium. Named in course partners more than diaries.unknown.unknown.Best: terminal+agent. Worst: quiet vs CLI agents.
52KiroAWS coding IDE/CLI path (Q successor in prior lane).Coding IDEClosedAWSAmazon-backedLow. Giveaway spam + one OSS “Kiro Crew” workspace (may be unrelated).unknown.unknown.Best: AWS-native. Worst: not the X default.
53v0Vercel UI generator.App builderClosedVercelVercel-backedLow-medium in vibe stacks.unknown.unknown.Best: UI. Worst: not a repo agent.
54BoltPrompt-to-hosted-site.App builderClosedStackBlitzunknownLow-medium.unknown.unknown.Best: hosted site fast. Worst: not enterprise coding.
55OllamaLocal model runner; pairs with OpenCode/Continue/Open WebUI.Local inferenceOpenOllamaunknown on XMedium as the local floor, not a harness.unknown.unknown.Best: local weights. Worst: not an agent.
56ZedFast editor; AI mentioned as respected.EditorOpen coreZedunknownLow in agent discourse.unknown.unknown.Best: editor. Worst: not agent-default.
57SWE-agentResearch GitHub-issue agent; mini-SWE successor path (prior).Research coding agentOpenPrinceton/communityAcademicLow. OSS lists only.unknown.SWE-bench is the point.Best: research. Worst: not daily driver.
58PlandexLong-running coding engine.Coding agentOpenPlandexunknownLow. OSS lists.unknown.unknown.Best: long tasks. Worst: thin X.
59DeerFlowByteDance LangGraph harness; 81.3k stars cited.Agent harnessOpen (MIT cited)ByteDanceByteDance-backedMedium as listicle star.unknown.Stars 81.3k cited.Best: research+chat apps. Worst: jurisdiction; listicle.
60Microsoft Agent FrameworkAutoGen/SK successor.Agent SDKOpenMicrosoftMicrosoft-backedLow-medium. Named as the new path.fact: AutoGen maintenance → this (@kv1nsiii).unknown.Best: new MSFT builds. Worst: early.
61OpenAI Agents SDKOfficial Python agents SDK.Agent SDKOpenOpenAIOpenAI-backedLow vs Codex CLI.unknown.unknown.Best: OpenAI-shaped apps. Worst: not the coding CLI.
62StrandsAWS agent SDK; DynamoDB memory post.Agent SDKOpenAWSAmazon-backedLow.fact: strands-dynamodb-storage blog cited (@VKazulkin).unknown.Best: AWS memory. Worst: thin X.
63LlamaIndexData/agent framework.FrameworkOpen + cloudLlamaIndexunknown on XLow this window (listicles).unknown.unknown.Best: private data. Worst: not coding CLI.
64HaystackRAG/agent framework.FrameworkOpendeepsetunknownLow.unknown.unknown.Best: RAG. Worst: quiet.
65FlowiseVisual LLM builder.WorkflowOpenFlowiseunknownLow. Not in 30d coding-agent hits.unknown (prior lane had sunset risk — not X-confirmed here).unknown.Best: visual. Worst: miss.
66PromptfooEval/red-team for prompts/agents.EvalOpenPromptfoounknownLow. Not retrieved this window.unknown.unknown.Best: eval. Worst: miss.
67CodeRabbitPR review agent.Review agentClosedCodeRabbitunknown on X (prior $143M not restated)Low this window.unknown.unknown.Best: review. Worst: miss.
68Chrome DevTools MCP / Playwright MCPBrowser debugging MCP for agents.Browser MCPOpenGoogle / Microsoftvendor-backedMedium. One 2026-09-03 explainer.fact: Chrome DevTools MCP for Antigravity, Claude, Cursor, Copilot (@FReza1984).unknown.Best: observable browser debug. Worst: still MCP, not a harness.
69Open InterpreterLocal computer-use / code agent.Computer-useOpenOpen InterpreterunknownLow. OSS lists.unknown rewrite (prior lane).unknown.Best: talk-to-computer. Worst: thin 30d.
70ManusCloud general agent; Meta-owned (prior).General agentClosedMetaunknown priceLow this window vs coding CLIs.unknown.unknown.Best: general agent. Worst: miss vs Claude/Codex.
71ChatGPTConsumer/enterprise assistant; outage today.AssistantClosedOpenAIunknown on XVery high as infra, not harness.fact: down 2026-09-03.unknown.Best: chat. Worst: coding-agent users bounce to Codex CLI.
72Claude.aiConsumer/enterprise assistant.AssistantClosedAnthropicunknownVery high.fact: login/API/Code outage.unknown.Best: chat+code. Worst: capacity.
73GeminiGoogle assistant.AssistantClosedGoogleGoogleHigh. 3.8 Flash cost-per-task debate.ToS/account-ban adjacency.AA cost/task up 40% same sticker price.Best: cheap Flash. Worst: account blast radius.
74PerplexitySearch assistant.AssistantClosedPerplexityunknownLow in coding-agent lane.unknown.unknown.Best: research. Worst: not a coding harness.
75M365 CopilotOffice agent.Enterprise assistantClosedMicrosoftMicrosoftMedium office, low coding.fact: 2026-08-31 M365 auth incident cited in one roundup.unknown.Best: Outlook/Teams. Worst: not SWE.
76AQ“Harness of harnesses” — multiplayer over Claude Code/Codex/OpenCode/Cursor/Devin.Meta-harnessunknown@aqdotdevunknownLow-medium. Launch 2026-09-02.unknown.unknown.Best: team + many harnesses. Worst: brand-new.
77CavemanToken-saving “talk like caveman” skill.Agent skillOpenJuliusBrussee (list URL)unknownLow. Skill lists, not platform.unknown.unknown.Best: fewer tokens. Worst: not the 100k-star platform in this window.
78TabbySelf-hosted completion.Coding assistantOpenTabbyMLunknownLow. OSS lists.unknown.unknown.Best: self-host complete. Worst: agent era passed it.
79HerdrCLI to manage multiple coding-agent sessions.OrchestratorOpenherdrdevunknownLow-medium. Claude/Hermes/Grok Bot plugins.unknown.unknown.Best: many terminals. Worst: another layer.
80oh-my-hermesSkills/routing pack for Hermes Agent.Harness packOpen@rlaopeunknownLow-medium. Author-heavy.unknown.unknown.Best: Hermes coding power-up. Worst: one-maintainer.
81LiteLLMModel gateway.GatewayOpenBerriAIunknownLow this window except security roundups.Named next to Langflow RCE targets.unknown.Best: campus router. Worst: attack surface.
82MCP (protocol)Tool protocol under AAIF/LF.ProtocolOpenAAIF / Anthropic originfoundationHigh as plumbing.A2A also under AAIF (2026-08-20 cited).unknown.Best: common tools. Worst: not a product.
83Kimi CodeMoonshot coding agent; Paperclip adapter.Coding agentOpen claimedMoonshotunknownLow-medium.Paperclip adapter fact.unknown.Best: Kimi quota. Worst: thin global X.
84ZiteApp builder that uses your Claude/ChatGPT/Cursor sub.App builderClosedZiteunknownLow. Launch today vs Lovable/Replit double-pay.unknown.unknown.Best: no extra credits. Worst: new.
85Funes (HF)Session memory across Claude Code/Codex/Pi/Hermes.Memory layerunknownHugging FaceNVIDIA deal todayLow. One Chinese explainer.unknown.unknown.Best: don’t re-archaeology sessions. Worst: new; HF acquisition noise.

Source: Grok Build X scan, work id uvu-x-harness-trends-v2, September 3, 2026 — 85 tools, source log of hits and misses in the working papers. Quotes are limited to 25 words; a miss is not proof of absence.

How this sweep repeats

The method is written down step by step — discovery sweep, repository snapshot, adoption signals, funding and ownership, stability and controls, benchmark handling — so it can be re-run monthly and the deltas tracked. It lives with the plan's working papers; the standing rule is that any landscape question gets the trending, funding, growth, and social-signal dimensions by default, never adoption alone.

Survey: working paper 35, September 3, 2026 (171 entries, source log with hits and misses). GitHub data: API snapshot, same day. X signal: Grok Build scan, same day (its first run died on the provider's capacity errors after forty minutes; the rerun wrote as it went and completed).

N

Adoption, audited

Our faculty-adoption plan checked against the six sources of influence and against what actually moved mass adoption cheaply, elsewhere

Appendix N in five lines

  1. QuestionHow can Utah Valley University get teachers to use artificial intelligence at low cost and prove it helped?
  2. AnswerThe program has the right kinds of support, but it must test finished work and repeat use instead of counting sign-ups or classes.
  3. Deciding numbers$20,000 first waveest110–150 verified adoptersest$133–$182 per verified adopterest
  4. What the plan doesChange the current program to run the cross-campus first wave, pay peer helpers and part-time teachers, start at course setup, and defer the coach bot.
  5. Still unknownNo local test has run: the cost per verified adopter, the 30-day repeat rate, and adjunct parity are all estimates until the first wave measures them.

The sponsor's first worry is that professors won't use it; the university's standing constraint is money. So the adoption program in §07 was put through a second, adversarial review: a fresh reader audited it source by source against Influencer's six sources of influence, then went looking for cases across public health, hospitals, government, agriculture, schools, workplaces, consumer products, and higher education where mass adoption was achieved at low cost — and graded each one for evidence quality (A causal or systematic review, B strong but non-causal, C descriptive or vendor, D anecdote). The verdicts below changed the guide. Every claim keeps its grade: fact est unknown.

The short version

  • fact The current plan covers all six sources and is unusually strong on faculty choice, adjunct access, real tasks, and privacy.
  • fact Its weak point is proof: activation, training, artifact completion, repeat use, teaching use, and learning results are often treated as if they were the same outcome.
  • fact The strongest evidence favors action placed inside normal work, hands-on practice, local opinion leaders, fast help, and timely prompts.
  • fact Messages and social norms usually produce modest gains; mandatory checklists can reach near-universal reported compliance without changing outcomes.
  • est Change the first vital behavior from “finish a demo” to “apply, check, and use or reject one result within seven days.”
  • est Replace “repeat weekly” with “repeat at the next natural occurrence within 30 days.” Faculty work is often episodic.
  • est Make course-shell creation, not the first week of class, the main adoption moment.
  • est Add a peer behavior: each nominated catalyst supports several colleagues and shares both a useful result and a rejected result.
  • fact The lowest directly relevant faculty-AI payment with a measured completion result found here was $500; its causal effect is still unknown.
  • est Use a $20,000 first wave to learn UVU’s real cost per verified adopter, then scale only if the result, trust, and adjunct-parity gates pass.
  • est STOP making the custom coach bot a launch dependency. Start with three templates, existing Power Hours, and human help.
  • unknown No public evidence supports a promise of one-term majority faculty adoption at UVU for $0, $20,000, or $60,000.

The audit — source by source

Source of influenceWhat the plan doesWhat it assumesWhere the evidence is thin
Personal motivationplan Uses a real faculty task, discipline examples, skeptic stories, integrity framing, personal choice, and a no-AI route.est Immediate usefulness and a credible failure example will overcome risk, workload, and identity concerns.fact Kantar’s 85% daily-use claim is vendor evidence. unknown Ithaka does not establish that discipline workshops draw more faculty than generic workshops. The cited concern percentages do not prove that the proposed framing changes behavior.
Personal abilityplan Teaches a short draft-ground-check cycle, gives starter tasks, provides feedback, and proposes a coach bot.est A short session teaches transferable verification skills and the bot will be safe, accurate, and easier than current support.unknown The bot has no adoption precedent. Completing a practice artifact does not prove the faculty member used it in real work. Temple’s Canvas transition was hands-on but eventually mandatory and lacks a causal comparison.
Social motivationplan Uses confidential peer nomination, includes skeptics and adjuncts, publishes only truthful norms, and avoids selecting champions solely by enthusiasm.est The nomination questions find task-specific influence and visible participation will not feel coercive.fact Opinion-leader trials support the broad idea, with a median 10.8-point practice gain, but do not establish the best leader-selection method or its cost. Social-comparison messages have also produced null and boomerang effects.
Social abilityplan Uses micro-cohorts, Power Hours, adjunct-friendly formats, a shared artifact bank, and human escalation.est Existing staff and peers can supply fast, discipline-specific help across seven colleges.unknown There is no UVU capacity model for review, accessibility, curation, or one-business-day help. An adjunct-first program can become unpaid adjunct labor if participation happens outside compensated work.
Structural motivationplan Pays for a completed artifact, proposes prompt payment, recognizes useful contributions, and rejects login-based rewards.est A $500 award produces adoption and repeat use rather than only course completion.fact CSUB had 37 of 43 spring participants complete a $500 course, but there was no unpaid comparison. Hawai‘i awarded 50 $1,000 incentives but has not published a completion rate. OER programs show that faculty development work can greatly exceed the value of a small stipend.
Structural abilityplan Proposes Canvas placement, one-click sign-on, three starter tasks, safe defaults, local workflow saving, and help where faculty already work.est Canvas changes, identity integration, privacy review, accessibility work, and maintenance will have little incremental cost.unknown None of those UVU systems were inspected. Forced LMS rollouts do not prove voluntary AI adoption. A removable policy block is reasonable, but any default affecting teaching needs prior faculty-governance approval.

The vital behaviors — keep, change, add

  1. V1 — Modify.
  2. V2 — Tighten.
  3. V3 — Change the cadence.
  4. V4 — Add a peer diffusion behavior.

The moments that matter — re-timed

MomentJudgment
Course-shell creation or copy, 4–6 weeks before termest Make this the main moment. Faculty can still change assignments, the syllabus, and course structure. Place the policy block, three task cards, and support link here.
Adjunct contract and onboardingest Correct high-reach moment only if it happens before syllabus deadlines, is asynchronous, and is paid when outside normal duties. Late onboarding is a poor learning window.
7–14 days before the first graded assignmentest Better than the day the assignment opens. Prompt the assignment-level rule and one safe-use or non-use example.
First failed or uncertain attemptest Offer human help within the person’s work window. A bad first result can end adoption unless recovery is easy.
Next recurrence within 30 daysest This is the repeat-use test. The cue comes from the work, not from a platform streak.
First week of termest Use only for confirmation and student communication. It is too late and too busy to be the principal faculty-adoption moment.
Post-term/course-copy periodest Ask faculty to reuse, revise, or retire the workflow and optionally share the artifact.

The devil's-advocate case against the plan

  • est The plan may recruit the same 120 enthusiasts repeatedly while the overloaded middle remains untouched.
  • fact Its strongest examples usually prove seats, course completion, or activity—not sustained faculty practice or better student learning.
  • est A custom bot and full Canvas integration could turn a cheap behavior program into an expensive software project.
  • est “Voluntary” use can still feel mandatory when senior leaders, default course blocks, public norms, and stipends all point one way.
  • est Discipline-specific support is valuable but becomes costly if every department creates and maintains separate material.
  • unknown The plan has no tested method for moving from a small paid cohort to a majority of 2,112 instructors.
  • unknown Board-reported employee AI use rising from 61% to 76% has no published denominator, method, faculty-only result, or causal tie to this service.
  • fact The CSU wording needs correction: the public record says at least 250,000 activations by spring 2026, not “at most half.” The 0.7% figure is student training completion; the reported faculty figure was 16%. Neither is repeated use.

Cross-domain cases — what moved adoption, at what cost, with what evidence

Codes for the sources of influence: PM personal motivation · PA personal ability · SM social motivation · SA social ability · XM structural motivation · XA structural ability. The mappings are our reading, not claims made by the studies.

CaseDomainWhat moved adoptionSourcesCostResultGradeSource · date
Delancey StreetResidential rehabilitationfact Long immersion, resident-run work and education, shared responsibility, and peer teaching.est All sixunknown No complete economic cost or cost/graduate; resident labor and businesses are real inputs.fact Foundation reports more than 18,000 graduates; entrant count, attrition, comparison group, and independently reproducible outcome are absent.DDelancey Street, current institutional account
Guinea-worm campaignPublic healthfact Village volunteers, filters, water treatment, containment, surveillance, and reporting rewards.est All sixunknown Forty-year total and cost/adopter not published.fact Estimated human cases fell from 3.5 million in 1986 to 10 reported in 2025.B; bundled long-run trendWHO, 1986–2025
Thailand 100% Condom ProgramPublic healthfact Uniform establishment policy, free supply, peer coordination, STI services, media, monitoring, and sanctions.est All sixunknownfact Commercial-sex condom use rose from 14% in early 1989 to over 90% from 1992; surveillance also recorded a large STI decline.B; no causal isolationProgram review, 1989 onward
Geneva hand hygieneHospital safetyfact Bedside alcohol rub, posters, training, observation, feedback, and leadership.est PM, PA, SM, SA, XAfact Crude three-year cost under SFr380,000; cost/new compliant worker unknown.fact More than 20,000 observed opportunities; compliance rose 48%→66%, while infections fell 16.9%→9.9%.B; uncontrolled bundleLancet record, 2000
WHO surgical checklistHospital safetyfact Oral team pause, local adaptation, training, champions, and feedback.est PA, SM, SA, XM, XAunknown Original eight-site cost.fact Complications fell 11%→7% and death 1.5%→0.8% among 7,688 before/after patients.BNEJM, 2009
Ontario checklist rolloutHospital safety; important nullfact Required checklist use and reporting, without the same intensive team implementation.est PA, SM, XM, XAunknownfact Across 101 hospitals, adjusted death and complication changes were not significant despite very high reported checklist use.B, nullNEJM, 2014
UK tax lettersPublic administrationfact One truthful local social-norm sentence added to an existing collection letter.est SM, XAfact The text change had no reported added implementation cost; total program cost unknown.fact In a cleaner 1,400-person trial, payment rose 38.7%→45.5%.ABIT debt trials, 2011–12
UK organ-donor promptPublic administrationfact A short prompt appeared immediately after an online transaction.est PM, SM, XAunknownfact In 1,085,322 allocations, reciprocity increased registration from 2.3% to 3.1%; a norm-plus-photo version underperformed the control.A with allocation caveatTrial report, trial 2013; published 2018
Rajasthan immunizationVaccinationfact Reliable local camps plus small food incentives that offset travel and time.est All sixfact Corrected cost per fully immunized child: $27.94 with incentives versus $55.83 with reliable camps alone.fact Full immunization: 39% incentive, 18% reliable-camp-only, 6% control; 134 villages and 1,640 children.ABMJ, 2010; cost correction, 2016
Rutgers opt-out appointmentsVaccinationfact Employees received a prebooked vaccination appointment they could freely change or cancel.est XAunknown Incremental scheduling cost.fact Vaccination was 45% with opt-out scheduling versus 33% with opt-in scheduling among 478 employees.AJAMA, 2010
Local clinical opinion leadersProfessional practicefact Locally recognized clinicians educated or influenced peers.est PA, SM, SAunknown No included study reported cost-effectiveness.fact Twenty-four randomized studies; median adjusted practice-compliance gain 10.8 points across 18 studies.A, systematic reviewCochrane, 2019
Haryana “information spreader” nominationsVaccination/network diffusionfact Residents nominated people good at spreading information; nominees received calls and texts about camps.est SM, SA, XAunknownfact In 521 villages, nominated spreaders produced about 4.9 more vaccinated children per village-month than random seeds’ 18.11. General “trust” nominations were not clearly better.AReview of Economic Studies, 2019
Malawi lead farmersAgricultural extensionfact Two network-positioned farmers learned and demonstrated pit planting; multiple exposures mattered.est PA, SM, SA, XM, XAfact $8 in-kind gift per seed farmer; network census and downstream cost unknown.fact In 200 villages and roughly 5,600 households, non-seed adoption gained 3.6 points over a 3.8% benchmark in year two.AAmerican Economic Review, 2021
Rogers and “trigger the middle”Diffusion theoryfact Relative advantage, compatibility, trialability, visibility, networks, and organizational readiness help explain diffusion.est All sixunknownunknown No universal tipping percentage or rule says a fixed “middle” segment will trigger mass adoption. Field trials support task-specific spreaders and repeated exposure, not a magic adopter share.C as a frameworkRogers; health diffusion review
George Mason CanvasHigher-ed technologyfact Opt-in migration, 20 faculty mentors, course help, training, office hours, and later mandatory cutover.est PA, SM, SA, XM, XAfact Mentor stipend $2,000/semester; total cost unknown.fact 892 instructors and 19% of sections used Canvas in fall 2024; 1,502 instructors and 38% in spring 2025.BOfficial final report, 2025
AUT and Temple CanvasHigher-ed technologyfact Templates, champions, hands-on course building, drop-ins, communities, and migration help.est PA, SM, SA, XAunknownfact AUT moved 1,837 courses; Temple reported more than 1,000 pilot participants and later majority use, without a denominator. Both transitions were headed toward required institutional use.B/CAUT, 2023; Temple vendor case, rollout began 2017
Google Classroom in 2020School technologyfact No-charge core service embedded with existing Google tools during emergency remote schooling.est PM, PA, XAunknown Devices, support, and rollout cost.fact Google reported 50 million students and educators in March 2020, 100 million by January 2021, and 150 million by May 2021. No teacher-only or repeat-use denominator.CGoogle, Jan. 2021; Google, May 2021
Turnitin and clickersTeaching technologyfact Common systems, training, one-to-one help, question design, and peer discussion.est PA, SA, XAunknownfact One Turnitin case had 454 accounts but only 24 of 47 departments using it. At Colorado, 70 clicker faculty—3% of faculty—reached 44% of undergraduates through large courses.BTurnitin case, 2010–13; clicker study, 2007
OER programsHigher-ed teaching practicefact Release time, grants, library/design help, reusable resources, and no-cost course tags.est PA, SA, XM, XAfact One multi-college review valued average course development at about 180 hours and $12,600, versus an average $1,500 stipend or release award.fact Affordable Learning Georgia reports more than $143 million in student savings over 1.1 million enrollments; causal faculty-adoption effect remains unknown.BALG 2022 report; SRI OER evaluation, 2022
GitHub Copilot at AccentureDeveloper toolsfact Suggestions appeared inside the IDE, with almost immediate first value.est PM, PA, XAunknownfact GitHub reported 81.4% same-day installation and 96% same-day first acceptance among installers; one acceptance is activation, not useful adoption.CGitHub/Accenture, 2024
Embedded customer-service AIWorkplace AIfact Three-hour onboarding, suggestions inside live chats, scheduled access, and ordinary coaching.est PA, SA, XAunknownfact In 5,172 agents and more than 3 million chats, access raised resolved issues/hour by 15%; effects were much larger for less-skilled workers.A−, staggered rolloutQuarterly Journal of Economics, 2025
UK government CopilotWorkplace AIfact Office integration, short training, workshops, central resources, and department support.est PA, SA, XAunknownfact Of 20,000 licenses, “active” meant one interaction in 30 days; adoption reached 83% and remained near 80%. Application use varied widely and some use fell after its peak.BGOV.UK report, 2025
Slack and enterprise InnerSourceWorkplace/open sourcefact Free trialability, familiar workflow placement, integrations, open backlogs, contribution guides, maintainers, and peer help.est PA, SM, SA, XAunknown Free licenses did not remove support, maintenance, or contribution costs.fact Slack reported more than 10 million daily users in 2019. Ericsson reported over 1,000 internal reuse instances, but also few early contributions where work time was not funded.B/CSlack S-1, 2019; Ericsson InnerSource study, 2024
Duolingo, Strava, PelotonConsumer habitsfact Private progress, reminders, finite challenges, instructors, streaks, and social cues.est PM, SM, SA, XM, XAunknown Cost/adopter and causal contribution of each feature.fact These services report high repeat activity, but their users are self-selected and most evidence does not show that streaks or rankings caused retention. A student RCT found personalized reminders improved first use more than streak messages.A for reminder experiment; B/C productsReminder RCT, 2026; Duolingo SEC filing; Peloton filing
California State UniversityHigher-ed AIfact Systemwide licenses, training, local projects, and $3 million for faculty proposals.est PA, SA, XM, XAfact Initial contract $17 million/18 months; reported renewal $13 million/year.fact More than 93,000 activations by June 2025 and at least 250,000 by spring 2026. Voluntary training completion was reported as 0.7% of students and 16% of faculty. Activation and training are different measures.B/CCSU launch, 2025; CalMatters, 2026
CSUB and Hawai‘iHigher-ed AI incentivesfact Payment was attached to bounded training or an implemented, shared assignment.est PM, PA, SA, XMfact CSUB: $500/completer. Hawai‘i: 50 awards of $1,000; actual payout unknown.fact CSUB spring 2026 completion was 37/43, with a waitlist. Hawai‘i has no published cohort completion rate. Neither isolates the incentive’s effect.B/CCSUB Senate report, 2026; Hawai‘i program, 2025
Virginia Tech, Manchester, CedarvilleHigher-ed AI usefact Training or supported cohorts plus institutionally approved access.est PA, SA, XAunknown Comparable total cost/adopter.fact Virginia Tech recorded 78% typical weekly activity among 425 measured pilot users. Manchester reported 90% 30-day adoption but did not publish its exact definition. Cedarville reported 82% activation and average weekly activity among 75% of active accounts.B/B/CVirginia Tech, 2025; Manchester, 2026; Cedarville, 2026
ASU, Michigan, ArizonaHigher-ed AIfact Proposal challenges, locally built tools, Canvas placement, workshops, and course-specific assistants.est PM, PA, SM, SA, XAunknown Comparable faculty cost/adopter.fact ASU reported more than 500 completed or active projects; Michigan reported 43,800 U-M GPT users and more than 500 courses; Arizona’s AI-VERDE pilot recorded 78 users and 97,658 calls. Only Michigan published an adjacent historical course comparison.B/CASU, 2025; Michigan, 2025; Arizona paper, 2025
Miami Dade, Florida, UT AustinHigher-ed AIfact Faculty training, institution-wide curriculum, course tools, vetted activities, and local AI platforms.est All six in varying mixesfact Miami Dade later received a $2 million expansion grant; UF disclosed an $800,000 recurring annual QEP budget and charges $500 for external Academy enrollment; UT cost is unknown.fact Miami Dade reports 1,056 faculty trained but its outcome claims lack denominators. UF reports 200+ AI courses and 300+ AI-focused faculty. UT Sage reported 366 tutors in 90 Canvas courses; Copilot training reached 2,400+ users.B/CMiami Dade; UF; UT Austin, 2024–26
Ivy Tech and MiddleburyHigher-ed AIfact Ivy Tech used workshops, departmental guides, and a 60-user pilot. Middlebury provides faculty choice, syllabus templates, workshops, consultations, and $1,000 AI mini-grants.est PM, PA, SA, XM, XAfact Middlebury mini-grants up to $1,000; other adoption costs unknown.unknown Ivy Tech published no activity or outcome result. Middlebury has useful student studies but no measured faculty-rollout denominator or cost/adopter.CIvy Tech account, 2024; Middlebury resources, 2025–26

The low-cost design for a constrained university

Behaviors

  1. est Apply one checked result. Finish a real, bounded task and use, edit, or reject the result within seven days.
  2. est Publish one clear course rule. Add an assignment-level permission, disclosure example, and checking rule before students begin.
  3. est Repeat at the next real occurrence. Reuse or revise the workflow within 30 days.
  4. est Help peers perform the behavior. Each catalyst supports three to five colleagues and shares a worked success plus a failure.
  5. est Retire what does not help. A faculty member may record “not useful,” “not safe,” or “not suitable” without being counted as resistant.

Moments

  • est Primary: course-shell copy, 4–6 weeks before term.
  • est Equity gate: adjunct contract/onboarding, with paid asynchronous completion.
  • est Teaching gate: 7–14 days before the first graded assignment.
  • est Recovery: immediately after the first failed or uncertain result.
  • est Repeat: the next occurrence of the same task, within 30 days.
  • est Reuse: post-term course-copy and artifact revision.
  • est Avoid launching tools during finals or using the first week as the main training period.

People — finding the real spreaders for $0

  • est Add two questions to an existing department, senate, or teaching-center pulse: Who spreads useful teaching practices? Who do you consult when a teaching tool may not be ready?
  • fact Haryana’s trial suggests “information spreader” nominations may be more useful than broad trust or title.
  • est Nominate at least two people per teaching cluster because complex behaviors may require more than one credible exposure.
  • est Include adjuncts, skeptics, high-enrollment instructors, and people with strong peer connections—not only public AI enthusiasts.
  • est Use the 120 Academy completers as a candidate pool, not as an automatic champion list.
  • est This can have $0 incremental cash cost when inserted into existing processes. unknown Staff administration and faculty time still have economic cost.

Structures that cost nothing

  • est Add a removable Canvas block when a new or copied shell is created: active choice among permitted, limited, or prohibited use; one example; one support link.
  • est Put three discipline task cards beside the normal work: course preparation, assessment/rubric work, and student feedback.
  • est Default visibility on but keep tool launch, Academy enrollment, reminders, telemetry, and public sharing opt-in.
  • est Reuse Canvas Commons, existing Power Hours, approved sign-on, and existing department meetings before building a new portal.
  • est Keep a small reviewed artifact library. Adapt a shared pattern rather than creating a separate program for every discipline.
  • est Put human help beside the task. Do not require a faculty member to leave Canvas, search a catalog, or create another account.
  • est Do not build the coach bot until support logs show a repeated problem that templates and people cannot solve cheaply.

Incentives — what the evidence supports

  • fact CSUB’s $500-on-completion course produced 37 completions among 43 enrolled participants, but there was no unpaid control.
  • fact Hawai‘i offered $1,000 for building, using, and sharing an assignment, but published no overall completion rate.
  • fact George Mason paid Canvas mentors $2,000 per semester; its total adoption cost was not reported.
  • unknown The minimum effective faculty-AI stipend is not known. No credible public test here shows that $100, $200, recognition, or early access alone changes sustained faculty behavior.
  • est Pay catalysts for a reusable artifact plus peer help, not for attendance or message volume.
  • est Use completion awards when the work is outside ordinary duties. Recognition, early access, and voluntary public commitment can supplement pay but must not replace compensation for adjunct labor.
  • est Avoid leaderboards, streaks, prizes for frequent use, and public lists of “non-adopters.”

Measures that don't surveil

A verified adopter is a faculty member who completes a checked real-work cycle, applies or edits or rejects the result within seven days, publishes one course-use rule (or completes another approved applied workflow), and repeats a related workflow within 30 days. The funnel is measured separately:

MeasurePrivacy-preserving method
Reachest Aggregate invitations, Canvas card views, and clinic seats.
First applied useest Content-free completion receipt plus apply/edit/reject choice.
Course adoptionest Voluntarily submitted policy or assignment artifact count; do not scrape course content.
Repeat useest Opt-in random evaluation token with no SSO/HR link and short retention. If governance rejects any longitudinal token, report repeat use UNKNOWN.
Peer diffusionest Aggregate unique clinic participants and voluntarily reported referral source.
Qualityest Small faculty review sample using a published rubric; never feed results into evaluation or rehire.
Equityest PT/FT and college participation/completion only in cells of at least 20.
Trustest Anonymous usefulness, pressure, autonomy, and surveillance-concern pulse.
Safetyest Aggregate privacy, access, accuracy, and accessibility incidents.
Economicsest Cash spend and estimated staff/faculty hours divided by verified adopters; publish both.

Cost per verified adopter at three budgets

With 2,112 instructors, a strict majority is 1,057. These are planning-capacity estimates, not adoption forecasts; they assume the AI service itself is already funded, and faculty time is not free.

Cash budgetProposed useSupported capacityVerified-adopter assumptionCash cost per verified adopterWhat it can honestly prove
$0est Ask each of the 120 existing Academy completers to help one new colleague during already-paid work; reuse Power Hours and templates.est 120 new peer-support places.est 60–84 new adopters, using a 50–70% planning completion range.est $0 cash; full economic CPA UNKNOWN.est Whether existing social capacity produces a measurable second wave. It is not a credible one-term majority plan.
~$20kest Ten catalysts × $1,000 for an artifact and two clinics; twenty adjunct/faculty artifact awards × $500.est Ten catalysts plus 200 unique clinic places.est 110–150 verified adopters: ten catalysts plus 50–70% of clinic places.est $133–$182.est A cross-campus first wave large enough to compare colleges, PT/FT participation, repeat use, and local CPA.
~$60kest Thirty catalysts × $1,000; sixty artifact awards × $500.est Thirty catalysts plus 600 unique clinic places.est 330–450 verified adopters.est $133–$182.est A material campus wave—about 16–21% of instructors—but not a one-term majority.

At the modeled $133–$182 per verified adopter, reaching a majority directly would take roughly $141,000–$192,000; peer diffusion could lower later waves, but that has to be measured, not assumed. Recommendation carried into the guide: run the $20,000 first wave, and continue to the $60,000 design only if cost per verified adopter stays at or below $200, 30-day repeat is credible, adjunct completion is not materially lower than full-time, reported pressure or surveillance concern is not rising, and artifacts pass faculty quality review. est

Ranked changes — what the guide now does differently

#ChangeEvidence and reasonIncremental cash
1Modify V1 to require applied use or reasoned rejection within seven days.fact Checklist and account cases show that formal completion can exist without faithful practice or useful results.est $0
2STOP making the custom coach bot a launch dependency.unknown No precedent here shows that a bot improves faculty adoption. fact Michigan’s locally built tools show possible scale, but their institutional cost is unknown.est Saves or defers build cost
3Replace weekly repetition with the next natural recurrence within 30 days.fact Embedded workplace tools work at the task; faculty work is not uniformly daily or weekly.est $0
4Move the main trigger to course-shell copy and assignment construction.fact BIT, vaccination, and teacher-nudge evidence supports prompts at a natural action point.est $0 cash if the Canvas template path already exists; staff time unknown
5Add the catalyst peer behavior and task-specific nomination questions.fact Opinion-leader review: median 10.8-point practice gain. Haryana: information-spreader nominations outperformed random seeds.est $0 for nomination; paid catalyst budget as selected
6Use the $20,000 budget for catalysts, adjunct artifacts, and clinics—not 30 isolated $500 completions.est The same cash creates more supported opportunities and reusable capacity.est Budget-neutral reallocation
7Separate policy clarity from AI adoption.est A clear prohibition can satisfy V2’s current wording without any adoption. Record clarity, safe AI use, and justified non-use separately.est $0
8Run a stepped or randomized invitation test.fact The strongest cheap-adoption evidence comes from randomized rollout. est Compare early versus later invitation at department or course-cluster level.est Near-zero cash if built into rollout
9Resolve the telemetry contradiction.plan The bot is supposed to beat self-service on 30-day repeat, while the current hard rule forbids durable identification. est Use opt-in, pseudonymous, short-retention evaluation tokens—or report repeat use UNKNOWN.unknown Governance and evaluation time
10Correct the evidence claims before leadership use.fact CSU’s “at most half” is unsupported; its 0.7% number concerns students, not faculty. Hawai‘i announced awards, not verified payouts. Manchester did not prove cohort support caused its 90% figure.est $0
11Make adjunct compensation a launch gate.fact OER evidence shows major hidden faculty workload. est Do not call unpaid extra work “free adoption.”est Included in the $20k/$60k designs
12Retire nudges that show no incremental effect.fact Large education and social-comparison nudge trials show that low-cost messages can be null or backfire.est Saves staff attention

Risks specific to influence programs — and how the cases handled them

RiskHow it failsWhat the cases showUVU control
Manipulation perceptionDefaults, norms, and “trusted peers” can look like concealed pressure.fact The UK organ-donor norm/photo prompt underperformed control.est Disclose the program owner, purpose, options, and measure. Keep refusal easy. Ask anonymously whether people felt pressured.
Faculty autonomy and senate reactionCentral purchasing or template defaults may be read as a teaching mandate.fact CSU’s rollout produced reported consultation and governance objections.est Faculty Senate approves the choice architecture before launch. Faculty actively choose permitted, limited, or prohibited use.
Union concernsNew required work may be added without workload recognition or bargaining.fact Large OER projects found substantial hidden development time.est Distinguish voluntary exploration from required work. Bargain or approve workload terms before requiring training or artifacts.
Adjunct exploitationEvening clinics and artifact production become uncompensated labor tied to perceived rehire risk.fact Small stipends often cover only a fraction of course-development work.est State pay, time estimate, ownership terms, and no-rehire consequence before enrollment. Offer asynchronous paid routes.
Surveillance fearsUsage logs become a list of “resistant” faculty or enter evaluation.fact Workplace and university platforms can expose named activity even when only aggregate reporting is promised.est Separate evaluation from SSO, HR, tenure, discipline, and rehire. Store no prompts or outputs. Suppress cells under 20.
AI-mandate backlashAccess, default visibility, repeated messages, and public norms combine into a soft mandate.fact Ontario showed that mandatory reported compliance did not guarantee outcomes.est Default visibility only. Keep participation, reminders, telemetry, testimonials, and public sharing opt-in.
Norm boomerangLow participation normalizes non-use; high-user comparisons discourage or shame others.fact Teacher social-comparison evidence found no overall gain and possible downward movement among above-average groups.est Use a norm only when true, local, recent, clearly defined, and helpful. Otherwise omit it.
Incentive gaming or crowd-outParticipants optimize for a certificate or payout rather than useful practice.fact CSUB proves completion, not repeat use.est Pay for a reviewed artifact and bounded peer contribution; measure later repeat separately. Do not pay for prompts or logins.
Champion burnout or elite captureThe same visible enthusiasts receive every role and central support becomes a bottleneck.fact InnerSource evidence reports bottlenecks and weak contribution when work time is not funded.est Cap cohort size, pay defined work, rotate roles, include skeptics, and publish help capacity.
Accuracy, privacy, and safetyA smooth first experience can normalize unchecked or sensitive use.fact UK Copilot users reported limits with nuanced, complex, and sensitive work.est Require data classification, verification, and a safe exit. Keep clinical and other high-risk tasks in separate reviewed lanes.
Gamification cringeStreaks, badges, and rankings feel infantilizing or punitive.fact Consumer products show repeat activity, but seldom prove that public rankings caused it.est Transfer only private progress, self-chosen reminders, finite tasks, and a lapse-recovery path. No public streaks or leaderboards.
Equity and accessibilityFull-time faculty capture the support while adjuncts and disabled faculty face time or access barriers.fact Workplace AI reports accessibility benefits, but also learning curves and role differences.est Test accessibility before launch; provide asynchronous and human routes; report PT/FT parity only in protected aggregate cells.

Review: working paper 37, September 3, 2026 — sources dated and graded; nothing here contacted UVU faculty, inspected UVU systems, or accessed person-level data. The three budget ranges are planning assumptions; only a local staged test turns them into evidence.

O

The two Mac lineups, side by side

Every configuration priced on the newest lineup and on the generation it replaced — what each cost at launch, what the old one costs today (if you can get it), and how much faster the new one should be

Appendix O in five lines

  1. QuestionShould the plan buy the new Mac computers or the older line?
  2. AnswerBuy the new line by default because matched old stock is not proved and its lower price does not fully pay for its slower speed.
  3. Deciding numbersSeptember 22, 2026 availabilityfact14.4% more costest10% more total answer-writing speedest
  4. What the plan doesPrice the new computers, test four after September 22, and buy old ones only with a written quote for the exact set and date.
  5. Still unknownNew desktop tests, university prices, stock for 4 or 26 matching old units, and exact warranty prices are NOT_RUN.

A fair question from the first read-through: the guide prices the machines on Apple's newest lineup, announced August 25, 2026 — so what did the generation before it cost, what does it cost today, and how much faster should the new one really be? This appendix answers that with dated sources, every number graded. Research cutoff September 3, 2026; prices are US dollars before tax. fact means a dated source states or displays it; est means we derived it and show the basis; unknown means the proof was not available.

What this changes in the plan

1. The plan stays priced on the newest lineup. The previous machines are no longer sold new — Apple's own store and its education store have dropped them, and the big education resellers show backorders. Single refurbished units exist at prices near or above what they cost new, and nobody can prove a matched set of 4 or 26 is available until it is reserved. 2. One price corrected. The 256GB Ultra needs Apple's top chip, so it is about $10,799 retail (about $9,900 education, an estimate), not the $9,499 an earlier pass carried; it only affects the later, optional heavyweight node, and the simulator and configurator now use the corrected figure. 3. Priced both ways everywhere. The budget simulator has a generation switch and the configurator shows the previous generation beside every configuration you build, using the refurbished prices below and speed credits deliberately below Apple's claims (+10% per box for the M5 Pro and M5 Max, +40% for the M5 Ultra). 4. An acceptance test before scale. A four-unit run after September 22 must re-prove the speed credit before any purchase beyond the validation. 5. Old units only against paper. Procurement may take previous-generation units only with a written quote for the exact quantity, drive size, and delivery date (Apple Education/NASPO Utah, CDW-G, or SHI). est

The bottom line

  • fact Apple is taking orders for the new Macs; normal availability starts September 22, 2026. The 512GB M5 Ultra is due in late October.
  • unknown No new Mac mini or Mac Studio serving benchmark can exist before delivery. Laptop M5 Pro/Max results are same-chip proxies only.
  • est For large-model decode, budget +10% throughput for M5 Pro/Max and +40% for M5 Ultra. These credits are intentionally below Apple’s headline AI claims.
  • fact Previous models are absent from Apple’s normal retail and consumer education catalogs.
  • fact Apple Refurbished currently shows individual previous-generation units, but unknown whether 4 or 26 matched units can be reserved.
  • fact CDW-G lists prior units as backordered; SHI reports zero stock/backorder. Current institutional fulfillment is unknown.
  • A 48GB fleet priced from Apple education costs est 14.4% more than the visible M4 Pro refurb alternative and is expected to provide 10% more aggregate decode.
  • Recommendation: price the plan on the newest generation, require a four-node acceptance test after September 22, and treat old inventory only as a quote-backed fallback.

The newest lineup — every configuration, priced

Apple’s Mac mini announcement and Mac Studio announcement give the starting prices and dates. Memory and networking come from the current Mac mini specifications and Mac Studio specifications.

Minimum storage is 256GB for M6, 512GB for M5 Pro/Max, and 1TB for M5 Ultra unless noted.

FamilyCPU/GPU tierMemoryBandwidthRetail priceEducation-store priceAvailability10GbEAppleCare+ premium
Mac mini M612/1216GBfact 153 GB/sfact $899fact $799fact 2026-09-22fact +$100 retail; +$90 eduunknown
Mac mini M612/1224GBfact 170 GB/sfact $1,099unknownfact 2026-09-22fact +$100; +$90unknown
Mac mini M612/1232GBfact 170 GB/sest $1,299 = $899 + $400 memoryunknownfact 2026-09-22fact +$100; +$90unknown
Mac mini M5 Pro15/1624GBfact 307 GB/sfact $1,699fact $1,599fact 2026-09-22fact +$100; +$90unknown
Mac mini M5 Pro15/1648GBfact 307 GB/sest $2,299 = $1,699 + $600est $2,139 = $1,599 + $540fact 2026-09-22fact +$100; +$90unknown
Mac mini M5 Pro15/1664GBfact 307 GB/sest $2,699 = $1,699 + $1,000est $2,499 = $1,599 + $900fact 2026-09-22fact +$100; +$90unknown
Mac mini M5 Pro18/2024GBfact 307 GB/sest $1,899 = $1,699 + $200 chipest $1,779 = $1,599 + $180fact 2026-09-22fact +$100; +$90unknown
Mac mini M5 Pro18/2048GBfact 307 GB/sest $2,499est $2,319fact 2026-09-22fact +$100; +$90unknown
Mac mini M5 Pro18/2064GBfact 307 GB/sest $2,899est $2,679fact 2026-09-22fact +$100; +$90unknown
Mac Studio M5 Max18/3236GBfact 460 GB/sfact $2,499fact $2,299fact 2026-09-22fact includedunknown
Mac Studio M5 Max18/4048GBfact 614 GB/sest $3,099est $2,839fact 2026-09-22fact includedunknown
Mac Studio M5 Max18/4064GBfact 614 GB/sest $3,499est $3,199fact 2026-09-22fact includedunknown
Mac Studio M5 Max18/40128GBfact 614 GB/sest $5,099 at 512GB SSD; fact $5,399 at 1TBest $4,639 at 512GB SSDfact 2026-09-22fact includedunknown
Mac Studio M5 Ultra30/6496GBfact 1.2 TB/sfact $5,499fact $5,099fact 2026-09-22fact includedunknown
Mac Studio M5 Ultra30/64256GBfact not configurableN/AN/AN/Afact includedN/A
Mac Studio M5 Ultra30/64512GBfact not configurableN/AN/AN/Afact includedN/A
Mac Studio M5 Ultra36/8096GBfact 1.2 TB/sest $6,799 = $5,499 + $1,300 chipunknownfact 2026-09-22fact includedunknown
Mac Studio M5 Ultra36/80256GBfact 1.2 TB/sest $10,799 at 1TB; fact $11,299 at 2TBunknownfact 2026-09-22fact includedunknown
Mac Studio M5 Ultra36/80512GBfact 1.2 TB/sunknownunknownfact late October 2026fact includedunknown

Price notes:

  • est values add Apple’s displayed option deltas to a sourced base price. They are not institutional quotes.
  • AppleCare+ eligibility is displayed, but an exact new-desktop premium was not exposed in the retrievable store pages. Apple says education buyers may save up to 10%; the actual premium remains unknown.
  • Configured-to-order delivery can be later than the family’s September 22 availability date.

The previous lineup — what it cost, and whether it can still be bought

Launch facts come from Apple’s 2024 Mac mini announcement and 2025 Mac Studio announcement.

Previous familyChip tier(s)MemoryBandwidthLaunch price: retail / educationApple retail and consumer edu todayApple Refurbished seen 2026-09-03Institutional/reseller state
Mac mini M410-core CPU/GPU16GBfact 120 GB/sfact $599 / $499fact not listedfact $679, 16GB/256GBOlder Apple institution list: fact $899 for 16GB/512GB; present fulfillment unknown
Mac mini M410-core CPU/GPU24GBfact 120 GB/sunknown exact CTO; family starts $599/$499fact not listedfact $1,019, 24GB/512GBOlder institution list: fact $1,099 for 24GB/512GB; fulfillment unknown
Mac mini M410-core CPU/GPU32GBfact 120 GB/sunknown exact CTOfact not listedfact $1,269, 32GB/512GB/10GbECurrent large-quantity price unknown
Mac mini M4 Pro12/16 or 14/2024GBfact 273 GB/sBase 12/16: fact $1,399 / $1,299; top-tier exact unknownfact not listedfact $2,889 with 4TB; CDW-G base listing $1,599 backorderedOlder institution list: fact $1,499 base; current fulfillment unknown
Mac mini M4 Pro12/16 or 14/2048GBfact 273 GB/sunknown exact CTOfact not listedfact $1,949, 12/16, 48GB/512GB/10GbENo four- or 26-unit stock proof
Mac mini M4 Pro12/16 or 14/2064GBfact 273 GB/sunknown exact CTOfact not listedfact $2,379, 14/20, 64GB/512GB/1GbENo four- or 26-unit stock proof
Mac Studio M4 Max14/3236GBfact 410 GB/sfact $1,999 / $1,799fact not listedfact $1,949, 36GB/512GBCDW-G: fact $2,499, backordered; older institution list $2,299; fulfillment unknown
Mac Studio M4 Max16/4048GBfact 546 GB/sunknown exact CTOfact not listedfact $2,459, 48GB/512GBSHI stock zero/backorder; quote unknown
Mac Studio M4 Max16/4064GBfact 546 GB/sunknown exact CTOfact not listedfact $6,029 with 8TB SSDQuantity and useful-storage price unknown
Mac Studio M4 Max16/40128GBfact 546 GB/sunknown exact CTOfact not listedfact $4,159, 128GB/512GBQuantity unknown
Mac Studio M3 Ultra28/6096GBfact 819 GB/sfact $3,999 retail; fact $3,599 institution educationfact not listedfact $6,029 with 2TBCDW-G: fact $5,299, backordered; older institution list $4,899
Mac Studio M3 Ultra32/8096GBfact 819 GB/sest $5,499 retail = $3,999 + $1,500 chip; edu unknownfact not listedExact current match unknownFulfillment unknown
Mac Studio M3 Ultra28/60 or 32/80256GBfact 819 GB/sTop tier: est $7,099 = $3,999 + $1,500 chip + $1,600 memory; edu unknownfact not listedfact $8,149, 28/60, 256GB/2TBCDW-G 256GB listing: fact $11,639.99, backordered
Mac Studio M3 Ultra32/80512GBfact 819 GB/sest $9,499 = $3,999 + $1,500 chip + $4,000 memory; edu unknownfact not listedfact $17,419, 512GB/8TBQuantity unknown

University purchasing status

ChannelCurrent resultUsable price bandUniversity order conclusion
Apple retailfact previous models not listedN/ANew prior-generation units unavailable through the ordinary public catalog
Apple consumer education storefact previous models not listedN/ANo public education checkout path found
Apple Certified Refurbishedfact individual Add to Bag listingsServing-relevant examples: $1,949–$17,419A single unit may be purchasable; matched quantity is unknown until reserved
Apple institution/NASPO documentsfact old SKUs remain in pre-refresh documentsExamples: $899–$4,899Documents prove catalog pricing, not September acceptance, stock, or delivery
CDW-Gfact public listings; backordered$1,599–$11,639.99Order entry may exist; delivery date and UVU contract price are unknown
SHIfact search results show backorder/zero stockConflicting public pricesRequire a written UVU quote; current fulfillment is unknown
Apple Education/NASPO Utahfact Utah participation path existsContract-specificPrevious-generation availability and UVU price require a current quote

Apple warns refurbished supply is limited. An “Add to Bag” page is not evidence that 4 or 26 matched computers can be delivered.

Speed evidence, and the credits we use for planning

Apple's own claims — the ceiling, not the plan

These are selective “up to” claims, not guaranteed LLM serving gains. “Peak GPU AI” refers to GPU Neural Accelerators; it is not a Neural Engine percentage.

UpgradeCPU claimGPU claimNeural Engine claimBandwidth changeApple LM Studio claim
M6 vs M4fact “up to 40 percent faster CPU performance”fact “up to 2x faster graphics”fact “up to 2x faster Neural Engine”fact 120 → 153/170 GB/s; +27.5%/+41.7%fact up to 4.8x prompt processing
M5 Pro vs M4 Profact up to 30%fact up to 20%; ray tracing up to 35%; peak GPU AI over 4xfact Apple says faster; numeric percentage unknownfact 273 → 307 GB/s; +12.45%fact up to 4x prompt processing
M5 Max vs M4 Maxfact up to 15%fact up to 20%; ray tracing up to 30%; selected Studio GPU workload up to 50%Numeric percentage unknownfact 546 → 614 GB/s; +12.45%fact up to 3.9x prompt processing
M5 Ultra vs M3 Ultrafact multi-core up to 30%; single-core up to 25%fact selected graphics up to 1.8x; generic chip comparison up to 40%; peak GPU AI 4.3–4.5xNumeric percentage unknownfact 819 → 1,200 GB/s; +46.52%fact up to 4x prompt processing

Sources: Apple’s M6/M5 Ultra chip announcement, M5 Pro/Max announcement, and the two product announcements above.

Independent measurements

PP is prompt processing or prefill. TG is token generation. Results are only comparable when model, quantization, software build, prompt, and batching match.

Hardware and testModel/stackPrefillSingle-stream generationConcurrent aggregateEvidence state
M4 Pro 16-GPU → M5 Pro 16-GPU laptop; same llama.cpp build 8e672efSame Q4 llama-bench row364.06 → 403.19 t/s; +10.75%49.64 → 60.04 t/s; +20.95%Not reportedfact same-build chip proxy
M4 Max 40-GPU → M5 Max 40-GPU laptop; same buildSame Q4 llama-bench row885.68 → 990.53; +11.84%83.06 → 104.03; +25.25%Not reportedfact same-build chip proxy
M4 Pro mini 64GBgpt-oss-20B MXFP4, llama.cpp build 6280700.78 @2K; 618.59 @8K; 534.95 @16K; 419.74 @32K63.34Not reportedfact
M4 Pro mini 64GBQwen3.8-27B-oQ4e, oMLX 0.6.x, MTPGeneration test at 1K/4K/8K contexts24.9 / 20.5 / 20.8120 t/s at 8-widefact field report
M4 Max 128GBQwen3.5-35B-A3B 4-bit, oMLX1,216122.2205.3 batch-2; 283.6 batch-4fact public benchmark
M3 Ultra 256GB, 60-GPUgpt-oss-120B oQ8, oMLX868.168.094.3 / 129.3 / 152.5 at batch 2/4/8fact public benchmark
M3 Ultra 512GB, 60-GPUgpt-oss-120B MXFP4, oMLX (verified result)600.7 @32k68.4 @1k · 57.8 @8k · 35.2 @32k187.1 at batch 8 (short context)fact public benchmark — corrected 2026-09-03: the earlier row (1,405.9 / 79.2) cited a Qwen3.5-122B-A10B result
M3 Ultra 512GBQwen3.8-27B-oQ4e, oMLX 0.6.1, MTP420–44265–75 at 8K259–274 at 8-widefact four-node field report
New Mac mini/Studio chassisAny serving stackunknownunknownunknownNot deliverable until 2026-09-22+

The same-build llama.cpp evidence is in the project’s Apple Silicon benchmark discussion. The larger-model M4 Pro result is in the gpt-oss llama.cpp guide. The Qwen field comparison is documented in oMLX discussion #2811.

The planning credits we use

Large-model decode is commonly limited by memory bandwidth. Prefill uses more compute and can benefit more from CPU/GPU changes and software-specific accelerators.

New tierDecode expectationPlanning creditPrefill expectationConcurrent serving expectationReason
M6 16GBest +20–30%est +25%est +20–80%; absolute speed unknownAggregate capacity est +25%; actual streams unknownBandwidth rises 27.5%; Apple’s 4.8x LM Studio claim is not treated as general serving speed
M6 24/32GBest +30–40%est +35%est +20–100%; absolute speed unknownAggregate capacity est +35%; actual streams unknownBandwidth rises 41.7%; more RAM also permits larger models/KV caches
M5 Proest +10–15%est +10%est +10–20%Aggregate capacity est +10%Bandwidth +12.45%; same-build small-model PP +10.75%, TG +20.95%
M5 Maxest +10–15%est +10%est +10–20%Aggregate capacity est +10%Bandwidth +12.45%; same-build PP +11.84%, TG +25.25%
M5 Ultraest +35–50%est +40%est +25–50%Aggregate capacity est +40%Bandwidth +46.52%; larger Apple GPU-AI claims are not assumed to transfer to ordinary decode

“Concurrent capacity” means aggregate token throughput. It does not promise an equal percentage increase in integer sessions because scheduling, context length, KV cache, and latency targets still matter.

Side by side, role by role

Serving rolePrevious generation: model, memory, bandwidth, launch/current price, availability, measured servingNewest generation: model, memory, bandwidth, price, date, expected servingPrice deltaPerformance deltaUniversity choice
Workhorse mini 48GBM4 Pro 12/16, 48GB, 273 GB/s. Launch exact unknown; family base $1,399/$1,299. Current Apple refurb: fact $1,949, 512GB SSD and 10GbE; quantity unknown. M4 Pro 64GB proxy: Qwen3.8 20.8 t/s @8K, 120 aggregate @8-wide.M5 Pro 15/16, 48GB, 307 GB/s. est $2,299/$2,139; with 10GbE $2,399/$2,229. Ships 9/22. est 22.9 t/s, 132 aggregate @8-wide-equivalent; actual chassis unknown.Versus refurb with 10GbE: edu +$280, +14.4%; retail +$450, +23.1%est +10%Buy new by default. It gives supported supply, warranty, and a modest decode gain. Use old only with a written matched-quantity quote.
64GB miniM4 Pro 14/20, 64GB, 273 GB/s. Launch exact unknown. Current refurb: fact $2,379, 512GB, only 1GbE. Measured 64GB mini: 20.8 t/s @8K, 120 batch-8.M5 Pro 15/16, 64GB, 307 GB/s. est $2,699/$2,499; 10GbE adds $100/$90. Ships 9/22. est 22.9 t/s, 132 batch-8-equivalent.Without network parity: edu +5.0%; retail +13.5%. With new 10GbE: +8.8%/+17.7%, but old remains 1GbE.est +10% decode; prefill est +10–20%Buy new. Choose the base CPU/GPU tier for bandwidth-bound decode; pay for 18/20 only when prefill testing proves value.
Heavyweight Studio 128GBM4 Max 16/40, 128GB, 546 GB/s. Launch CTO unknown. Current refurb fact $4,159. Qwen3.5-35B: 122.2 TG, 283.6 aggregate batch-4; same-build llama row 83.06 TG.M5 Max 18/40, 128GB, 614 GB/s. est $5,099/$4,639 at 512GB SSD; ships 9/22. Planning proxy: est 134.4 TG, 312.0 batch-4 aggregate; chassis result unknown.Edu vs refurb: +$480, +11.5%; retail: +$940, +22.6%Fleet credit est +10%; small-model same-build evidence showed +25.25% TGBuy new for planned deployments. Old refurb is attractive only if exact quantity and storage are confirmed.
256GB UltraM3 Ultra up to 32/80, 256GB, 819 GB/s. Top-tier launch est $7,099, edu unknown. Current refurb: $8,149 for lower 28/60 tier with 2TB. gpt-oss-120B on 60-GPU: 868.1 PP, 68 TG, 152.5 batch-8.M5 Ultra 36/80, 256GB, 1.2 TB/s. est $10,799 at 1TB; fact $11,299 at 2TB; edu unknown. Ships 9/22. est 95.2 TG, 213.5 batch-8 aggregate.Like-for-like 2TB new vs lower-tier refurb: +38.7%. Top-tier launch comparison: +52.1%est +40%Buy new only when models exceed 128GB. The bandwidth gain is material, but obtain an institutional price and validate the exact model first.
512GB UltraM3 Ultra 32/80, 512GB, 819 GB/s. Launch est $9,499; current 8TB refurb fact $17,419. Qwen3.8: 65–75 TG, 259–274 batch-8; gpt-oss-120B: 68.4 TG at 1k, 35.2 at 32k (corrected).M5 Ultra 36/80, 512GB, 1.2 TB/s. Price unknown; late October. Qwen proxy est 91–105 TG, 363–384 batch-8 aggregate.unknownest +40%Wait. Do not budget a purchase until Apple posts the price and a post-delivery serving run confirms thermals, speed, and memory behavior.

The fleet priced both ways

This comparison uses the workhorse configuration:

  • Previous: Apple-refurbished M4 Pro 48GB/512GB/10GbE at fact $1,949.
  • New education: M5 Pro 15/16, 48GB/512GB/10GbE at est $2,229.
  • New retail: the same configuration at est $2,399.
  • Performance basis: measured previous aggregate 120 t/s at 8-wide; new planning estimate 132 t/s, or +10%.
  • AppleCare, tax, racks, switches, storage upgrades, and spares are excluded.
FleetPrevious refurb costNew education costNew retail costEdu deltaRetail deltaPrevious aggregate throughputNew aggregate throughput EST
4-mini validationfact (list price) $7,796 = 4×$1,949est $8,916 = 4×$2,229est $9,596 = 4×$2,399+$1,120; +14.4%+$1,800; +23.1%est 480 t/s = 4×120est 528 t/s = 4×132
26-mini faculty fleetfact (list price) $50,674 = 26×$1,949est $57,954 = 26×$2,229est $62,374 = 26×$2,399+$7,280; +14.4%+$11,700; +23.1%est 3,120 t/s = 26×120est 3,432 t/s = 26×132

The previous totals are counterfactual list-price calculations. Availability of 4 or 26 matched refurbished units is unknown.

Cost per stream-capacity unit

A “stream-capacity unit” is 15 aggregate tokens/s. It is useful for cost comparison, but 8.8 units does not prove nine latency-compliant simultaneous sessions.

OptionAggregate throughput per node15-t/s capacity unitsHardware cost per capacity unitChange from previous
Previous M4 Pro refurbfact (proxy) 120 t/sest 8.0est $243.63 = $1,949 / 8Baseline
New M5 Pro educationest 132 t/sest 8.8est $253.30 = $2,229 / 8.8est +4.0%
New M5 Pro retailest 132 t/sest 8.8est $272.61 = $2,399 / 8.8est +11.9%

Plan-ready statement: At education pricing, the newest 48GB fleet costs about 14.4% more and is expected to deliver about 10% more aggregate decode, raising hardware cost per stream-capacity unit by about 4.0%.

Power and cooling, both ways

Apple maximum power is an electrical envelope, not expected LLM draw. The only directly measured wall result found for the target mini class was 46.2W on an M4 Pro 64GB while generating a 70B model.

Formula: BTU/h = watts × 3.412142.

HardwarePer-node power basisEvidence26-unit demand26-unit heat
M4 mini65Wfact Apple maximum wall power1.69 kW5,767 BTU/h
M4 Pro mini46.2Wfact measured wall power under Llama 70B generation1.201 kW4,099 BTU/h
M4 Pro mini140Wfact Apple maximum wall power3.64 kW12,420 BTU/h
M6 or M5 Pro miniActual LLM load unknown; 155Wfact Apple maximum continuous power4.03 kW maximum13,751 BTU/h maximum
M4 Max Studio145Wfact Apple maximum wall power3.77 kW12,864 BTU/h
M3 Ultra Studio270Wfact Apple maximum wall power7.02 kW23,953 BTU/h
M5 Max or M5 Ultra StudioActual LLM load unknown; 480Wfact Apple maximum continuous power12.48 kW maximum42,584 BTU/h maximum

For electrical planning, the 26-mini maximum envelope rises from 3.64kW to 4.03kW, or est +10.7%. Do not compare the new 155W maximum directly with the old 46.2W measured workload result.

Sources: Apple’s Mac mini power-consumption page, Mac Studio power-consumption page, current technical specifications, and the measured M4 Pro LLM report.

Source log

All dynamic catalogs and store pages were accessed 2026-09-03 MDT.

IDSourcePublished or document dateUsed for
A01Apple: M6/M5 Pro Mac mini announcement2026-08-25Starting prices, availability, Apple performance claims
A02Apple: M5 Max/M5 Ultra Mac Studio announcement2026-08-25Prices, dates, workload claims, 512GB timing
A03Apple: M6 and M5 Ultra chip announcement2026-08-25CPU, GPU, AI and bandwidth claims
A04Apple: M5 Pro and M5 Max announcement2026-03M5 Pro/Max percentage claims
A05Apple: current Mac mini specificationsCurrent 2026-09-03Chips, memory, bandwidth, networking, maximum power
A06Apple: current Mac Studio specificationsCurrent 2026-09-03Chips, allowed memory tiers, bandwidth, 10GbE, maximum power
A07Apple retail Mac mini storeCurrent 2026-09-03Retail bases, option prices, date
A08Apple education Mac mini storeCurrent 2026-09-03Education bases and option deltas
A09Apple retail Mac Studio storeCurrent 2026-09-03Retail configurations and availability
A10Apple education Mac Studio storeCurrent 2026-09-03Education configurations
A11Apple: M4/M4 Pro Mac mini launch2024-10-29Previous starting prices and launch availability
A12Apple Support: 2024 Mac mini specifications2024 modelPrevious memory and bandwidth
A13Apple: M4 Max/M3 Ultra Mac Studio launch2025-03-05Previous Studio launch prices and date
A14Apple: M3 Ultra announcement2025-03-05M3 Ultra specifications
A15Apple Support: 2025 Mac Studio specifications2025 modelPrevious chip tiers, memory and bandwidth
A16Apple US Education Institution Price List2026-07-15Pre-refresh institutional catalog prices
A17Apple NASPO PSS catalog2026-06Cooperative-contract catalog
A18Apple education contracts: UtahAccessed 2026-09-03Utah purchasing path
A19Apple refurb M4 Pro 48GB/10GbEAccessed 2026-09-03$1,949 current listing
A20Apple refurb M4 Pro 64GBAccessed 2026-09-03$2,379 current listing
A21Apple refurb M4 Max 128GBAccessed 2026-09-03$4,159 current listing
A22Apple refurb M3 Ultra 256GBAccessed 2026-09-03$8,149 current listing
A23Apple refurb M3 Ultra 512GBAccessed 2026-09-03$17,419 current listing
A24CDW-G M4 Pro 24GB listingAccessed 2026-09-03$1,599, backordered
A25CDW-G M4 Max 36GB listingAccessed 2026-09-03$2,499, backordered
A26CDW-G M3 Ultra 96GB listingAccessed 2026-09-03$5,299, backordered
A27CDW-G M3 Ultra 256GB listingAccessed 2026-09-03$11,639.99, backordered
A28SHI Mac mini searchAccessed 2026-09-03Backorder/zero-stock evidence
A29SHI Mac Studio searchAccessed 2026-09-03Backorder/zero-stock evidence
A30MacRumors: maximum M3 Ultra configuration2025-03-05Dated M3 Ultra chip/memory option prices
A31Tom’s Hardware: M3 Ultra memory-upgrade price history2026-03-06Prior 256GB upgrade price
P01llama.cpp Apple Silicon performance tableLiving discussion; accessed 2026-09-03Same-build M4/M5 Pro and Max PP/TG comparisons
P02llama.cpp gpt-oss guideLiving discussion; accessed 2026-09-03M4 Pro 64GB gpt-oss-20B benchmark
P03oMLX Qwen3.8 field report2026-08-18 and 2026-08-20M4 Pro 64GB and M3 Ultra 512GB serving
P04oMLX M4 Max 128GB benchmarkAccessed 2026-09-03Qwen3.5-35B serving
P05oMLX M3 Ultra 256GB benchmarkAccessed 2026-09-03gpt-oss-120B serving
P06oMLX M3 Ultra 512GB benchmarkAccessed 2026-09-03gpt-oss-120B serving
W01Apple: Mac mini power consumptionAccessed 2026-09-03Previous maximum wall power
W02Apple: Mac Studio power consumptionAccessed 2026-09-03Previous maximum wall power
W03Eastkode M4 Pro LLM wall-power measurement2026; accessed 2026-09-0346.2W Llama 70B generation measurement

Not run — the checks that need the machines in hand

  • not run Physical serving benchmarks on M6, M5 Pro mini, M5 Max Studio, or M5 Ultra Studio; hardware is not available until September 22 or later.
  • not run UVU-specific Apple Education, NASPO, CDW-G, or SHI quote.
  • not run Stock reservation for 4 or 26 identical previous-generation units.
  • not run Exact AppleCare+ premium quote for each new configuration.
  • not run Matched end-to-end benchmark using the same model, quantization, serving version, context, batch, and power meter across both chassis generations.
  • not run Validation of the owner’s private M3 Ultra benchmark; no numerical receipt was supplied in this lane.
  • not run Purchase, order, deployment, or other live-system action.

Review: working paper 36, September 3, 2026 — every price and date checked against the source listed; nothing here contacted UVU or any vendor on UVU's behalf. Configured-to-order prices are Apple's displayed option deltas added to a sourced base price, not institutional quotes. The guide's §04 carries the plain-language summary and the fleet priced both ways.

P

Every angle, and where it lives

The coverage matrix — thirty-eight angles a university leader would raise, graded, with the appendix or section that answers each

Appendix P in five lines

  1. QuestionHas the plan checked every issue a university leader may raise, and where is each answer?
  2. AnswerAlmost; the current table has 37 rows marked COVERED and 1 marked OPEN.
  3. Deciding numbers38 anglesest37 COVERED rows, a derived countest1 OPEN row, a derived countest
  4. What the plan doesUse this matrix as the plan’s index, keep the open quality issue as a test gate, and grade it again after each wave.
  5. Still unknownQuality loss from compressed models stays open until the delivered machines are benchmarked.

The owner's test for this plan is not whether each number is right but whether every angle a university leader would raise has been looked at. This page is the answer, kept honest: each angle graded COVERED (analyzed with evidence), THIN (mentioned, not analyzed), or OPEN (not yet workable), with the place it lives. First graded September 3, 2026 by searching the shipped guide and appendices for each angle; regraded the same afternoon after the five engineering, safety, law, and money reviews (Appendices Q–U) landed. It is regraded every wave; the rows that were THIN at the first grading became lanes of their own and are now COVERED, which leaves 37 COVERED and one OPEN.

The owner's test for this plan is not whether each number is right but whether every angle a university leader would raise has been looked at. This page is the answer, kept honest: each angle is graded COVERED (analyzed with evidence), THIN (mentioned, not analyzed), or OPEN (not yet worked), with the place it lives. First graded on 2026-09-03 by searching the guide and appendices for each angle, then regraded twice the same day as Appendices Q–U and V–Z were added. It is regraded whenever new evidence lands.

AngleGradeWhere it lives
Model choice per memory tier; the most intelligent open modelsCOVERED§11, Appendix L, Appendix Q
Capacity: conversations per box per model; intelligence per dollarCOVERED§04, Appendix Q, configurator (by model and by workload)
Two or more models resident on one machine; swap vs residentCOVERED§04, Appendix Q
512GB frontier node: when it earns its placeCOVERED§04 three-path decision, Appendix Q
Workload classes: chat, coding, documents, agents, speech, imagesCOVERED§06, Appendix R
Prompt-reading (prefill) limits on Apple Silicon; long documentsCOVERED§06, §08, Appendix R
Batching, prefix-cache reuse, speculative decoding, quantization lossCOVEREDAppendix R
Service levels (time to first token, p95) and a real queue modelCOVERED§06, Appendix R
Coding-agent loads: their own queue, cap, overflowCOVERED§06, §08, §10, Appendix R
Minors: the 18,163 concurrent-enrollment studentsCOVERED§08, §10, Appendix S (eleven gates)
Health data: HIPAA vs FERPA, placements, business associatesCOVEREDAppendix S
Crisis disclosures, duty of care, human handoffCOVERED§08, §10, Appendix S
Prompt injection, exfiltration, model supply chain, abuse limitsCOVERED§08, Appendix S (twenty graded controls)
Failover, single points of failure, backups, recovery targets, incident runbookCOVERED§08, §10, Appendix S
Utah AI disclosure law and the Office of AI PolicyCOVERED§08, §09, Appendix T
ADA Title II web accessibility rule (WCAG 2.1 AA by April 26, 2027)COVERED§09, §10, Appendix T
FERPA and open-records status of chat logs; retention; legal holdsCOVERED§08, Appendix T
State higher-education policy, USHE task force and credential, legislature, governorCOVERED§07, §09, Appendix T
Model license fitness for a public university, by exact checkpointCOVERED§11, Appendix T
Utah peers and similar public universities elsewhereCOVERED§12, Appendix T
Five-year cost, refresh cycle, resale valueCOVERED§04, §12, Appendix U, simulator
Lease vs buyCOVEREDAppendix U
Who pays: central, colleges, chargeback, student feesCOVERED§09, Appendix U
Student operations staffCOVERED§09, Appendix U, simulator and configurator switch
Grants and partnershipsCOVERED§09, Appendices F and U
Electricity at UVU's rateCOVERED (proxy; UVU rate unknown)Appendix U
Adoption, behavior, catalystsCOVERED§07, Appendix N
Procurement path, UVU policies 445/447/452COVERED§09, Appendix H
Contingencies (28 branches), failure-modes tableCOVERED§10, Appendix G
Free-tier cloud, harness marketCOVEREDAppendices L, M
Hardware generations, prices, availabilityCOVERED§04, Appendix O
Campus map, department demandCOVEREDmap, Appendices D, I
Pedagogy: assessment redesign and integrityCOVERED§07, §10, Appendix V (160-plus sources)
Pedagogy: learning outcomes and how to measure themCOVERED§07, Appendix W (learning-outcomes contract)
Human factors: trust, sources, error handling in the interfaceCOVERED§08, Appendix X (trust contract, test plan)
Accessibility and language: screen readers, WCAG for chat, Spanish, neurodivergent, mobileCOVERED§08, §09, Appendix Y
Roadmap: Apple cadence, open-model trajectory, buy-vs-wait, wavesCOVERED§09, Appendix Z (24-month watchlist)
Quantized-model quality vs published scores; delivered-unit benchmarksOPEN until hardware shipsAppendix Q and R acceptance-test lists
Q

Model × machine: the matrix

Ten open models against six machines — fit, speed, conversations at once, cost per conversation, intelligence per dollar; two models per box; the 512GB decision

Appendix Q in five lines

  1. QuestionShould the university buy a 512-gigabyte computer for the smartest open model, and how many people can each model-and-machine pair serve?
  2. AnswerKeep smaller Mac minis as the default and keep the 512-gigabyte box as a research choice only after its price and tests pass.
  3. Deciding numbers2 × 60 = 120 score-conversations on the large boxest40 × 52 = 2,080 on five minisestabout $100,000 and no more than 15% of hardware spending as the buying gateest
  4. What the plan doesKeep minis as the campus default, use two-model pairs that fit, and hold the large research box behind price and delivered-machine tests.
  5. Still unknownThe 512-gigabyte price, new-Mac tests, full 32,000-token batch tests, and compressed-model quality tests are NOT_RUN.

The owner's question was direct: why not run the most intelligent models on 512GB Studios, and did we do the multivariate analysis — every model against every machine, capacity included, with more than one model per box? The plan had one capacity number ("four conversations per box") applied to everything. This appendix replaces it with the matrix: ten open models × six machines, each cell with memory fit, single-conversation speed, prompt-reading speed, conversations at ≥10 and ≥20 tokens/s, cost per conversation, and an intelligence-per-dollar index; then the two-models-per-box analysis, the same-money comparison, where an eight-point score gap actually matters, and the three-path decision on a 512GB frontier node. Research cutoff September 3, 2026. fact dated measurements · est derived, basis shown · unknown no defensible number.

What this changes in the plan

1. "Four per box" is retired as a fleet-wide constant. It was conservative for the daily 27B model on a mini (about eight conversations at ≥10 tokens/s by the measured proxies) and for the sparse 35B fast model, fair for coding, and optimistic for the biggest models: on a 512GB node the best open model (GLM-5.3, score 60) fits only two 32k-context conversations with the preferred quantization, at about 18 tokens/s each. The engine keeps 4 as the blended admission cap for mixed campus traffic (Appendix R explains why), and the configurator now shows capacity by workload and by model. 2. Same money, both ways. The 256GB Ultra's price buys five 48GB minis: five minis carry about 40 conversations of a score-52 model; one 512GB node carries two conversations of a score-60 model — 2,080 versus 120 "score-conversations". The eight points matter for advanced coding (Terminal-Bench 88.2 vs 73.0; DeepSWE 66.9 vs 42.2) and complex professional artifacts (a 217-Elo lead on GDPval); for tutoring chat, summaries, and ordinary drafts no matched result shows a difference. 3. The 512GB node becomes a conditional research node (Path 1): kept as an unpriced option behind the router, bought only when the hardware pool reaches about $100,000, it takes no more than 15% of hardware spend, Apple has posted the price, and a delivered-unit test passes. What breaks first on it: long-document reading (a 131k-token prompt took about 23 minutes on the previous generation) and a single point of failure. 4. Two models on one machine is real and now designed in: a 48GB mini holds the 27B daily model plus a 9B fast model comfortably; the 64GB mini is the one that holds the 27B plus the sparse 35B fast model with production margin — that, not speed, is what the extra $360 buys; a 128GB Studio holds a 120B-class model plus the 27B; a 256GB Studio holds GLM-5.3-Flash plus the 27B; a 512GB node running GLM-5.3 at the preferred quantization has no room for a companion. Swapping a 27B takes about three seconds; a 428GB model about 75 seconds — not interactive. llama-server (router mode), oMLX, LM Studio, and vLLM-MLX all support keeping several models resident today. 5. Names fixed. The 256GB anchor is GLM-5.3-Flash (320B total, 18B active, score 57, MIT), not an unnamed "235B-class"; the Appendix O row that credited gpt-oss-120B with 1,405.9 prompt / 79.2 writing tokens/s belonged to a different model and is corrected below (the verified gpt-oss-120B result on the previous 512GB Ultra: 68.4 tokens/s at 1k context, 35.2 at 32k, 187 aggregate at batch eight). est

Research cutoff: 2026-09-03 MDT.

fact means a dated source states or measures the number. est means arithmetic or a proxy, with its basis stated. unknown means no defensible number was available.

Executive verdict

  • The current “four conversations per box at ≥10 t/s” rule is not valid as one fleet-wide constant.
  • est It is conservative for Qwen3.8-27B and Qwen3.6-35B-A3B workhorses.
  • est It is too optimistic for Qwen3.8-Flash-Next on 128GB when four 32k contexts are required.
  • est It is too optimistic for the quality-preferred 427.7GB GLM-5.3 build: only two 32k contexts fit under the production memory rule.
  • FACT/EST A 512GB GLM node buys real gains on advanced coding and complex professional artifacts, but minis buy far more concurrent capacity per dollar.
  • Recommendation: retain minis as the campus default. Keep a 512GB node as a conditional research lane only after Apple posts the price and UVU verifies the exact model on delivered hardware.

1. Candidate models

Practical-memory rule

est I reserve 15% of unified memory for macOS, Metal/MLX buffers, the server, allocator variance, and safety. Model weights plus every reserved 32k KV cache must fit inside the remaining 85%.

Quantized-model scores remain a caveat: Artificial Analysis scores the model, not each community Apple quantization. Quality retained by a given quant is unknown until tested.

Model ledger

ModelTotal / activeAA Intelligence IndexLicense and inputsPractical Apple weightsBF16 KV for one 32k conversationWeights + one cache
Qwen3.5-9Bfact 9.7B densefact 22, accessed 2026-09-03Apache-2.0; text/image/videofact 5.98GB MLX 4-bitest 1.074GBest 7.05GB
Qwen3.8-27Bfact 27B densefact 52Apache-2.0; text/image/videofact 16.1GB MLX 4-bitfact (math) 2.147GBest 18.25GB
Qwen3.6-35B-A3Bfact 35B / 3B; AA rounds total to 36Bfact 32Apache-2.0; maker supports text/image/video; AA records text/imagefact 20.7GB MLX DWQfact (math) 0.671GBest 21.37GB
Qwen3.8-Flash-Nextfact 180B stored / 6B activefact 56Qwen Community 1.0; text/image/videofact 106.2GB mixed MLXest 0.830GBest 107.03GB
gpt-oss-120bfact 117B / 5.1Bfact 24Apache-2.0; textfact 62.4GB MLX MXFP4-Q4est 1.213GBest 63.61GB
Mistral Small 4fact 119B / 6.5Bfact 20Apache-2.0; text/imagefact 67.8GB MLX 4-bitfact (math) 0.755GBest 68.55GB
GLM-5.3-Flashfact 320B / 18Bfact 57MIT; maker supports text/image/video/filesfact 181.9GB preferred mixed MLX; uniform is 177.6GBest 0.392GBest 182.29GB
GLM-5.3fact 753B checkpoint / 40B activefact 60Custom GLM-5.3 license; textfact 427.7GB preferred mixed MLX; uniform is 418.6GBFACT/EST 3.121GBest 430.82GB
Kimi K3fact 2.8T / 104Bfact 60Custom Kimi K3 license; maker supports text/image/videofact 1.56TB official MXFP4; fact 928.6GB quality GGUFfact (math) 0.906GBest 929.51GB using GGUF
DeepSeek V4 Pro 0813fact about 1.57T / 48B; AA rounds to 1.6T/49Bfact 53MIT; textfact about 850GB Q4_K_XL GGUFfact (math) 0.331GB optimized hybrid cacheest 850.33GB

Model facts and scores come from the current Artificial Analysis model pages and the makers’ model cards: Qwen3.8-27B, Qwen3.6-35B-A3B, Qwen3.8-Flash-Next, gpt-oss-120b, Mistral Small 4, GLM-5.3-Flash, GLM-5.3, Kimi K3, and DeepSeek V4 Pro.

KV-cache math

All arithmetic uses 32,768 tokens and two-byte BF16 cache values.

Qwen3.5-9B:
32,768 × 8 full layers × 2 K/V × 4 KV heads × 256 × 2 bytes
= 1,073,741,824 bytes

Qwen3.8-27B:
32,768 × 16 full layers × 2 × 4 × 256 × 2
= 2,147,483,648 bytes

Qwen3.6-35B-A3B:
32,768 × 10 full layers × 2 × 2 × 256 × 2
= 671,088,640 bytes

Qwen3.8-Flash-Next:
core = 32,768 × 12 full layers × 2 × 2 × 256 × 2
index estimate = (32,768 / 4) × 12 × 128 × 2
total = 830,472,192 bytes

gpt-oss-120b:
[(32,768 × 18 full layers) + (128 × 18 sliding layers)]
× 2 K/V × 8 KV heads × 64 × 2
= 1,212,678,144 bytes

Mistral Small 4 optimized MLA:
32,768 × 36 × (256 latent + 64 rope) × 2
= 754,974,720 bytes

GLM-5.3-Flash:
main = 32,768 × 11 MLA layers × 512 latent × 2
plus estimated compressed index = about 23MB
total = about 392MB

GLM-5.3:
main = 32,768 × 78 × (512 latent + 64 rope) × 2
indexers = 32,768 × 21 × 128 × 2
total = 3,120,562,176 bytes

Kimi K3:
32,768 × 24 MLA layers × (512 latent + 64 rope) × 2
= 905,969,664 bytes

DeepSeek V4 Pro:
30 c4 layers × 10,616,832 bytes
+ 31 c128 layers × 393,216 bytes
= 330,694,656 bytes

Hybrid linear-attention models also have fixed recurrent state. That state is not fully exposed by model configs; it is covered by the 15% reserve.

Releases since August 15

  • fact Qwen3.8-Flash-Next arrived August 26 and changes the 128GB quality ceiling.
  • fact GLM-5.3-Flash received its formal open-weight launch September 2 and changes the 256GB choice.
  • fact K2 Horizon launched September 3. Its 375B-A23B flagship scores 47, but the interesting 32B and 36B-A4B variants have no current independent score or proven Apple quant. est It is a watch item, not a matrix winner today. IFM announcement, AA flagship analysis.

2. Machines

MachinePhysical / usable memoryBandwidthPrice usedAvailability
Mac mini M5 Pro 48GBfact 48GB; est 40.8GB usablefact 307GB/sest $2,139 educationfact 2026-09-22
Mac mini M5 Pro 64GBfact 64GB; est 54.4GB usablefact 307GB/sest $2,499 educationfact 2026-09-22
Mac Studio M5 Max 128GBfact 128GB; est 108.8GB usablefact 614GB/sest $4,639 education, 512GB SSDfact 2026-09-22
Mac Studio M5 Ultra 256GBfact 256GB; est 217.6GB usablefact 1.2TB/sest $10,799 retail, 1TB SSD; education unknownfact 2026-09-22
Mac Studio M5 Ultra 512GBfact 512GB; est 435.2GB usablefact 1.2TB/sunknown retail and educationfact late October 2026
Mac Studio M3 Ultra 512GB proxyfact 512GB; est 435.2GB usablefact 819GB/sfact $17,419 current 8TB refurb listing; launch-equivalent est $9,499Available quantity unknown

Hardware and dates are from Apple’s Mac mini announcement, Mac Studio announcement, mini specifications, and Studio specifications.

The configured price sums are est, not UVU quotes. In particular, data.js treats some derived prices as facts and includes a $9,899 Ultra education estimate that is not supported by a current Apple quote.

3. Model × machine matrix

Estimation method

For measured or closely matched models, I use the published oMLX/MLX result.

For unmeasured MoE models:

EST single decode =
memory bandwidth ÷ active 4-bit bytes per token × 30% efficiency

est 30% is bracketed by approximately 34% for the measured Qwen 3B-active MoE and approximately 23% for measured gpt-oss MXFP4.

New-generation planning factors are:

  • est ×1.10 for M5 Pro/Max versus comparable M4 measurements.
  • est ×1.40 for M5 Ultra versus M3 Ultra.

S10/S20 means estimated simultaneous conversations at at least 10/20 generated tokens per second. Counts reserve one 32k cache each, but published batch tests generally use shorter active prompts. Therefore 32k memory fit is stronger evidence than 32k speed.

8 means “at least eight under the batch proxy”; no claim above eight is made.

Score-streams/$10k = AA score × S10 × 10,000 ÷ hardware price. It is only a capacity-quality index.

M5 Pro mini 48GB

ModelFitSingle decodePrefillS10 / S20Cost per S10 / S20Score-streams per $10k
Qwen3.5-9Best yesest 62 t/sunknownest 8 / 8est $267 / $267est 823
Qwen3.8-27Best yesest 22.9 t/s @8kest 110–141 t/s @4kest 8 / 1est $267 / $2,139est 1,945
Qwen3.6-35B-A3Best yesest 62 t/s; same-chip 20-GPU upper proxy is 68.7est about 1,600 @16kest 8 / 8est $267 / $267est 1,197
Other seven modelsFACT/EST no0 / 0

The 48GB machine can fit Qwen3.8-27B and Qwen3.6 together only with narrow headroom; see co-residency.

M5 Pro mini 64GB

Decode speed is the same as 48GB because bandwidth is unchanged.

ModelFitSingle decodePrefillS10 / S20Cost per S10 / S20Score-streams per $10k
Qwen3.5-9Best yesest 62unknownest 8 / 8est $312 / $312est 704
Qwen3.8-27Best yesest 22.9 @8kest 110–141 @4kest 8 / 1est $312 / $2,499est 1,665
Qwen3.6-35B-A3Best yesest 62est about 1,600 @16kest 8 / 8est $312 / $312est 1,024
Other seven modelsFACT/EST no0 / 0

The extra $360 buys co-residency and cache space, not more decode bandwidth.

M5 Max Studio 128GB

ModelFitSingle decodePrefillS10 / S20Cost per S10 / S20Score-streams per $10k
Qwen3.5-9Best yesest 135unknownest 8 / 8est $580 / $580est 379
Qwen3.8-27Best yesest 52unknownest 8 / 8est $580 / $580est 897
Qwen3.6-35B-A3Best yesest 110 @8kest 1,635 @8k, same-chip proxyest 8 / 8est $580 / $580est 552
Qwen3.8-Flash-Nextest yes, tightest 61unknownest 3 / 3, memory-limitedest $1,546 / $1,546est 362
gpt-oss-120best yesest 49 short-contextunknownest 8 / 6 short-contextest $580 / $773est 414
Mistral Small 4est yesest 57unknownest 8 / 7est $580 / $663est 345
Larger four modelsFACT/EST no0 / 0

Qwen3.8-Flash-Next permits three 32k caches by arithmetic, not four. Its current custom Apple artifact also excludes MTP and operates text-only despite carrying vision weights.

M5 Ultra Studio 256GB

ModelFitSingle decodePrefillS10 / S20Cost per S10 / S20Score-streams per $10k
Qwen3.5-9Best yesest 244unknownest 8 / 8est $1,350 / $1,350est 163
Qwen3.8-27Best yesest 98, range 91–105est about 603 @8kest 8 / 8est $1,350 / $1,350est 385
Qwen3.6-35B-A3Best yesest 205unknownest 8 / 8est $1,350 / $1,350est 237
Qwen3.8-Flash-Nextest yesest 120unknownest 8 / 8est $1,350 / $1,350est 415
gpt-oss-120best yesest 96 short-contextest about 841 @32kest 8 / 8 short-contextest $1,350 / $1,350est 178
Mistral Small 4est yesest 111unknownest 8 / 8est $1,350 / $1,350est 148
GLM-5.3-Flashest yesest 40unknownest 8 / 4est $1,350 / $2,700est 422
GLM-5.3, Kimi K3, DeepSeek V4 ProFACT/EST no0 / 0

The plan’s unnamed “235B-class” anchor remains unknown: a parameter count does not specify active weights, artifact size, cache design, or batch behavior. GLM-5.3-Flash is the stronger named 256GB choice.

M5 Ultra Studio 512GB

Price-based columns are unknown until Apple posts the 512GB price, denoted P.

ModelFitSingle decodePrefillS10 / S20Cost per S10 / S20Score-streams per $10k
Qwen3.5-9Best yesest 244unknownest 8 / 8est P/8 / P/8unknown
Qwen3.8-27Best yesest 98est about 603 @8kest 8 / 8est P/8 / P/8unknown
Qwen3.6-35B-A3Best yesest 205unknownest 8 / 8est P/8 / P/8unknown
Qwen3.8-Flash-Nextest yesest 120unknownest 8 / 8est P/8 / P/8unknown
gpt-oss-120best yesest 96 short-contextest about 841 @32kest 8 / 8 short-contextest P/8 / P/8unknown
Mistral Small 4est yesest 111unknownest 8 / 8est P/8 / P/8unknown
GLM-5.3-Flashest yesest 40unknownest 8 / 4est P/8 / P/4unknown
GLM-5.3 mixed 4/8est yes, very tightest 18est about 132; GLM-5.2 131k proxyest 2 / 0, memory-limitedest P/2 / N/Aunknown
Kimi K3, DeepSeek V4 Profact no0 / 0

The lower-quality 418.6GB uniform GLM build has room for five calculated 32k caches and could reach est three ≥10 t/s streams. The preferred 427.7GB mixed build fits only est two.

M3 Ultra Studio 512GB measured proxy

ModelFitSingle decodePrefillS10 / S20Cost per S10 / S20Score-streams per $10k
Qwen3.5-9Best yesest 174unknownest 8 / 8est $2,177 / $2,177est 101
Qwen3.8-27Bfact yesfact 65–79; central 70fact 420–442; central 431FACT/EST 8 / 8est $2,177 / $2,177est 239
Qwen3.6-35B-A3Best yesest 141unknownest 8 / 8est $2,177 / $2,177est 147
Qwen3.8-Flash-Nextest yesest 82unknownest 8 / 8est $2,177 / $2,177est 257
gpt-oss-120bfact yesfact 68.4 @1k; 57.8 @8k; 35.2 @32kfact 600.7 @32kFACT/EST 8 / 8 short-contextest $2,177 / $2,177est 110
Mistral Small 4est yesest 76unknownest 8 / 8est $2,177 / $2,177est 92
GLM-5.3-Flashest yesest 27unknownest 7 / 2est $2,488 / $8,710est 229
GLM-5.3 mixed 4/8est yes, tightest 12est about 95; GLM-5.2 proxyest 1 / 0est $17,419 / N/Aest 34
Kimi K3, DeepSeek V4 Profact no0 / 0

Important source correction

findings-36 associates oMLX result mrqg3jhh with gpt-oss-120b, but that result is for Qwen3.5-122B-A10B.

The verified gpt-oss-120b M3 Ultra result is:

  • fact 68.4 t/s at 1k.
  • fact 57.8 t/s at 8k.
  • fact 35.2 t/s at 32k.
  • fact 187.1 t/s aggregate at batch eight in the short-context batch test.

The earlier 1,405.9 prefill / 79.2 decode claim should be removed.

Audit of “four conversations per box”

Tier and modelVerdictWhy
48GB or 64GB, Qwen3.8-27Best conservative at ≥10Old 64GB proxy reached 120–134 aggregate at batch eight. Full 32k decode is still unknown.
48GB or 64GB, Qwen3.6-35B-A3Best conservativeSparse active weights and same-chip results leave substantial throughput room.
128GB, Mistral or gpt-ossest conservative at normal contextBoth have memory for many caches and estimated S10 of eight.
128GB, Qwen3.8-Flash-Nextest optimisticOnly three 32k caches fit under the rule.
256GB, GLM-5.3-Flashest conservative at ≥10; fair near ≥20Estimated eight S10 but only four S20.
256GB, unnamed “235B-class”unknownNo exact model or quant was specified.
512GB, preferred GLM-5.3est optimisticMixed quant permits only two 32k contexts; uniform quant permits more but has lower known quant quality.
512GB, GLM-5.3-Flashest conservativeEight S10 is plausible with comfortable memory.

4. More than one model on one machine

Pairs that fit

Figures include one 32k cache for each model.

MachinePairMemoryVerdict
48GB miniQwen3.8-27B + Qwen3.5-9Best 25.30GBComfortable
48GB miniQwen3.8-27B + Qwen3.6-35Best 39.62GB of 40.8GB usableFits by arithmetic; too little operating margin for production
64GB miniQwen3.8-27B + Qwen3.6-35Best 39.62GB of 54.4GBGood dual-resident case
128GB Studiogpt-oss-120b + Qwen3.8-27Best 81.86GBComfortable
128GB StudioMistral Small 4 + Qwen3.8-27Best 86.80GBComfortable
128GB StudioQwen3.8-Flash-Next + another useful modelMore than 108.8GBDoes not fit under production rule
256GB StudioGLM-5.3-Flash + Qwen3.8-27Best 200.54GBFits with about 17GB usable headroom
512GB Studiopreferred mixed GLM-5.3 + Qwen3.5-9Best 437.87GBDoes not fit under 85% rule
512GB Studiouniform GLM-5.3 + Qwen3.5-9Best 428.77GBFits, but uses the lower-quality GLM quant
512GB StudioGLM-5.3 + Qwen3.6-35BAt least est 452GB using preferred GLMDoes not fit

The proposed “GLM-5.3 plus a fast MoE lane” is therefore not valid with the preferred mixed quant and the production reserve. A small companion requires the uniform GLM build or a looser safety rule.

Speed cost when both models are active

Memory bandwidth is shared. A simple first-order estimate is:

active speed ≈ solo speed × assigned bandwidth share
Machine and pairEqual 50/50 splitBetter routing choice
48GB: Qwen27 + Qwen9est 11.5 + 31 = 42.5 aggregateKeep Qwen27 above 10; send short/simple work to Qwen9
64GB: Qwen27 + Qwen35est 11.5 + 31 = 42.5Qwen35 handles high-volume chat; Qwen27 gets quality requests
128GB: gpt-oss + Qwen27est 24.5 + 26 = 50.5Both remain responsive
256GB: GLM-Flash + Qwen27est 20 + 49 = 69Strong coding/research lane plus workhorse
512GB: uniform GLM + Qwen9est 9 + 122Give GLM at least 56% of bandwidth: est 10 GLM + 107 Qwen9

These are ideal bandwidth splits, not scheduler measurements.

Serving-stack support today

  • fact llama-server has router mode, model-specific child processes, dynamic load/unload, autoload, and a default maximum of four loaded models. A current LRU issue can unload an active model, so production validation is required.
  • fact oMLX has a multi-model engine pool with LRU, TTL, and explicit load/unload. An M4 Pro 64GB field run kept about 37.73GB of two models resident and switched between them in fact 0.23–0.9 seconds.
  • fact LM Studio can hold multiple models when JIT auto-evict is off or models are loaded manually. Its default auto-evict behavior keeps at most one JIT model.
  • fact vLLM-MLX provides a named multi-model registry, lazy loading, preloading, LRU, and wait/fail/preempt policies. Its memory budget counts weights, so KV, activations, and OS reserve must be added by the operator.
  • fact Native mlx-lm server replaces its primary model when a different model key is loaded; it is not a multi-resident router.
  • fact Ollama can keep multiple models resident when memory permits, controlled separately from per-model parallelism.

Swap times

Prior internal-SSD measurements are fact 5.1–5.8GB/s. New M5 desktop SSD speed is unknown.

ModelRaw-read floorOperational meaning
Qwen3.5-9Best 1.0–1.2sEasy to swap
Qwen3.8-27Best 2.8–3.2sOften acceptable on demand
Qwen3.6-35Best 3.6–4.1sOften acceptable
gpt-oss-120best 10.8–12.2sNoticeable
Mistral Small 4est 11.7–13.3sNoticeable
Qwen3.8-Flash-Nextest 18.3–20.8sKeep resident if used regularly
GLM-5.3-Flashest 31.4–35.7sKeep resident during active periods
GLM-5.3 preferred mixedest 73.7–83.9sSwapping is poor interactive service

An oMLX incident loaded fact 403.96GB in 73 seconds, or 5.53GB/s effective, although the storage medium was not proved to be the internal SSD.

Residency verdict

  • Tutoring, summaries, and first drafts: pin Qwen3.6-35B or Qwen3.5-9B.
  • General faculty work: pin Qwen3.8-27B.
  • Advanced coding and complex document work: route to Qwen3.8-Flash-Next, GLM-Flash, or the metered frontier pool.
  • Keep two models resident when both receive frequent, alternating traffic.
  • For rare large-model requests, swapping a 60–70GB model can be acceptable. Swapping a 428GB model is not.

5. Frontier questions

Same money: 512GB GLM versus 48GB minis

Let the unknown 512GB M5 Ultra price be P.

Same-money line: est GLM preferred quant: 2 conversations × score 60 = 120 score-conversations, cost P/2 each; Qwen27 minis: 8 × floor(P / $2,139) conversations × score 52 = 416 × floor(P / $2,139) score-conversations, cost about $267 each.

For illustration only, using the cheaper 256GB Ultra’s est $10,799 price—not a 512GB quote—buys five 48GB minis:

GLM node:     2 × 60 = 120 score-conversations
Five minis: 40 × 52 = 2,080 score-conversations

This multiplication is a routing-capacity heuristic, not a measure of completed academic work.

Where the eight-point gap matters

The comparison is AA 60 versus 52. The AA methodology weights agentic work, coding, scientific reasoning, and general knowledge. Eight index points do not mean eight percentage points on every task.

UVU workEvidenceVerdict
Advanced codingfact Official cards report GLM versus Qwen: Terminal-Bench 2.1 88.2 vs 73.0, DeepSWE 66.9 vs 42.2, NL2Repo 58.0 vs 42.3Gap matters
Complex research and professional artifactsfact Independent GDPval-AA v2 reports 1763 vs 1546 Elo, a 217-Elo GLM lead across work products such as memos, slides, and spreadsheetsGap matters
Math and scientific reasoningunknown Qwen reports GPQA 89.2, while GLM-5.3 does not publish a directly comparable current GPQA figureDo not claim a proven advantage
Tutoring chatunknown No AA component directly measures ordinary campus tutoringRoute to cheaper model unless a UVU test shows a difference
Basic summariesunknown No matched task result proves the GLM advantageQwen-tier model is sufficient until tested
Ordinary draftsunknown GDPval supports complex professional work, not simple proseQwen-tier model is the default

Official task-level results: GLM-5.3 card, Qwen3.8-27B card, and GDPval-AA.

When a 512GB node earns a place

Path 1 — Recommended: conditional shared research node

  • Result: keep it in the plan as an unpriced option behind the router.
  • Budget trigger: est total hardware pool of about $100,000 or more, assuming an eventual node price near $15,000 and a policy that one frontier node consumes no more than 15% of hardware spend.
  • Benefit: local advanced coding and professional-artifact work.
  • Risk: price, quant quality, concurrency, thermals, and long-context behavior remain unproved.
  • Undo: remove the option before purchase if price or acceptance results fail.

Path 2 — Ring-fenced research purchase

  • Result: one department or grant buys it once P is posted.
  • Benefit: earlier access for research that has demonstrated need.
  • Risk: low utilization and no campus-service redundancy.
  • Undo: repurpose it for batch research, but the purchase itself is hard to undo.

Path 3 — No 512GB node

  • Result: spend the same money on minis, 128GB nodes, and the metered frontier pool.
  • Benefit: much higher concurrency and multiple failure domains.
  • Risk: some advanced local work remains cloud-routed.
  • Undo: easy; add a later model generation when evidence improves.

Approve Path 1, the conditional shared research-node option? Yes or no.

What breaks first

Long-document prefill breaks first.

The closest measured proxy is GLM-5.2, which shares GLM-5.3’s base architecture:

  • fact A 131k-token cold prefill took about 1,385 seconds, roughly 23 minutes.
  • fact A 196k attempt was rejected at an estimated 494.81GB peak.
  • fact Its steady KV allocation was only about 20GiB.
  • est M5 Ultra’s planning gain does not turn a multi-minute 131k prefill into an interactive request.

That is transient prefill memory and latency failing before steady 32k KV capacity. One node is still a structural single point of failure: a hardware or service fault removes fact 100% of the local frontier lane.

Source: oMLX GLM long-context report.

What the plan should change

  • 1. Replace the blanket four-stream assumption. Store streams10 and streams20 by exact model, quant, context, runtime, and machine.
  • 2. Keep Qwen3.8-27B as the 48GB default. It has the best estimated score-capacity value in that tier.
  • 3. Buy 64GB minis only for co-residency or longer contexts. They have the same 307GB/s bandwidth and expected single-stream speed as 48GB.
  • 4. Use Mistral Small 4 or gpt-oss as the stable 128GB shared service. Treat Qwen3.8-Flash-Next as a three-context pilot, not a four-stream default.
  • 5. Name GLM-5.3-Flash as the 256GB candidate. Remove the unspecified “235B-class” capacity claim.
  • 6. Do not count four GLM-5.3 conversations. Use est two for the preferred mixed quant or est three for the smaller uniform quant, pending measurement.
  • 7. Make the 512GB node conditional and research-oriented. Do not use it as the campus workhorse.
  • 8. Correct the gpt-oss benchmark citation and price evidence states in findings-36 and data.js.
  • 9. Pin small daily models and route advanced requests. Avoid making two large resident models generate simultaneously unless latency is not important.
  • 10. Run acceptance tests after delivery. The minimum matrix is 1k/8k/32k prefill and decode at batch 1/2/4/8, cold-load time, peak memory, and a 24-hour mixed-traffic soak.

Source log

NOT_RUN

  • not run Any M5 Pro mini, M5 Max Studio, or M5 Ultra Studio benchmark; the desktops are not delivered yet.
  • not run GLM-5.3 or GLM-5.3-Flash on the target Apple machines.
  • not run Quantized-model quality comparison against the full-precision Artificial Analysis scores.
  • not run Full 32k batch 1/2/4/8 tests for every model.
  • not run New-machine SSD load tests.
  • not run UVU institutional price quote.
  • not run 512GB M5 Ultra price; Apple has not posted it.
  • not run Power, thermal, failure-rate, or 24-hour mixed-traffic tests.
  • not run Contact with UVU, Apple, a model maker, or any vendor.
  • not run Purchase, deployment, publication, or live-system change.

The data behind this appendix

In the underlying file, memGB is the practical weight of a model plus one reserved 32,000-token cache, and a missing prefill figure means unknown. Every conversation count in it is a planning estimate at the benchmark settings described above.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/38-model-machine-matrix.md in the research pack.

R

Service engineering

Seven workload classes, what binds each on Apple Silicon, the capacity multipliers and what they measured, service levels, the queue model, coding agents, and the replacement for four-per-box

Appendix R in five lines

  1. QuestionHow should the service handle seven kinds of work, and what replaces four conversations per box?
  2. AnswerCold documents and coding agents fail first, so the service needs separate lines, saved document work, strict agent limits, and overflow.
  3. Deciding numbers8 chat conversations per 48-gigabyte miniest2 cold or 4 warm document conversations per miniest1 agent session per miniest
  4. What the plan doesMake separate queues, index each document once, reuse saved work on the same box, limit agents, and test four new nodes after September 22.
  5. Still unknownActual university traffic, saved-work reuse, exact M5 capacity, the four-node test, speech and image tests, and cloud overflow are NOT_RUN or UNKNOWN.

A campus AI service does not carry one kind of traffic. This appendix builds the engineering view the plan lacked: seven workload classes with measured token profiles, which ones are bound by writing speed and which by prompt reading (Apple's weak spot), the capacity multipliers available in today's Apple serving software and what they actually measured, proposed service levels per class, a queue model that replaces the plan's fixed 30-second service time, what coding agents do to a fleet, and a replacement for "four conversations per box". Research cutoff September 3, 2026. fact measured · est derived · unknown not yet measurable on the unshipped M5 desktops.

What this changes in the plan

1. Documents are the workload that breaks first. A 30-page upload read cold by the 27B daily model takes about 69–167 seconds to the first word on a mini (36–76 on a 128GB Studio); the sparse 35B fast model does it in 10–25 seconds. So the service indexes every upload once, retrieves 2,000–6,000 relevant tokens per question instead of replaying the file, caches exact course prefixes, keeps a student's later questions on the box holding that cache, and sends cold large prompts to the fast lane — never the dense model. 2. The multipliers are real but must not be multiplied together. Continuous batching: 2.4–2.6× aggregate throughput at four to eight conversations (measured); reusing a cached prefix: a 115k-token prefix went from about 20 minutes cold to 43 seconds warm (measured); speculative decoding: 2.3–2.6× for one conversation on the exact model but only 1.06× at batch four; 4-bit versus 8-bit compression: 1.75 quality points lost for 36% more speed and 46% less memory — use 4-bit for the daily lane and 6-bit where writing or code correctness matters. 3. The fixed 30-second service time understates a finals-week mix by about 70% (weighted mean 51 seconds; 93 occupied conversations instead of 54 for the same arrivals). That sits inside the plan's 9× stress envelope, so the fleet sizes stand — but the reason is now stated honestly. 4. Coding agents change everything if admitted freely. One agent turn consumes about 580,000 prompt tokens (94% cached) — roughly 1,200× a chat; if 5% of users start agents in the peak hour, a 26-mini fleet runs at 133% utilization and p95 waits reach 16 minutes. Agents therefore get their own queue, a metered allowance (illustrative: 64K in + 8K out per agent-active user-day, about 0.84× the whole chat baseline), session affinity so caches survive, and overflow to the metered frontier pool before utilization reaches 85%. 5. "Four per box" becomes a table: chat 8 / coding 4 / documents 2 cold or 4 warm / agents 1 / research writing 2 per 48GB mini (Studios about double on the daily model); speech and image work are measured in seconds per job, not conversations. Every number is an admission cap pending the four-node acceptance test after September 22: cold and warm prompts, 1/2/4/8 concurrency, p50/p95 time-to-first-token, per-user and aggregate speed, cache reuse, memory, thermals. est

Workload classes

“Request” means one external user turn unless noted. These are planning inputs, not measured UVU usage.

WorkloadInput per requestOutput per requestTypical contextRequests per user-dayEvidence
Tutoring/chatest 3,000 processed tokensfact 1,115 mean tokensest 6,400 tokensfact (derived) 7.0 per weekly-active-user calendar dayCedarville reported fact 49 interactions/week among active users. ShareChat reported fact 135-token user messages, fact 1,115-token assistant messages and fact 4.62 turns/conversation. Cedarville, Summer 2026, ShareChat, revised 2026-05-17
Coding helpest 4,000 tokensest 800 tokensest 8,000 tokensest 8 per coding-active dayStudyChat measured fact 2,214 conversations, fact 16,851 student utterances, and fact 7.6 utterances/conversation from fact 203 students. Token sizes were unknown. StudyChat, revised 2026-03-04
Document questionsest 4,000 retrieved tokens/question; cold file est 6,667–66,667 tokens for fact supplied 10–100 pages at est 500 words/pageest 500 tokensest 16,000 tokensest 4, rounded from fact (derived) 3.70Prof. Leodar recorded fact 12,330 queries from fact 34 consenting students over fact 14 weeks. Frontiers, 2026-06-16
Agentic sessionsfact 582,500 cumulative prompt tokens/external turn, including fact 545,300 cachedfact 4,000 cumulative tokens/turnfact about 68,000 prompt tokens/internal callfact (derived) 4.2455 external turns/calendar dayGitHub’s production trace covered fact 3.2M users, fact 13.5M sessions, fact 95.1M turns, fact 760.5M model calls, and fact 774.7M tool calls in fact one week. Agentic Coding in the Wild, 2026-07-30
Research writingest 2,000 tokensest 1,000 tokensest 16,000 tokensfact 8.5 prompts per controlled writing-task dayA fact 45-minute study with fact 60 students recorded fact 512 prompts; actual API token counts were unknown. Frontiers, 2026-03-19; corrected 2026-05-11
Lecture transcriptionInput is audio, so text input tokens are UNKNOWN/not applicableest 5,260–7,380 transcript tokens/lecturefact 30-second audio windowsunknown; planning placeholder est 1 per submitting-user dayMedical lectures measured fact 30.71–45.71 minutes and fact 3,945–5,535 words. Whisper processes fact 30-second chunks. Medical lecture study, 2025-02-04, Whisper, 2022-09-21
Image tasksfact 256–1,280 visual tokens/image plus est 50 text tokensImage generation tokens are UNKNOWN/not applicable; image understanding est 250 text tokensest 2,000 working tokens; model maximum fact 32,768unknown; planning placeholder est 2 per image-active dayQwen2.5-VL’s model card supplies visual-token geometry, not campus demand. Qwen2.5-VL, January 2025

The page conversion uses fact approximately 1 token per 0.75 English words, with model and language variation. OpenAI token guidance, accessed 2026-09-03

For document service, index the upload once and retrieve est 2,000–6,000 relevant tokens/question. Replaying the entire est 6,667–66,667-token file on every question is not a workable default.

Bound analysis on Apple Silicon

The planned M5 Pro mini and M5 Max Studio do not ship until fact 2026-09-22. Their desktop serving performance is therefore unknown on fact 2026-09-03. The estimates below use shipped M4 desktop results and an explicit est +10% M5 planning credit supported by same-build laptop-chip comparisons. Apple Mac mini announcement, 2026-08-25, Apple Mac Studio announcement, 2026-08-25

All bound classifications below are est engineering classifications.

Class48GB mini64GB mini128GB StudioMain risk
Tutoring/chatDecode-boundDecode-boundDecode-boundLong answers and too many active decoders
Coding helpMixed; prefill-bound when code/repository text is pastedMixed, with more KV headroomMixed; usually decode-bound on short helpRepository-scale prompts
Document questionsPrefill-bound cold; mixed warmPrefill-bound cold; extra RAM does not materially raise bandwidthPrefill-bound cold; mixed warmFull-file replay and cache misses
Agentic sessionsPrefill/cache-boundPrefill/cache-boundPrefill/cache-bound; decode also matters with a heavy modelRepeated internal calls and resident KV state
Research writingMixed; output-heavy after sources loadMixedMixed/decode-boundLong source packs and long output
Lecture transcriptionAudio encoder-bound, not text decode-boundSameSameReal-time factor and audio memory
Image tasksVision-prefill or diffusion-boundSameSame, with more model headroomImage size, model type and GPU memory

Cold first token for a 30-page upload

est 30 pages × 250–500 words/page ÷ 0.75 words/token = 10,000–20,000 prompt tokens. The range excludes parsing, OCR, queueing and retrieval.

BoxDense 27BSparse 35B-A3B MoEResult
Shipped M4 Pro 48/64GB proxyest 76–184 secondsest 11–28 secondsPrefill-bound
Planned M5 Pro 48/64GBest 69–167 secondsest 10–25 secondsDirect result unknown
Shipped M4 Max 128GB proxyest 39–84 secondsest 9–17 secondsPrefill-bound
Planned M5 Max 128GBest 36–76 secondsest 8–16 secondsDirect result unknown

The measurements behind those estimates include:

  • M4 Pro/48GB dense Qwen3.8-27B: fact 126.1 prompt tokens/s at 16K. oMLX, 2026-08-27
  • M4 Pro/48GB sparse Qwen3.5-35B-A3B: fact 890.2 prompt tokens/s at 8K, fact 830.3 at 16K, and fact 713.2 at 32K. oMLX, 2026-04-07
  • M4 Max/128GB dense Qwen3.8-27B: fact 255.9 prompt tokens/s at 16K and fact 238.7 at 32K. oMLX, 2026-08-19
  • M4 Max/128GB sparse Qwen3.5-35B-A3B: fact 1,160 prompt tokens/s at 32K and fact 28.247-second TTFT for 32,768 tokens. oMLX, 2026-03-17

The router should therefore:

  • 1. Put short chat and short coding turns in a fast lane.
  • 2. Parse and index uploads before the first question.
  • 3. Preserve exact course prefixes and route later turns back to the box holding that cache.
  • 4. Use the sparse MoE lane for cold large prompts.
  • 5. Send unusually large, uncached or deadline-sensitive work to the metered cloud pool.

Capacity multipliers available today

Serving-stack support

StackFACT support available by 2026-09-03Measured gainPlanning treatment
llama-serverParallel slots, continuous batching enabled by default, prompt caching, cache reuse, quantized KV, and several speculative modesControlled Apple batching/cache gain unknown; one M1 Max MTP report measured fact 11%–28% slowerUse batching, but give speculation est 1.0× capacity credit until the exact model passes an A/B test. Server docs, MTP issue, 2026-05-27
MLX-LMContinuous-batch server, separate prompt/decode concurrency, chunked prefill, LRU prompt cache and draft modelsOfficial controlled Apple batch multiplier unknownCurrent code makes the draft-model path non-batchable, so batch and speculation gains must not be multiplied. MLX-LM server, accessed 2026-09-03
LM Studio / MLX engineDraft-model speculation, disk-backed KV checkpoints and continuous VLM batchingShort-chat version comparison: fact 2.2× output gain. Repeated image: fact 3.5× faster. Long-prompt extra RAM: fact 82% lowerUseful evidence for caching and batching, but not a universal stream multiplier. LM Studio, 2026-06-05
oMLXContinuous batching, prefix-cache preservation, SSD cache, Lightning MTP and DFlashExact-model MTP: fact 2.33×–2.62× single-stream gains; deeper MTP gave est 1.50× at one stream but only est 1.06× at batch fourBest current evidence for these models. Test model by model. oMLX releases, 2026-08
vLLM-MLXCommunity vLLM-style stack with continuous batching, paged KV, prefix sharing and SSD cacheM4 Max/128GB Qwen3-30B-A3B: fact 98.1 → 233.3 aggregate tokens/s, or fact 2.38×, at fact five requestsPromising pilot option; production reliability remains unknown. Guide, accessed 2026-09-03

What batching does to conversations per box

Measured proxySingle decodeBatched aggregatePer streamReading of current est 4/box
M4 Pro/48GB, sparse 35B-A3Bfact 66.9 t/sfact 163.8 t/s at eight-wideest 20.5 t/sConservative for short decode-heavy work
M4 Pro/48GB, dense 27Bfact 15.4 t/sfact 59.4 t/s at four-wideest 14.9 t/sFair-to-conservative
M4 Pro/64GB, dense 27B with MTPfact about 20.8 t/s at 8Kfact about 120 t/s at eight-wideest 15.0 t/sConservative for short work
M4 Max/128GB, dense 27Bfact 29.0 t/sfact 217.4 t/s at eight-wideest 27.2 t/sVery conservative when the Studio runs the daily model
M4 Max/128GB, sparse 35B-A3Bfact 122.2 t/sfact 267.8 t/s at four-wideest 67.0 t/sVery conservative for decode-only work

These are throughput tests, not service-level tests. A machine can show higher aggregate output while cold prompts wait too long.

Prefix/KV reuse

Caching has no cold-request gain.

  • LM Studio measured fact 3.5× on a repeated image prompt with fact 3,584 cached and fact 145 uncached tokens.
  • An oMLX field report reduced an approximately fact 115K-token agent prefix from fact about 20 minutes cold to fact about 43 seconds warm; the arithmetic ratio is est about 28×. oMLX field report, 2026-08-18/20
  • Exact course-material capacity gain is unknown until UVU measures reusable-prefix share, hit rate, restore cost and session-affinity failures.

Quantization

Official MLX results below used the same Qwen3-30B-A3B model on M4 Max/64GB. Quality is MMLU-Pro, not the Artificial Analysis Index. MLX benchmark, build dated 2025-10-08

PrecisionQualityDecodeMemoryChange from q8
q8fact 72.46fact 83.16 t/sfact 33.46GBBaseline
q6fact 72.41fact 94.14 t/sfact 25.82GBest -0.05 quality point, est +13.2% decode, est -22.8% memory
q4fact 70.71fact 113.33 t/sfact 18.20GBest -1.75 quality points, est +36.3% decode, est -45.6% memory

Recommendation: use q4 for the high-volume daily lane only after course tests show acceptable answers. Use q6 where writing, code correctness or nuance warrants the small measured quality advantage. The stream-count multiplier from bit width alone is unknown.

Service levels

All targets are est proposed service levels, not current UVU commitments.

ClassTTFT or first-result targetQueue budgetSpeed target
Tutoring/chatest p95 ≤10 secondsest p95 ≤2 secondsest ≥15 output t/s
Coding helpest p95 ≤20 secondsest p95 ≤3 secondsest ≥15 output t/s
Document questionsest upload acknowledgment ≤2 seconds; cold answer est ≤30 seconds; warm answer est ≤10 secondsest p95 ≤5 secondsest ≥12 output t/s
Agentic sessionest acknowledgment ≤2 seconds; first model action est ≤30 seconds; progress every est 15 secondsest start wait ≤10 secondsest ≥10 output t/s per model call
Research writingest p95 ≤30 secondsest p95 ≤5 secondsest ≥12 output t/s
Lecture transcriptionest acknowledgment ≤2 seconds; start est ≤30 seconds; finish est ≤0.5× recording lengthAsyncText t/s not applicable
Image taskest acknowledgment ≤2 seconds; first result est ≤60 secondsest start ≤15 secondsText t/s not applicable

Replacement for the fixed 30-second service time

Use:

slot occupancy=prefill and scheduling time+ {output tokens{decode rate

For planning, model each class as an est lognormal service-time distribution:

ClassEST meanEST coefficient of variationEST p95
Tutoring/chat24 seconds0.6051.2 seconds
Coding help66 seconds0.70152.8 seconds
Chunked document QA94 seconds0.80233.4 seconds
Agentic session536 seconds1.001,490.7 seconds
Research writing104 seconds0.80258.3 seconds
Transcription scenario input225 seconds0.70520.8 seconds
Image scenario input45 seconds0.5087.5 seconds

The supplied finals model gives est 6,534.553 requests/hour. An est finals class mix of chat 57.89%, coding 15.79%, documents 12.63%, research 8.42%, speech 2.11%, and image 3.16% produces:

  • est weighted mean service = 51.11 seconds.
  • est offered load = 92.76 occupied streams.
  • The old est 30-second assumption produces only est 54.45 streams, understating this class mix by est 70.4%.

Small queue simulation

Method: est Poisson arrivals, class-specific lognormal service, one shared first-come queue, est 500 seeded one-hour runs. The method follows standard multi-server queue concepts. MIT queueing notes, Spring 2026

Because the brief does not define \(k\), this report defines it as est the number of Studios replacing minis in a fixed 26-node fleet. When every node runs the same approximately est 30B daily model:

EST capacity=4(26-k)+8k=104+4k

MixEST capacityEST utilization without agentsEST median-run p95 queue waitEST jobs queued after hour
est k=0: 26 minis104 streams89.2%4.6 seconds0
est k=4: 22 minis + 4 Studios120 streams77.3%0.0 seconds0
est k=8: 18 minis + 8 Studios136 streams68.2%0.0 seconds0
est k=13: 13 minis + 13 Studios156 streams59.5%0.0 seconds0

The est 26-mini case is stable, but its est 4.6-second p95 queue delay plus the measured-proxy fact 8.6-second short-prompt TTFT gives est 13.2 seconds, missing the proposed est 10-second chat target.

Finals with 5% starting agents in the peak hour

est 5% × 6,123 = 306.2 agent sessions/hour.

At est 536 seconds of slot occupancy, agents add est 45.59 occupied streams, taking total offered work to est 138.35 streams.

MixEST utilizationEST median-run p95 waitEST median queued after hour
est k=0133.0%965.0 seconds or 16.1 minutes1,469
est k=4115.3%400.3 seconds or 6.7 minutes701
est k=8101.7%63.8 seconds100
est k=1388.7%0.0 seconds0

For est k=0, est k=4, and est k=8, utilization exceeds fact 100%, so no steady state exists; waits continue growing if the peak continues.

If est 306.2 daily agent users spread evenly over an est eight-hour day, agents add only est 5.70 streams. The est 26-mini fleet reaches est 94.7% utilization, but median-run p95 wait still rises to est 14.9 seconds.

If Studios are fixed to the plan’s est four-stream 120B role, replacing minis does not increase nominal capacity. No tested static \(k\) mix serves both the interactive queue and the concentrated agent queue. The service needs dynamic routing plus cloud overflow, not only a fixed hardware ratio.

What fails first

  • 1. Cold document TTFT fails before total decode capacity: a mini proxy already measured fact about 75 seconds at an 8K prompt.
  • 2. Agent admission makes the shared queue unstable.
  • 3. Chat then waits behind multi-minute work.
  • 4. KV memory and fresh prefill become the agent bottleneck.
  • 5. A rigid Studio-only heavy lane can starve the mini chat lane.

Agentic and coding loads

A controlled SWE-bench study gives the cleanest chat comparison:

WorkloadAverage total tokens/taskInput/output ratio
Code chatfact 3,390 tokensfact 1.33
Agentic codingfact 4.17M tokensfact 153.85

That is fact (derived) 1,230×; the authors report approximately fact 1,200×. The same task varied by as much as fact 30× across runs, and higher token use did not reliably improve success. How Do AI Agents Spend Your Money?, revised 2026-04-29

The production GitHub trace adds:

  • fact (derived) 7.04 user turns/session.
  • fact mean 6.6 model calls/user turn.
  • fact 582,500 prompt tokens/user turn, of which fact 545,300 were cached.
  • fact (derived) 37,200 fresh prompt tokens/user turn.
  • fact (derived) about 15.7× more fresh-prefill work if cache reuse is lost.
  • fact mean 396.3 seconds/turn.

For the supplied est 6,123-user staff case:

  • Agent users: est 306.2.
  • Agent turns/day at the production rate: est 1,299.98.
  • Agent token work: est 614.3M–762.4M tokens/day.
  • Existing demand-model baseline: est 26.36M tokens/day.
  • Added agents: est 23.3×–28.9× the entire baseline.
  • Combined work: est 24.3×–29.9× baseline.

The agent paper contains conflicting aggregate and median values. Its exact median token count is therefore unknown; the range above uses its aggregate-derived lower case and Table 4 mean upper case.

Recommended policy:

  • Cap locally admitted agent jobs.
  • Meter cumulative input, cached input, output, model calls and elapsed time separately.
  • Preserve session affinity and cache state.
  • Give agents their own queue.
  • Route overflow and very long jobs to the metered frontier pool.
  • Start with an illustrative est one 64K-input plus 8K-output local agent allowance per agent-active user-day. At est 306.2 users, that is est 22.05M tokens/day, or est 0.84× the existing baseline. Change the cap only after measured local traces.

Audit of the plan’s number

The audited number is in data.js: est four simultaneous streams at at least 10 t/s.

Replacement counts below are est admission caps pending a delivered-M5 acceptance test. Studio counts assume the daily est 30–35B model unless noted.

ClassCurrent 4/box judgment: 48GB mini64GB mini128GB StudioReplacement
Tutoring/chatConservativeConservativeConservative with daily model; fair with heavy modelest 8 / 8 / 8
Coding helpFair short; optimistic at repository scaleFairConservative short; fair longest 4 / 4 / 8
Document QAOptimistic cold; fair warmOptimistic cold; fair warmFair when chunked/warmest 2 cold or 4 warm / 2 cold or 4 warm / 4
Agentic sessionsOptimisticOptimistic, but extra RAM helps KVOptimistic or unknown with 120B modelest 1 / 1–2 / 2–4
Research writingFair-to-optimisticFairFairest 2 / 4 / 4
Lecture transcriptionInvalid unitInvalid unitInvalid unitunknown; separate ASR benchmark
Image tasksInvalid unitInvalid unitInvalid unitunknown; separate vision/image benchmark

Even if concurrency stays at est four, one mini’s request rate varies from:

  • Chat: est 600 requests/hour.
  • Coding: est 218.2/hour.
  • Document QA: est 153.2/hour.
  • Research writing: est 138.5/hour.
  • Agent sessions: est 26.9/hour.

That is an est 22.3× span hidden by the phrase “four conversations.”

What the plan should change

  • 1. Replace the single stream count with the class table above. Size the fleet from \(\sum \lambda_jE[S_j]\), p95 TTFT and per-user decode speed.
  • 2. Create separate queues for interactive text, documents/research, agents/API, and speech/image. Protect chat from long work.
  • 3. Cap and meter agents from launch. Spill excess jobs to the frontier pool before projected utilization reaches est 85%.
  • 4. Index documents once. Use retrieval, exact-prefix caching and box affinity. Do not replay whole files.
  • 5. Require a mixed-load acceptance test on est four delivered M5 nodes after fact 2026-09-22: cold and warm prompts, est 1/2/4/8 concurrency, p50/p95 TTFT, per-user and aggregate output, cache reuse, memory, thermals and failures.
  • 6. Treat continuous batching as a measured curve. Do not multiply batching, cache, speculation and quantization gains.
  • 7. Start with q4 for the daily lane and q6 for quality-sensitive work, subject to UVU course evaluations.
  • 8. Collect per-class UVU telemetry before moving from the pilot to the est 26-node fleet.

Source log

NOT_RUN

  • not run Physical M5 Pro mini or M5 Max Studio benchmark: hardware availability begins fact 2026-09-22.
  • not run UVU trace replay: no class, token, upload, cache, queue or peak trace was supplied.
  • not run Four-node acceptance test.
  • not run Token-scheduler-level continuous-batching simulation; virtual streams are an approximation.
  • not run ASR and image-generation benchmarks on the planned boxes.
  • not run Cloud-valve quota, privacy, latency or failover test.
  • not run Like-for-like Artificial Analysis quantization comparison; official task benchmarks were used instead.
  • not run Academic-research helper scripts because their local requests dependency was unavailable and the workspace prohibited installation.
  • not run Contact with UVU or any vendor, as instructed.
  • unknown Actual UVU class mix, student-agent behavior and cache-hit distributions.
  • unknown Exact M5 chassis stream counts and exact 120B Studio capacity.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/39-service-engineering.md in the research pack.

S

Duty of care and resilience

Minors, health data, crisis disclosures, security controls, single points of failure, recovery targets, and the incident runbook — from the statutes and named campus examples

Appendix S in five lines

  1. QuestionWhat must be ready before the service handles minors, patient data, crisis messages, or a system failure?
  2. AnswerMinors and real patient data stay out until privacy, human help, security, and recovery checks all pass.
  3. Deciding numbers18,163 concurrent-enrollment studentsfacteleven named conditionsestseventeen security gapsest
  4. What the plan doesStart with practice health cases, block outside tools and web searches, provide a real human crisis path, and test backup routes and the incident guide.
  5. Still unknownStudent ages and the live system design are UNKNOWN; legal review, system checks, and attack, failover, restore, and crisis drills are NOT_RUN.

What happens when a 16-year-old concurrent-enrollment student uses the tutor, when a nursing student pastes a patient's chart, when someone types that they want to hurt themselves, when an uploaded file carries hidden instructions, or when the router dies during finals? The plan had a routing rule and an isolation test; it did not answer those questions. This appendix does, from the statutes, federal guidance, and named campus examples. Research cutoff September 3, 2026. It is planning research, not legal advice. fact dated law and guidance · est design controls and targets · unknown what only UVU's counsel or records can settle.

What this changes in the plan

A. Minors and real health data stay closed until specific gates pass. A concurrent-enrollment student controls their own UVU record even under 18 (FERPA), while survey-type questions stay under parental rights until 18 (PPRA); Utah's student-data law governs the school-district side; COPPA reaches only under-13s. Phase two opens only when eleven named conditions pass — parental permission alone is not one of them. Placement patient records never enter the service (the clinical host's HIPAA duties continue); UVU's own clinics and counseling are FERPA records; the service uses synthetic cases, and real patient data fails closed at the router. UVU's public pages disagree about whether its own health records are FERPA or HIPAA — a question for UVU, not a finding of fault. B. A crisis handoff is built before broad student use. The bot is not an employee, a therapist, or a campus security authority, so Title IX, Clery, and Utah duty-to-warn do not attach to it — they attach to the trained person who receives the alert (and Utah's child-abuse reporting duty attaches to that reviewer). Design: a separate detector, an honest message ("I'm an AI tutor, not a crisis service… call or text 988"), an acknowledged human queue rather than an unread email, a restricted case record, graduated handling for low-confidence signals, and no discipline on classifier output alone. C. The first pilot removes the dangerous features: no tools, no web fetching, no shared document retrieval, no automatic cloud spill, outbound network denied from inference and parsing workers. D. "Content logging off" becomes a verified data map covering histories, caches, retrieval stores, telemetry, backups, support access, incident capture, retention, and deletion. E–I. The control plane gets finished — redundant routers, database and queue recovery, identity-outage behavior, model mirrors, a same-class spare for the one Ultra, tested recovery targets (a failed mini: 5 minutes; total local loss for public data: 30 minutes via the metered route; the restricted route: fails closed, 4 hours), a change freeze from seven days before finals through the grade deadline, a model admission record with pinned hashes and licenses, named owners (a service owner, a cross-trained backup, central IT after hours), and an incident runbook that is a launch gate. Twenty controls are graded: three exist in the plan, seventeen were gaps, each now with a fix and a cost class. est

Research date: FACT — September 3, 2026 MDT.

Statute, policy, model, section, emergency-service, and source-log numbers are identifiers. Every measured quantity below is marked fact, est, or unknown.

Minors

Who is in scope

fact — 18,163 concurrent-enrollment students, dated September 2, 2026. This comes from the supplied engine assumptions. Their age distribution is unknown. The plan’s decision to keep them out of the first phase is sound.

What the laws require

FERPA

A student attending UVU controls their UVU education records even when the student is younger than eighteen. The parent keeps FERPA rights over the high-school record and can inspect UVU records that UVU sends back to the school. UVU may disclose records to a tax-dependent student’s parent under a FERPA exception, but it is not required to do so. A parent’s concurrent-enrollment permission is not FERPA permission to inspect the student’s UVU chat record. U.S. Department of Education dual-enrollment guidance, updated August 2026.

An AI operator receiving education records without student consent must fit FERPA’s school-official exception: it must perform a real institutional function, remain under UVU’s direct control over record use and maintenance, use the information only for the approved purpose, observe redisclosure limits, and meet UVU’s annual-notice criteria. A privacy statement alone is insufficient. FERPA school-official guidance.

PPRA

fact — federal rule current through 2026: PPRA rights normally remain with a parent until the student is eighteen or emancipated. Postsecondary attendance by itself does not transfer those rights.

A normal open-ended tutoring chat is not automatically a PPRA survey. Risk arises when UVU or a participating school requires the bot to ask, analyze, or evaluate protected subjects such as mental health, sexual behavior, illegal or self-incriminating conduct, religion, family relationships, privileged relationships, or income. Required activities can require prior written parental consent; some other school-administered activities require notice and an opt-out. Each intake form, tutor script, wellness check, evaluation, and research use must therefore receive a PPRA classification. Department of Education PPRA guidance.

Utah Student Data Privacy Act, Title 53E Chapter 9

The main “education entity” definition covers Utah’s state and local public-school bodies, not UVU. It nevertheless applies to the school-district side of concurrent enrollment. Among other things, it governs parent notice, optional-data consent, disclosure, contractor controls, and protected-topic examinations or surveys.

The Utah law requires parental consent for specified disclosures of an under-eighteen pupil’s defined educational data and creates a process for school-to-higher-education data sharing. UVU and each participating school must agree which organization owns each record and which rule controls it. Current Utah Code Title 53E Chapter 9.

UVU also has direct duties under Utah’s higher-education student-data law. Those duties cover contracted purposes, collection, storage, sharing, deletion, audits, secondary uses, affiliates, advertising, and sale. Utah Code Title 53H Chapter 14 Part 5.

COPPA

fact — federal threshold current September 3, 2026: COPPA concerns covered online operators collecting personal information from children younger than thirteen. It does not cover every minor. A commercial AI vendor may still be covered even if UVU itself is outside the FTC’s ordinary nonprofit jurisdiction.

A school may consent for a covered operator only for a school-authorized educational use, with no separate commercial purpose. The operator remains responsible for notice, security, collection limits, review, deletion, and stopping further use. FTC COPPA guidance.

What peer institutions do

fact — 2025–2026 practice: Utah’s concurrent-enrollment process uses parent permission for participation, while UVU’s FERPA-protected college records remain controlled by the student. Weber State follows the same student-owned-record approach. Existing concurrent-enrollment permission should not be treated as consent for AI prompt processing. UVU concurrent-enrollment information and Weber State parent information.

Salt Lake Community College uses a student-and-guardian agreement while separately telling families that college FERPA rights belong to the student. The University of Arizona and Worcester Polytechnic Institute likewise state that college records belong to the student regardless of age. Arizona also tells users to use vetted enterprise AI, remove personal information, and avoid putting student records into external AI tools. SLCC agreement, Arizona FERPA guidance, and WPI FERPA guidance.

Age-appropriate design

These are est design controls, based on UNICEF’s fact December 2025 and June 2026 guidance, not separate Utah legal mandates:

  • Give recurring, plain-language notice that the service is AI.
  • Use maximum privacy settings by default.
  • Collect an age band rather than a full birth date unless exact age is necessary.
  • Do not present the tutor as a therapist, friend, confidant, or exclusive relationship.
  • Do not use emotional dependency, streaks, guilt, or re-engagement nudges.
  • Keep sexual, violent, manipulative, and self-harm testing in the release gate.
  • Give an equal non-AI route without penalty.
  • Explain before use that narrow safety content may be sent to trained people.
  • Let students see, correct, export, and delete records where applicable.
  • Write notices separately for students and parents; do not hide them inside general terms.

UNICEF Guidance on AI and Children.

Exact conditions for opening phase two

These are est launch conditions. All must pass:

  • Population gate: Block children younger than thirteen unless a documented COPPA path passes. Confirm vendor minimum-age terms.
  • Rights-holder matrix: State when the student, parent, school district, or UVU controls consent and access. Do not give parents automatic access to UVU chat records.
  • Separate permission: Obtain clear student opt-in and the promised parent or guardian permission. State that this permission does not transfer the student’s FERPA rights.
  • Data map: List every prompt, output, identity field, history item, safety event, retrieval item, cache, log, backup, and cloud transfer.
  • PPRA screen: Remove or separately approve required questions that function as protected-topic surveys, analyses, or evaluations.
  • Contract pass: Prohibit model training, targeted advertising, sale, profiling, unrelated product improvement, and unauthorized redisclosure. Name subprocessors, deletion, audit, breach, access, and export duties.
  • Age-safe product pass: Test the exact interface and model for manipulation, sexual content, self-harm, bias, jailbreaks, false crisis alerts, and age-gate bypass.
  • Health exclusion: Prohibit real patient information, counseling notes, disability files, and real clinical cases.
  • Logging proof: Prove prompt text is absent from ordinary logs, caches, backups, telemetry, and support paths while keeping useful security metadata.
  • Human safety route: Prove that a serious alert reaches an acknowledged trained person rather than an unattended form or email.
  • Governance: Obtain written approval from UVU counsel/privacy, Registrar, CISO, accessibility, the institutional data owner, and each participating school body.

Phase two cannot open on parental consent alone.

Health data

When HIPAA applies

Health-related words do not make a record subject to HIPAA. HIPAA applies when a covered health plan, clearinghouse, covered provider, or business associate creates, receives, maintains, or transmits protected health information in the covered role.

ContextGoverning boundary
Synthetic or properly de-identified teaching caseUsually not HIPAA PHI. FERPA, ethics, license, and re-identification controls may still apply.
Real patient information from a clinical-placement hostThe host’s HIPAA duties normally continue. Placement access does not authorize copying the patient’s information into UVU’s AI.
UVU student clinic or student counseling record maintained by UVUGenerally a FERPA education or treatment record excluded from HIPAA PHI.
Nonstudent patient treated by a covered university clinicNormally HIPAA.
University hospital treating people without regard to student statusNormally HIPAA.
Student disability-services recordNormally FERPA plus disability-law confidentiality and need-to-know controls.
Employee accommodation recordADA confidentiality applies; HIPAA depends on where the record came from and which entity maintains it.

HHS expressly says postsecondary student clinic and counseling records are generally on the FERPA side, even when the university is also a HIPAA covered entity. Nonstudent clinic records can remain HIPAA records. HHS FERPA/HIPAA clinic guidance.

UVU’s public pages create an unknown that must be resolved: Accessibility Services says its disability records are FERPA records, while Student Health says it operates under HIPAA. This may reflect voluntary practices, a hybrid-entity boundary, a covered transaction, or public wording; the reviewed evidence does not settle it. Do not label UVU noncompliant.

Clinical placements

Students may see patient records under the clinical host’s approved training workflow. That does not authorize transfer into a campus AI service.

Until each host approves otherwise:

  • Students must not paste or upload patient information.
  • UVU should use synthetic cases.
  • Real patient data must fail closed at the router.
  • A host-specific use would need authorization, minimum-necessary controls, security review, and possibly a business associate agreement.

HHS guidance on trainees’ access to patient information.

Business-associate question

A business-associate question arises when the service performs work for a HIPAA-covered component or clinical host and creates, receives, maintains, or transmits PHI.

  • If an external AI or cloud provider processes PHI for that covered entity, it is generally a business associate or subcontractor and needs an appropriate agreement.
  • A cloud provider storing encrypted PHI can still be a business associate even when it lacks the decryption key.
  • If UVU components are inside one legal covered entity, the issue may instead be the university’s hybrid-component boundary and internal safeguards.
  • If the record is a FERPA student record rather than HIPAA PHI, the HIPAA business-associate concept generally does not control. FERPA and Utah contract controls still do.

HHS now gives a third-party AI chatbot operating on a patient portal as an express business-associate example. HHS Business Associates, reviewed July 30, 2026.

What “content logging off” solves

It helps by reducing stored copies of conversations.

It does not solve:

  • Whether the disclosure was authorized.
  • Prompt and output data in memory.
  • Browser history, screenshots, downloads, or device traces.
  • Retrieval stores, embeddings, caches, queues, backups, and crash reports.
  • Vendor telemetry, support access, or model-training use.
  • Cross-user leakage.
  • Role-based access, minimum necessary use, or deletion.
  • Required business-associate terms.
  • Security-event logging and incident evidence.

The plan should replace “content logging stays off” with a tested data-retention statement. Ordinary logs should omit raw content but retain request identifiers, pseudonymous identity, time, route, model and configuration hash, access decision, tool decision, error, safety event, and administrative action.

Crisis and duty of care

What is required

Title IX

fact — current federal position reviewed January 29, 2025: the federal government is enforcing the 2020 Title IX rule after the later rule was vacated. Federal “actual knowledge” generally requires notice to the Title IX coordinator or another university official with authority to act. An unread disclosure to a general tutoring bot is an unknown legal trigger.

UVU’s own policy is broader for people: an employee aware of covered conduct must report within fact twenty-four hours under the policy last reviewed October 9, 2025. The bot is not an employee. UVU must name the human who receives and acts on alerts. Federal Title IX status and UVU Policy 162.

Generic self-harm or violence is not automatically Title IX. The disclosure must involve covered sex discrimination, sexual harassment, or a listed offense.

Clery

Clery is not a general crisis-chat law. It applies to defined crimes, geography, and reports received by campus police, security, designated reporting offices, or officials with significant campus responsibility. Timely warnings and emergency notifications have different threat tests. Self-harm alone is not a listed Clery crime. A tutor bot is not automatically a campus security authority. Current Clery regulation.

A trained person must decide whether an alert creates a Clery crime-report, warning, emergency-notification, or no-Clery outcome.

Utah duty to warn and abuse reporting

Utah’s named duty-to-warn rule applies to specified licensed therapists when a client communicates an actual threat of physical violence against an identifiable victim. The therapist discharges the statutory duty by making reasonable efforts to warn the victim and notifying law enforcement. A tutoring bot is not a statutory therapist. Utah duty-to-warn statute.

fact — Utah law effective May 1, 2024: a person with reason to believe that a child is being or has been abused or neglected must report immediately to child protection or law enforcement, subject to narrow exceptions. A software classifier is not the reporting person. A human reviewer can acquire that duty after seeing the disclosure. Utah child-abuse reporting law.

FERPA emergency disclosure

FERPA permits disclosure to appropriate parties when necessary to address an articulable and significant threat. When UVU relies on that exception, it must record the threat basis and recipients. It is not a blanket reason to circulate every flagged chat. FERPA emergency-disclosure record rule.

Safe messaging pattern

est design based on NIMH, SAMHSA, and UVU crisis guidance:

I’m sorry you’re dealing with this. I’m an AI tutor, not a crisis service. If you or someone else may be in immediate danger, call 911 or UVU Police now. For suicide or emotional crisis, call or text 988. You do not need to repeat the details here.

The bot should:

  • Stop ordinary tutoring after a serious flag.
  • Use calm, direct, nonjudgmental language.
  • Ask only the minimum needed to distinguish immediate danger from nonacute distress.
  • Offer a trusted person and trained human help.
  • Never promise confidentiality, active monitoring, dispatch, or a response time unless each promise is true.
  • Say that it is sending an alert only after an acknowledged route exists.
  • Avoid long disclaimers while a person may be in danger.

UVU already distinguishes immediate emergencies, crisis support, and nonemergency Behavioral Assessment Team reports. Ordinary email or an unacknowledged form is not an acute handoff. UVU crisis services.

What campus bots do today

  • UNCW Sammy — FACT, checked September 3, 2026: says it is not for emergencies, directs users to emergency and crisis services, warns against sharing private information, and can connect unanswered ordinary questions to a person. Its crisis-alert backend is unknown. UNCW Sammy.
  • University of Houston Wayhaven — FACT, launched July 23, 2025: is described as a bridge to care, not a clinician. The vendor says it screens messages, interrupts normal coaching, shows crisis resources, and can alert designated contacts. Houston’s exact alert settings and acknowledgement path are unknown. University of Houston Wayhaven.
  • NJIT Charlie — FACT, published August 10, 2026: uses a platform whose default safety process emails designated contacts after certain risk flags. NJIT’s exact configuration and after-hours coverage are unknown. Email delivery alone is not human acknowledgement. NJIT Charlie.

Router behavior: detect → human handoff → record

Detect — EST

Use a separate safety gate, not the tutor’s ordinary judgment alone. Classify imminent self-harm, violence toward others, child or vulnerable-adult abuse, sexual misconduct, and nonacute distress separately. Combine rules with a tested classifier. Do not use keywords or one model alone.

Human handoff — EST

Send the smallest useful packet to an acknowledged human queue:

  • Identity and contact details if known.
  • Exact triggering text.
  • Time and service route.
  • Category and classifier version.
  • The message shown to the student.
  • Delivery receipt and acknowledgement state.

A trained person—not the model—decides clinical risk, police contact, Title IX, Clery, abuse reporting, and any outside disclosure. Immediate-risk routing needs a genuinely staffed destination. Lower-risk concerns can go to a normal student-support queue.

Record — EST

Keep a restricted safety-case record containing the trigger, response, recipient, acknowledgement, action, legal basis for any outside disclosure, and closure. Keep it separate from tutoring analytics. Do not retain the entire chat merely because one part created a case.

False-positive costs

No validated public precision or false-positive rates were found for the named campus systems. Those rates are unknown.

Likely costs are est:

  • Student fear, shame, or loss of trust.
  • Disclosure of private text to more people.
  • Unnecessary police or emergency involvement.
  • Alert fatigue that delays real emergencies.
  • Unequal flagging of dialect, disability, quoted coursework, or culturally different speech.
  • Students avoiding tutoring or counseling.

Use graduated handling: offer private resources for low-confidence signals, ask a short clarifying question for ambiguous risk, and reserve automatic human alerts for defined serious cases. Never use classifier output alone for discipline.

Security

Existing controls and present gaps

“Existing” below means present in the plan or supplied findings, not proven in a live deployment.

ControlStatusWhat it coversGap and concrete fix
Isolation canary batteryExisting design; live proof not runCross-user leakage around caches and concurrent sessionsExtend testing through UI history, databases, retries, restarts, timeouts, node loss, RAG, and cloud failover. Drain the route on any canary leak.
Routing by data classExisting design; live proof not runKeeps sensitive and restricted requests on approved routesEnforce the decision outside the model, show the destination before submission, use egress allowlists, and fail closed when identity or classification fails.
No content loggingExisting intent; actual state unknownReduces stored prompt and response copiesInventory databases, histories, caches, traces, backups, crash reports, and support access. Prove raw text is absent while retaining useful security metadata.
Uploaded-document isolationGapMalicious instructions, malware, parser exploits, and oversized archivesQuarantine files; allowlist necessary formats; verify real type; cap compressed and expanded size; scan or disarm active content; parse in a no-egress sandbox.
Retrieval isolationGapPoisoned documents and cross-course or cross-user retrievalHash every source, record provenance, attach access rules to every chunk, separate indexes by data class, and recheck access at query time.
Tool controlGapExfiltration or unauthorized action through connectorsLaunch tutoring without tools. If tools are later added, use narrow read-only functions, per-user credentials, server-side authorization, parameter validation, and human confirmation for consequential actions.
Model supply chainGapAltered weights, unsafe serialization, unsupported licenses, and compromised runtimesPin weights, runtime, container, prompt, and configuration by cryptographic hash. Save source, model card, license, required notices, conversion recipe, scan receipt, test receipt, and approver.
Abuse limitsGapService denial, model extraction, queue starvation, and cloud overspendApply identity-based request, upload, context, output, concurrency, timeout, retry, tool-depth, and daily-cloud limits with a global breaker. Exact thresholds remain unknown until load testing.
Insider controlsGapAdmin misuse, secret theft, hidden content access, or log deletionUse unique administrator accounts, MFA, short-lived elevation, least privilege, separate approval and operating roles, protected central logs, and a second campus approver for releases and destructive administration.
Forensic loggingGapIncident detection and reconstructionRecord pseudonymous identity, time, route, data class, model/configuration hash, access decision, document hashes, tool decision, latency, error, safety event, cloud spend, and admin action. Keep raw content off by default.

OWASP identifies indirect prompt injection, retrieval poisoning, excessive tool authority, sensitive-information disclosure, supply-chain weaknesses, and unbounded consumption as core AI-service risks. OWASP GenAI risks.

Main exfiltration paths

  • Another user’s chat, cache, history, workspace, or retrieval index.
  • A poisoned uploaded file or retrieved page.
  • A tool using broad service credentials.
  • Automatically fetched links, images, HTML, or URLs in model output.
  • Outbound DNS or web traffic from the model or parser.
  • Silent cloud overflow to an unapproved provider.
  • Logs, traces, backups, crash reports, or administrative consoles.
  • Model or system prompts containing secrets.
  • A compromised weight, runtime, package, container, or conversion script.
  • An insider with broad production and log access.

The smallest safe pilot has no external tools, no arbitrary web fetching, no shared retrieval uploads, and outbound access denied from inference and parsing workers.

Supply-chain gate

est required release record:

  • Repository owner and immutable revision.
  • Download source and file list.
  • Cryptographic hashes.
  • Signature status.
  • Exact license and required notices.
  • Model card, intended uses, and limits.
  • Serialization format and conversion process.
  • Scanner and evaluation receipts.
  • Named approvers and approval date.
  • Read-only local copy plus an independent mirror.

A matching hash proves byte identity. It does not prove safety, provenance, model quality, or legal permission. Exact license suitability remains unknown until the chosen models are reviewed.

Availability

Single points of failure

The exact deployed topology is unknown. The supplied red-team review identifies an incomplete control plane.

Failure pointResultConcrete fix
Router or ingressWhole service becomes unreachableUse redundant stateless router instances, independent health checks, and a tested failover address.
Identity providerNew sessions failFail closed for new access. Consider only a short, CISO-approved grace period for already authenticated sessions. Keep a public status page outside the login path.
The single Ultra nodeHeavy jobs stopKeep a compatible model mirror and documented degraded route. Buy or qualify a same-class spare before promising heavy-route continuity. A mini is not an equivalent replacement.
Database, queue, or cacheHistories, routing state, or jobs failReplicate only required state, keep inference nodes stateless where practical, and test restore.
Network switch or uplinkFleet becomes unreachableSeparate failure domains, use redundant network paths where available, and test a disconnected-node case.
Power or roomEvery colocated Mac stopsUse UPS-backed circuits and a warm spare in another campus location if the recovery promise requires site resilience.
DNS, certificates, or secretsService fails despite healthy MacsMonitor expiry, keep documented renewal and break-glass procedures, and test recovery.
One operatorVacations or simultaneous incidents halt recoveryCross-train a second campus operator and use central IT/security as the after-hours receiver.

What the spare policy buys

fact — plan assumption dated September 2, 2026: sparePolicy is set to est fifteen percent, with a comment describing about one spare for every seven and a minimum of one per deployment. The rounding rule and whether “spare” means cold, warm, or unused live capacity are unknown. Engine assumption.

est arithmetic:

  • One ready spare beside four active equal-capacity machines is twenty-five percent of active capacity.
  • One machine among seven equal machines is about fourteen-point-three percent of installed capacity.
  • A literal fifteen-percent margin may not equal a whole device in a small fleet.

The policy buys one-for-one protection only when the spare is the same class, configured, licensed, connected, patched, loaded with verified models, and regularly boot-tested. It does not protect against a router, identity, network, power, bad release, shared database, site, or staff failure. It also does not replace the unique Ultra node unless the spare is Ultra-class.

Fleet pattern

est recommended design:

  • Run minis active-active behind redundant routers.
  • Keep ordinary load low enough that one node can disappear without overload.
  • Store approved weights on each compatible serving node.
  • Keep a separate checksum-verified model mirror.
  • Make inference nodes replaceable from an immutable manifest.
  • Keep public-data cloud fallback separate from restricted local recovery.
  • Never send sensitive or restricted traffic to cloud merely because a Mac failed.
  • Prevent automatic replay of tool actions after a failed stream.

Backups and disaster recovery

Back up:

  • Router and data-class policy.
  • System prompts and safety messages.
  • Model, runtime, container, and configuration manifests.
  • Identity and access mappings.
  • Retrieval source inventory and access rules.
  • Security and incident metadata.
  • Required database state.
  • License and approval records.

Do not back up raw chats when the approved policy says they are not retained.

Keep an encrypted off-site copy of critical configuration and manifests, plus a separate model mirror. Test restoration into a clean machine. A backup without a successful restore test is not recovery evidence. NIST contingency guidance.

Proposed RTO and RPO

These are est targets pending a UVU business-impact review, load tests, and recovery drills.

FailureProposed target and basis
One inference Mac or router instanceest RTO: five minutes. Basis: automated health removal and redundant capacity. The in-flight request may be lost.
Total local compute loss for public dataest RTO: thirty minutes. Basis: a pre-approved metered route with tested credentials and budget.
Restricted local routeFail closed; est RTO: four hours. Basis: ready same-class spare and immutable rebuild. Current request may be lost.
Model or configuration compromiseest RTO: four hours; RPO: last approved manifest. Basis: clean image, pinned artifacts, and independent mirror.
Retrieval-index corruptionest RTO: four hours; RPO: twenty-four hours. Basis: daily protected index state while source documents remain authoritative.
Security-log collectorest RTO: four hours; RPO: five minutes. Basis: buffered or replicated event forwarding.
Campus-room lossest RTO: one business day only if UVU maintains a separate-site warm spare and tested restore. Without those controls, RTO is unknown.

Finals maintenance

est controls based on repair lead time, not published UVU service levels:

  • Fourteen days before finals: test node loss, router loss, identity failure, model mirror, backup restore, cloud quota, and the physical spare.
  • Seven days before finals through the grade deadline: freeze ordinary OS, model, prompt, retrieval, schema, router, and provider changes.
  • Permit emergency security changes only through the canary and rollback path.
  • Stop nonessential batch, conversion, and fine-tuning work.
  • Check spare readiness, backup freshness, capacity, provider budget, and escalation contacts each day.
  • Tune warning and overflow thresholds only after real load tests.

Incident response

One-page runbook outline

Detect

  • Receive a canary failure, route anomaly, safety alert, user report, provider notice, capacity alarm, or administrator alert.
  • Open a restricted incident record.
  • Record MDT and UTC time, symptoms, request identifiers, affected accounts and nodes, route, data class, model hash, document hashes, and actions.
  • Name one incident lead.

Contain

  • Drain or quarantine affected nodes.
  • Disable the affected upload, retrieval source, tool, model, cloud route, or account.
  • Fail sensitive and restricted traffic closed.
  • Stop abnormal metered spending.
  • Revoke affected sessions and rotate exposed credentials.
  • Preserve central logs and volatile evidence. Do not erase or casually restart a compromised system.

Notify

  • Follow UVU Policy 447 by escalating suspected or actual security incidents to the CISO/CITRM path and appropriate data owner.
  • Add privacy, General Counsel, Registrar, Title IX, Clery, campus safety, counseling, disability services, the school district, clinical host, or business associate only when the incident facts require them.
  • Do not let the local operator decide external legal notice alone.
  • For a crisis, use the acknowledged safety route immediately rather than waiting for the security-notification process.

Recover

  • Rebuild compromised machines from a known-good image.
  • Restore only verified configuration, models, and data.
  • Re-run identity, data-class routing, cross-user isolation, upload injection, retrieval access, tool denial, egress, rate-limit, logging, and fallback tests.
  • Canary the repaired route before normal traffic.
  • Record actual RTO, RPO, data affected, and remaining limits.

Learn

  • Complete a plain-language review.
  • Identify the root cause and missed control.
  • Assign an owner and due date.
  • Add a regression test where it protects the failed boundary.
  • Update notices, training, contracts, and runbooks.
  • Preserve or delete evidence under the correct records and legal-hold rules.

This follows the lifecycle in NIST incident-response guidance, April 2025.

On-call model

est — supplied operating assumption: one to one-and-a-half operating staff cannot provide safe continuous human coverage alone.

Use:

  • One named platform owner for normal operation, releases, and incident leadership.
  • One cross-trained backup for leave, drills, and simultaneous work.
  • Central IT/security as the after-hours technical receiver.
  • Existing campus safety and crisis services for immediate danger.
  • Title IX, Clery, privacy, and legal owners for decisions in their areas.
  • A published service window for ordinary faults.
  • Temporary pooled coverage around finals if UVU promises faster recovery.

The local operator may be pre-authorized to take reversible containment actions: drain a node, stop a tool, disable an upload, block cloud egress, and invoke tested failover. External notice and destructive recovery remain with the proper university authority.

How students are told

Use a public status page plus in-product and direct notices appropriate to the event.

Tell students:

  • What service is affected.
  • When it started and whether it continues.
  • What information or functions may be involved.
  • What UVU has done.
  • What the student should do.
  • Where to obtain help.
  • When the next update will appear.

Do not speculate, identify a victim, disclose crisis details, or promise that no data was affected before the review supports that claim. Make notices accessible and provide language support when needed.

What the plan should change

Ranked from highest priority:

Priority A — Keep minors and real health data closed. Open concurrent-enrollment access only after every phase-two gate passes. Continue synthetic clinical cases until the specific FERPA, HIPAA, hybrid-entity, clinical-host, and business-associate boundaries are approved.

Priority B — Build the crisis handoff before broad student use. Add the separate detector, honest user message, acknowledged human route, record rules, and tests for self-harm, violence, abuse, Title IX, and false alerts.

Priority C — Remove dangerous features from the first pilot. Begin without tools, arbitrary web access, shared document retrieval, or automatic cloud spill. Add each only after its own control and test gate.

Priority D — Replace privacy slogans with a verified data map. Change “content logging off” to an exact statement covering histories, caches, retrieval, telemetry, backups, support access, incident capture, retention, and deletion.

Priority E — Finish the control plane. Add redundant routing, database and queue recovery, identity-outage behavior, network and power failure handling, model mirrors, same-class spares, and tested RTO/RPO.

Priority F — Adopt a model admission record. Pin weights and runtimes, preserve provenance and license terms, scan untrusted artifacts, test before promotion, and retain a known-good rollback copy.

Priority G — Fund real operating ownership. Name the service owner, cross-trained backup, central after-hours receiver, crisis destinations, change approvers, and incident authorities.

Priority H — Make incident response a launch gate. Attach runbooks for cross-user leakage, cloud misrouting, poisoned retrieval, compromised weights, stolen credentials, excess spend, node failure, database exposure, and safety disclosures.

Priority I — Prove the controls. Run the isolation battery, data-class failover, upload and retrieval attacks, node and router loss, backup restoration, finals load, and incident tabletop before student launch.

Source log

Source-log ordinals are reference labels, not measured quantities.

NOT_RUN

  • not run — Contact with UVU, any school, clinical host, or vendor. None was attempted.
  • not run — Legal review by UVU counsel.
  • not run — Review of private contracts, business-associate agreements, placement agreements, or provider terms.
  • not run — Inspection of a live router, model server, identity system, database, retrieval index, logs, backups, network, power system, or cloud account.
  • not run — Cross-user isolation, prompt-injection, retrieval-poisoning, tool-exfiltration, penetration, load, failover, restore, disaster-recovery, or incident-tabletop testing.
  • unknown — Exact Mac count and classes, location, topology, model revisions, licenses, staffing assignments, supported hours, provider capacity, and current backup state.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/40-duty-of-care-resilience.md in the research pack.

T

Utah law, state policy, licenses, and peers

The AI disclosure law, the federal accessibility rule, records and retention, the state's programs, each model's license by exact checkpoint, and what Utah's peers have deployed

Appendix T in five lines

  1. QuestionWhich Utah and federal rules, state programs, model licenses, and peer work change the plan?
  2. AnswerThe service needs a clear artificial-intelligence label, a firm web-access gate, exact license checks, and proof before state aid counts.
  3. Deciding numbersApril 26, 2027 web-access deadlinefact36,629 potential campus accountsest$15,000,000 for a shared research data centerfact
  4. What the plan doesShow that the assistant is not human before and during each chat, meet web-access rules before launch, pin each model’s license, and add the state workforce course to onboarding.
  5. Still unknownThe university’s share of the $15,000,000 is UNKNOWN; legal opinions, private contract checks, and a live access test are NOT_RUN.

The plan mapped UVU's own policies and procurement path. This appendix adds what arrived from the state and Washington in 2024–2026, whether each model's license fits a public university, and what Utah's peers have actually deployed — with sources. Research cutoff September 3, 2026; planning research, not legal advice. fact dated law, policy, and public pages · est reasoned conclusions · unknown what only a counsel opinion or a private contract can settle.

What this changes in the plan

1. One universal label. Utah's AI disclosure law (effective May 7, 2025) requires disclosure when a person asks in a consumer transaction and before any high-risk interaction in a regulated occupation; whether a free campus assistant is a "consumer transaction" is undecided, so the service shows "UVU AI assistant — not a human" before and throughout every chat (the statute's safe harbor) and keeps counseling, clinical, legal, and financial advice out of the general assistant. Since May 6, 2026 a state-funded university can ask Utah's Office of AI Policy for a joint interpretation, or a 12-month regulatory mitigation agreement — useful for that one question, not a general approval. 2. Accessibility is a launch gate with a date. The federal ADA Title II web rule requires WCAG 2.1 AA; the 2026 interim rule moved UVU's deadline to April 26, 2027. Test the real chat, uploads, streaming answers, and exported documents with assistive technology; generated PDFs must be tagged. 3. Records get precise. A chat tied to a student and kept is a FERPA education record; retained chats and metadata are records under Utah's open-records law (classified, not "confidential"); retention follows content (advising records: five years after separation); legal holds can stop deletion. The promise becomes: "Chat content logging is off by default. UVU retains only the data listed in its records notice, except when a user saves a conversation or a scoped legal or records hold requires preservation." 4. Ride the state's programs. The Utah Board of Higher Education set statewide AI direction (December 2, 2025); its AI Task Force launched May 1, 2026 with Dr. Burns a named member; the free statewide AI Workforce Credential (available since July 1, 2026 to 50,000+ graduates) goes into onboarding instead of a competing certificate; the 2026 budget holds $15 million one-time for a shared AI research data center for Utah's public universities and $3 million ongoing for an AI workforce accelerator — UVU's share is unknown, so neither is counted as funding until an allocation exists. 5. Licenses pinned by exact checkpoint. Qwen3.8-27B (Apache-2.0), gpt-oss-120B (Apache-2.0), Mistral Small 4 (Apache-2.0), GLM-5.3-Flash (MIT), and DeepSeek V4 (MIT) fit; GLM-5.3 and Kimi K3 carry custom terms for counsel; Qwen3.8-Flash-Next is on hold — its Qwen Community License requires a separate license for commercial hosted services and student access may defeat the internal-use exception; "the Qwen3.8 family is Apache-2.0" was false and is corrected. Downloaded published weights sit outside export-control classification unless UVU fine-tunes at scale or gives foreign access. 6. The uniqueness claim narrows. All six Utah peers run cloud vendor tools (Copilot, Gemini, ChatGPT, Claude); UC Irvine, CSU Fullerton, and UC San Diego already run campus-built interfaces, a campus-hosted model, and a hybrid gateway. What no reviewed source documents is UVU's complete three-layer design with a university-owned Apple inference tier. Say that, and publish real adoption measures rather than "first". est

Research date: September 3, 2026, MDT.

This is planning research, not legal advice. Primary law and official government or university sources were used first.

est — 36,629 potential campus accounts. Basis: the supplied data.js adds 30,506 non-concurrent students and 6,123 employees, rounded in the brief to about 37,000. Student employees may create overlap.

1. Utah AI Policy Act and Office of AI Policy

Disclosure duties

Utah’s rules now have two tracks:

UseCurrent dutyLikely UVU effect
Consumer transactionfact — effective May 7, 2025. A supplier using generative AI in a consumer transaction must say it is AI, not human, when the person clearly asks.unknown. No Utah court or published Office interpretation located decides whether UVU’s no-charge campus assistant is part of the tuition-based education transaction. Public status and a zero-dollar user price are not clear exemptions.
Regulated occupationfact — effective May 7, 2025. A person providing services in a Utah-regulated occupation must disclose before a high-risk AI interaction and follow every professional rule.Applies if AI is used to provide licensed counseling, medical, mental-health, legal, financial, or similar professional services. Keep these services out of the general assistant.
Safe harborfact — current September 3, 2026. The disclosure enforcement safe harbor applies when the interface clearly identifies itself at the outset and throughout the interaction as generative AI, not human, or an AI assistant.Put “UVU AI assistant — not a human” above the first prompt and keep a visible label during the chat.
Penaltiesfact — current September 3, 2026. Up to $2,500 per violation and up to $5,000 per violation of an administrative or court order.Use one universal disclosure pattern instead of trying to classify each chat in real time.

“High-risk” includes collecting health, financial, or biometric data and providing personalized financial, legal, medical, or mental-health guidance likely to affect a significant personal decision. Current Utah disclosure chapter

Reasoned conclusion: Treat student-facing use as covered until Utah counsel obtains a contrary interpretation. Provide an obvious path to a person. Do not let the general assistant present itself as a licensed professional.

Regulatory mitigation and joint interpretation

fact — effective May 6, 2026. H.B. 320 expressly includes a state-funded higher-education institution within the Office of AI Policy program. UVU may apply for:

  • A joint interpretation agreement explaining how a Utah law or rule applies to the planned service.
  • A regulatory mitigation agreement temporarily adjusting an identified Utah rule through cure periods, reduced civil fines, safeguards, disclosures, or reporting.

An agreement can limit users, geography, scope, and duration. It can require safeguards, disclosures, reports, and audits. It is not state approval, and laws not expressly addressed remain in force.

fact — H.B. 320: The first term may not exceed 12 months. An extension request is due at least 30 days before the term ends, and the Office may grant up to 2 extensions. Current Office statute

Recommendation: If UVU wants certainty about “consumer transaction,” seek a joint interpretation. Use mitigation only after naming a specific Utah rule that blocks a bounded pilot. The state program cannot waive FERPA, the ADA, or other federal law.

Other 2026 change

fact — effective January 1, 2027. Utah’s Digital Content Provenance Standards Act covers a provider that produces a publicly accessible generative system with more than 1,000,000 monthly users or visitors. It requires provenance information for generated or substantially altered image, audio, and video content when technically feasible. H.B. 276

Reasoned conclusion: UVU’s planned 36,629-account EST service is below that provider threshold. The rule could still govern a large underlying model provider. Recheck if UVU later makes a public multimodal service available beyond campus.

2. ADA Title II web accessibility rule

Coverage and deadline

fact — DOJ rule published April 24, 2024. Public universities are covered, including services supplied through contractors and licensed platforms. The technical standard is WCAG 2.1 Level A and Level AA.

fact — 2026 correction. DOJ’s interim final rule, effective April 20, 2026, moved the large-entity deadline from April 24, 2026 to April 26, 2027. It moved the small-entity and special-district deadline to April 26, 2028. DOJ current fact sheet

fact — current DOJ guidance. A state university uses its state’s population, not its enrollment or city population. UVU therefore uses the large-entity deadline: April 26, 2027. DOJ first-steps guide

The extension does not suspend UVU’s existing duties to provide effective communication, reasonable changes, and equal access.

Chat interface

The production service should provide:

  • Full keyboard operation and a logical focus order.
  • Visible focus and clear labels, instructions, and error messages.
  • Screen-reader announcements for streaming answers, upload progress, and status changes.
  • Contrast, zoom, and reflow support.
  • Accessible sign-in, model selection, conversation controls, and time-limit controls.
  • Alternatives for visual, audio, and video information.
  • A real assistive-technology test using complete conversations. An automated page scan is not enough.

Uploaded documents

The upload button, instructions, accepted-file information, progress, errors, preview, extracted text, and generated answer are part of UVU’s service.

A document independently uploaded by a user may sometimes fit the unaffiliated third-party-content exception. That exception does not cover UVU’s upload interface or the summaries, conversions, and answers produced by UVU or its contractor.

PDF outputs

A newly generated PDF is not preexisting content. After the deadline, public, course, and general-service outputs normally must meet WCAG 2.1 AA.

Use accessible HTML as the primary result. If the user requests PDF, generate a tagged file with:

  • Headings and correct reading order.
  • Document title and language.
  • Alt text.
  • Accessible tables.
  • Selectable text and sufficient contrast.

A password-protected, individualized document may fit a narrow exception, but UVU must still provide the information promptly in an accessible form. A separate accessible version is not a routine escape hatch.

Section 508 overlap

fact — current September 3, 2026. Section 508 directly governs federal agencies’ information and communications technology. UVU does not become directly subject to Section 508 merely because it is public or receives federal funds. U.S. Access Board

Section 508 may enter through a federal contract or when UVU supplies technology to a federal agency. Section 504 is the broader federal-funding rule that applies to public colleges receiving federal assistance. A vendor VPAT or accessibility report is useful evidence but does not prove the complete UVU service is accessible.

3. Records

FERPA

A chat becomes a FERPA education record when it is directly related to a student and maintained by UVU or a party acting for UVU. Department of Education FERPA materials

Authenticated chats about advising, assignments, grades, progress, accommodations, discipline, financial aid, or student support are likely education records if retained.

A general anonymous chat is not automatically an education record. The answer depends on its content, whether it can be linked to a student, and whether it is maintained. Removing a name alone may not de-identify a record when the remaining facts identify the student.

Sending education-record content to a cloud model can be a FERPA disclosure even if the provider says it does not retain prompts. A provider relying on the school-official exception must perform an institutional function, remain under UVU’s direct control for record use and maintenance, observe redisclosure limits, and meet UVU’s published criteria. Department of Education school-official guidance

Local inference reduces third-party disclosure risk. It does not remove FERPA from saved history, exports, administrator access, backups, or support records.

Utah GRAMA

fact — current September 3, 2026. GRAMA covers state-funded higher-education institutions. Its record definition includes reproducible electronic data prepared, owned, received, or retained by the governmental entity. Utah GRAMA

Retained prompts, answers, account associations, saved chats, moderation events, exports, and usable logs can be GRAMA records. A vendor’s storage location does not necessarily remove them if UVU owns or controls them.

A “record” is not automatically a “public record.” It can be private, controlled, protected, or otherwise exempt. FERPA education records remain governed by FERPA. Other chats may be public unless their content supports a classification or exemption. UVU must classify, segregate, and redact them rather than promise that all chats are confidential.

Retention

unknown: No approved Utah or UVU schedule specifically titled “AI chat logs” was verified.

Retention follows the content and business purpose, not the file format. Possible current schedules include:

Possible functionPublished scheduleRetention
Educational advisingGRS-2042fact — 5 years after separation, then destroy.
Student disciplineGRS-1504fact — until the issue is resolved, then destroy.
Transitory correspondenceGRS-1759fact — until resolution, then destroy.
Transitory tracking, including website visitor informationGRS-1720fact — 1 year after final action, then destroy.
Routine state administrative correspondenceGRS-48fact — 7 years, then destroy.
Program and policy developmentGRS-1717fact — retain 3 years after final action, then transfer to the archives permanently.

These are possible mappings, not one automatic schedule for all AI data. Before launch, UVU’s records officer should map chat content, saved history, feedback, safety events, access logs, model-routing logs, exports, caches, and backups to approved series or obtain a new schedule. Utah Archives retention guidance

E-discovery

Utah civil discovery can reach electronically stored information within UVU’s possession, custody, or control. That can include retained chat content, metadata, answers, moderation records, and provider-held data when the contract gives UVU control.

Routine deletion may follow an approved schedule. Once UVU reasonably expects litigation or receives a preservation duty, it must be able to stop deletion for relevant records. Utah courts may act when a party destroys or fails to preserve electronic evidence in violation of that duty. Utah Rule of Civil Procedure 37

Meaning of “no content logging”

AreaRequired meaning
FERPAPrompts and answers are not persisted after the session unless the user saves them or an authorized record workflow requires them. This reduces maintained education records but does not authorize cloud disclosure.
GRAMAContent never retained generally cannot be reproduced as a content record. Retained metadata remains a record.
RetentionSaved chats, caches, backups, error traces, safety samples, exports, and provider copies must appear in the records map.
E-discoveryOrdinary minimization reduces existing evidence. A legal hold must preserve relevant existing and future data once a duty arises.

Replace “we never log anything” with:

Chat content logging is off by default. UVU retains only the data listed in its records notice, except when a user saves a conversation or a scoped legal or records hold requires preservation.

4. State higher-education policy

USHE and the Utah Board of Higher Education

fact — December 2, 2025. The Utah Board set statewide direction calling for human-centered AI literacy and responsible AI use in teaching, research, student services, and operations. It is strategic direction, not a detailed technical standard. USHE announcement

fact — May 1, 2026. USHE launched its statewide AI Task Force. UVU Chief AI Officer Barclay Burns is a named member. USHE task-force announcement

Its first project is an AI Workforce Credential:

  • fact — USHE dated May 1, 2026: More than 50,000 graduates from the classes of 2025, 2026, and 2027 are eligible.
  • fact: The credential became available at no cost on July 1, 2026.
  • fact: It uses a statewide online environment and involves USHE institutions, Talent Ready Utah, employers, and the governor’s Pro-Human AI group.
  • unknown: Public sources do not yet state its completion rules, platform ownership, or exact place in UVU degree programs.

UVU’s onboarding should carry this credential instead of creating a competing general AI-literacy certificate.

UETN and UEN

fact — checked September 3, 2026. UETN supplies statewide education networking, bulk purchasing, and technology training. Its public material does not establish a UETN-operated higher-education generative-AI platform. That capacity is unknown. UETN

UEN offers free self-paced courses including “AI and Student Learning” and “AI and the Writing Process,” plus a public AI toolkit. Its published programming is aimed mainly at school educators, although parts can support teacher preparation and faculty development. UEN courses

Use UETN as a network, purchasing, and training partner. Do not count it as model-hosting capacity without written evidence.

2025 legislative session

  • fact: H.B. 168 would have created an AI-in-education task force covering privacy, security, literacy, integrity, and equity. It appropriated $0 and ended “House filed,” so it did not pass. H.B. 168 record
  • fact: H.B. 265, Higher Education Strategic Reinvestment, was signed. Its implementation shifted institutional funds toward high-demand programs including applied AI and computer science. It was not a dedicated campus-chat appropriation. USHE reinvestment record
  • fact/UNKNOWN: A higher-education subcommittee recommended $2,000,000 one-time for UVU’s Applied AI Institute. A final matching award was not verified in this run, so UVU receipt is unknown. 2025 subcommittee recommendation

2026 legislative session

  • fact — budget dated March 3, 2026: $15,000,000 one-time for an Artificial Intelligence Public-Private Partnership Ecosystem, including a shared AI research data center for Utah public universities and researchers, with spending across 3 years.
  • fact — same budget: $3,000,000 ongoing for an Energy, Artificial Intelligence, and Deep Tech Workforce Accelerator.
  • unknown: The budget does not name a UVU allocation, access date, service level, or cost offset. Official 2026 budget index

No enacted standalone 2026 mandate requiring a USHE campus chatbot was located.

Governor initiatives

  • fact — March 10, 2025: UVU is already named in Utah’s NVIDIA education agreement. It offers teaching kits, workshops, accelerated-computing resources, an instructor-ambassador route, internships, and apprenticeships. Governor’s Office announcement
  • fact — February 25, 2026: The Pro-Human AI Task Force covers academic research, education, public protection, government, industry, and workforce development.
  • unknown: No open university application process was published. UVU’s practical path is through its USHE task-force seat and existing NVIDIA participation. Pro-Human AI Task Force

The shared research data center and statewide credential are the most concrete programs to pursue. Neither should be shown as UVU funding until a written allocation or agreement exists.

5. Model-license fitness

est — about 37,000 users. This is below all user-count branding thresholds found. Revenue, business-purpose, and third-party-use clauses remain separate issues.

“FIT” below covers the downloadable-weight license only. It does not approve model quality, privacy, security, accessibility, or procurement.

ModelCommercial use and thresholdsAttribution, indemnity, and other termsVerdict
GLM-5.3Custom license permits hosting, modification, derivatives, distribution, sublicensing, and sale. fact — 2026 license: A Model-as-a-Service business with affiliate revenue above $10 billion during any consecutive 12 months must pass Z.AI security review before commercial use. No user threshold.Keep copyright and license text with copies or substantial portions. No express patent grant, warranty, or indemnity.CONDITIONAL FIT. The scale is not a trigger. Counsel should decide whether the campus service is a “business” or commercial use.
GLM-5.3-FlashMIT; commercial use, modification, hosting, sublicensing, and sale allowed. No user or revenue threshold.Keep the copyright and permission notice with copies or substantial portions. No express patent grant, warranty, or indemnity.FIT. Cleaner default than full GLM-5.3.
Kimi K3Custom license permits commercial use and derivatives. fact — 2026 license: Separate agreement above $20 million affiliate revenue during any consecutive 12 months for a Model-as-a-Service business. A commercial service above 100 million monthly users or $20 million monthly revenue must show “Kimi K3” prominently.Keep the license notice. No express patent grant, warranty, or indemnity. Whether campus students count as third parties under the internal-use exception is unknown.CONDITIONAL FIT. The user threshold is not reached, but the low business-revenue trigger and internal-use wording need counsel review.
Qwen3.8-27BApache-2.0; commercial deployment, modification, and redistribution allowed. No user or revenue threshold.On redistribution: include the license, mark changes, retain notices, and pass through any NOTICE file. Express patent grant; no warranty or indemnity.FIT. This is the checkpoint named in data.js.
Qwen3.8-2.4T-A95BCustom Qwen3.8-Max license, not Apache-2.0. fact — 2026 license: Branding applies above 100 million monthly users or $20 million monthly revenue. A Model-as-a-Service or AI Work Assistant business above $50 million affiliate revenue during any consecutive 12 months needs a separate license.Keep notice and obey lawful-use and third-party-rights terms. No express patent grant, warranty, or indemnity.CONDITIONAL FIT. Do not inherit the 27B model’s Apache label.
Qwen3.8-Flash-NextQwen Community License 1.0, not Apache-2.0. Any commercial Model-as-a-Service or AI Work Assistant business needs a separate license; there is no revenue floor. fact: The 100 million-user or $20 million-monthly-revenue branding threshold also applies.Custom notice and use restrictions. Student access may defeat the internal-use exception; that is unknown.HOLD until counsel confirms noncommercial treatment or UVU obtains a separate license.
gpt-oss-120BApache-2.0; commercial use, hosting, modification, and redistribution allowed. No user or revenue threshold.Apache notice, change-marking, patent, warranty, and liability terms apply. OpenAI’s separate policy requires lawful use. No vendor indemnity.FIT, with a UVU acceptable-use layer and preserved notices.
Mistral Small 4Apache-2.0; commercial and noncommercial use allowed. No user or revenue threshold.Apache redistribution and patent terms apply. The model card also bars infringement or misuse of third-party rights. No warranty or indemnity.FIT.
DeepSeek V4 Pro-0813 and Flash-0731MIT; commercial use, hosting, modification, sublicensing, and redistribution allowed. No user or revenue threshold.Keep the MIT notice. No express patent grant, model-specific indemnity, or warranty in the repository license.FIT on license terms. Security and procurement review remain separate.

The plan’s “Qwen3.8 family — Apache-2.0” statement is false. Only the reviewed Qwen3.8-27B checkpoint uses Apache-2.0.

EAR note for downloaded weights

fact — current EAR reviewed through September 1, 2026. ECCN 4E091 Note 1 excludes AI parameters that have been “published” under EAR §734.7.

Reasoned conclusion: The publicly downloadable official weights appear to fit that published-parameter exclusion as downloaded. No repository supplied a vendor ECCN or CCATS, so a checkpoint-specific classification remains unknown.

fact — rule dated January 15, 2025. A substantially trained derivative can leave the exclusion after additional training above the greater of 2.5 × 10²⁵ operations or 25% of the training operations described in ECCN 4E091 Note 2. Ordinary inference does not cross that training rule.

The license is not an export authorization. Before foreign redistribution, foreign administrator access, or a major fine-tune, UVU should screen destination, end user, end use, sanctions, model-weight rules, encryption, and attached hardware. Do not label a checkpoint EAR99 without supported classification. Current BIS EAR Part 734

6. Peers

“Scale” means the verified eligible group or usage measure. It is not an assumed headcount.

Utah peers

InstitutionWhat is deployedLocal or cloudScale and public costDate
Utah State UniversityProtected Copilot, Zoom AI Companion, Box AI, Gemini, NotebookLM, and department-paid options for Microsoft 365 Copilot, Claude, and ChatGPT Business.Cloud; Microsoft, Zoom, Box, Google, Anthropic, OpenAI.Base tools for current students, faculty, and staff — fact; active use unknown. Base Copilot has no added user charge. Published add-ons: $30/user/month Microsoft 365 Copilot, $20 or $50/user/month Claude, and $20/user/month ChatGPT Business — fact, checked September 3, 2026. Institution total unknown.Live by September 3, 2026; original launch unknown. USU tools
BYUPublic/basic Copilot access; a course-level teaching bot built around a textbook and syllabus. No verified campus-managed general assistant.Cloud; Copilot vendor Microsoft. Course-bot vendor unknown.Students and faculty for Copilot — fact. The course bot supports about 2,500 students per year — FACT, 2025 annual report. Costs and active campus use unknown.Copilot evidence October 17, 2024; course bot began in early 2023. BYU annual report
Weber StateAuthenticated Gemini, NotebookLM, Copilot, and other approved vendor tools.Cloud; Google, Microsoft, and other vendors.Eligible students, faculty, and staff — fact; usage unknown. Published Gemini premium price $36/user/month — FACT, checked September 3, 2026; institution total unknown.Gemini and NotebookLM rollout May 5, 2025 — fact. Weber AI services
Southern Utah UniversityThor, an AI texting chatbot for main-campus undergraduates; SUU-provided Gemini appears in current course requirements.Cloud; Gemini by Google. Thor vendor unknown.Thor available to main-campus undergraduates — fact; campus-wide Gemini entitlement and usage unknown. Cost unknown.Thor and Fall 2026 course evidence current September 3, 2026; original launch unknown. Thor
Salt Lake Community CollegeBase Copilot for faculty and staff; public free Copilot for students; training-linked premium licensing.Cloud; Microsoft.Faculty and staff base access — fact; usage unknown. Base user price $0 — FACT. Premium license $209/user/year — FACT, current page; institution total unknown.Article created July 26, 2025; modified May 18, 2026. SLCC Copilot
Utah Tech UniversityUniversity-agreement Copilot Chat and Gemini; departments may purchase ChatGPT Business.Cloud; Microsoft, Google, OpenAI.Faculty, staff, and eligible students for Copilot — fact; usage unknown. Copilot has no added user cost. ChatGPT Business guidance lists $25/user/month annually or $30/user/month monthly — FACT. Institution total unknown.September 23, 2025 guidance. Utah Tech guidance

University of Utah context: fact — November 20, 2025. It launched campus ChatGPT Edu and added Gemini and NotebookLM on May 12, 2026. Students, faculty, and staff may request access. Usage and institution cost are unknown here. ChatGPT Edu announcement

Similar-size public universities outside Utah

InstitutionVerified deploymentArchitecture and vendorScale and costDate
University of California, IrvineZotGPT Chat, Gateway API, ClassChat, and no-code Creator.Campus-built interface using Microsoft Azure AI, Amazon AWS, and open-web software.Enrollment 36,621 — FACT, 2024–25. Available to all students, faculty, and staff. More than 1,000 custom bots and more than 20 public department bots — FACT, November 18, 2025. User charge $0 for Chat, Copilot Chat, and Gemini Chat; institution total unknown.Faculty/staff launch January 10, 2024; student access April 25, 2024. ZotGPT
California State University, FullertonTitanGPT plus opt-in ChatGPT Edu.TitanGPT is campus-hosted; ChatGPT Edu is cloud OpenAI.Enrollment 45,863 — FACT, Fall 2025. Both services support students, faculty, and staff. TitanGPT showed 9,603 authenticated users — FACT, October 2025 board material. Institution cost unknown.ChatGPT Edu live by April 2025; TitanGPT verified Fall 2025. CSUF technology guide
University of California, San DiegoTritonGPT with chat, documents, campus assistants, course tutors, model choice, and a shared model gateway.Started on local San Diego Supercomputer Center infrastructure; now combines approved enterprise cloud models and open models hosted on UC-controlled infrastructure.Enrollment 45,087 — FACT, Fall 2025. Available to faculty, students, and staff. Institution cost unknown.All campus and Health Sciences employees by Spring 2024; student access June 2025 — fact. TritonGPT overview

What UVU’s plan does that none of these peers documents

est — evidence-set conclusion: No reviewed public source documents UVU’s complete three-layer design:

  • 1. University-owned Apple Silicon running open models locally.
  • 2. Protected free cloud tools for ordinary work.
  • 3. A separate, centrally metered frontier-model pool.

This is not proof that no university has an unpublished version.

UVU should not claim that campus interfaces, multi-model gateways, local hosting, custom agents, document chat, or metering are new. UC Irvine, CSU Fullerton, and UC San Diego already demonstrate several of those elements. The defensible distinction is the Apple-owned inference tier and its place in the full routing and cost design.

What the plan should change

  • 1. Make accessibility a launch gate. Test the real chat, uploads, streaming answers, authentication, and exported documents against WCAG 2.1 AA. Record the fact deadline: April 26, 2027.
  • 2. Replace the broad “no logging” promise. Publish a complete data map, default deletion behavior, approved retention-series mapping, cloud-provider handling, FERPA basis, GRAMA classes, access roles, and legal-hold override.
  • 3. Add universal AI identification. Show “UVU AI assistant — not a human” before and throughout every chat. Route counseling, clinical, legal, financial, and other regulated work to people unless a separately approved service exists.
  • 4. Pin licenses by exact checkpoint. Replace “Qwen3.8 family — Apache-2.0” with “Qwen3.8-27B — Apache-2.0.” Archive each model’s repository commit, license, model card, and notice bundle.
  • 5. Hold custom-license risks. Prefer GLM-5.3-Flash, Qwen3.8-27B, gpt-oss-120B, Mistral Small 4, or DeepSeek V4 on license simplicity. Keep GLM-5.3 and Kimi K3 behind counsel review. Hold Qwen3.8-Flash-Next campus-wide.
  • 6. Tie the program to USHE. Put the statewide AI Workforce Credential into onboarding, use UVU’s task-force seat, and check future USHE guidance before production launch.
  • 7. Seek state compute terms before buying research-scale capacity. The $15,000,000 FACT shared-data-center budget may help, but UVU’s share, timing, and access remain unknown.
  • 8. Narrow the uniqueness claim and publish real adoption measures. Report eligible people, activated accounts, monthly active users, repeat users, local-versus-cloud routing, and frontier spend separately.
  • 9. Use an Office of AI Policy agreement only for a named issue. A joint interpretation is reasonable for the consumer-transaction question. Do not present the sandbox as a general compliance approval.
  • 10. Add an export and model-origin review gate. Recheck foreign access, sanctions, large fine-tunes, procurement rules, security, and exact model classification before weights leave UVU-controlled systems.

Source log

NOT_RUN

  • findings-05-governance.md: not run. The named file was not present beside the brief.
  • Vendor, university, Utah agency, and UVU contact: not run.
  • Login-only verification of model menus, entitlements, or usage: not run.
  • Contract, invoice, procurement-record, or institution-total-cost review: not run.
  • Legal opinion on “consumer transaction,” “commercial purpose,” “business,” or model-license “third party”: not run.
  • Vendor ECCN or BIS classification request: not run.
  • Accessibility audit of a working UVU interface or exported PDF: not run.
  • Review of final cloud-provider contracts, retention settings, abuse logs, backups, or legal-control clauses: not run.
  • Independent verification of unpublished peer deployments: not run.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/41-utah-law-policy-peers.md in the research pack.

U

Money angles

Five-year cost, refresh and resale, lease versus buy, who pays, a student operations team, grants and philanthropy, and the electricity rate

Appendix U in five lines

  1. QuestionWhat will five years cost, when should machines be replaced, who should pay, and can paid students help run them?
  2. AnswerCompare all five years, buy first unless a real lease quote wins, review in year three, and use central funding plus paid students.
  3. Deciding numbers$90,341 with no refreshest$122,438 with a year-three refreshest$128,233 with a year-four refreshest
  4. What the plan doesAdd the five-year cases, make year four the default replacement point, set $0 from general student fees, and put paid students under a staff lead.
  5. Still unknownA university lease quote, power bill, wage table, staffing test, and grant review are NOT_RUN; true rates and costs stay UNKNOWN.

The plan priced three years, assumed a purchase, funded centrally, and staffed with professionals. A chief financial officer will ask about years four and five, leasing, what the machines are worth when replaced, who pays, whether students can run it, and the real electricity rate. This appendix answers each with public numbers. Research cutoff September 3, 2026. fact dated source or direct arithmetic · est planning estimate with basis · unknown not public or not verified.

What this changes in the plan

1. Five-year cases beside the three-year one. For the 26-mini fleet: keep the first fleet five years, $90,341; refresh after year three, $122,438; refresh after year four, $128,233 — net of modeled resale, excluding spares and the value of the fleet still owned at the end. Three-year resale values from completed sales of the previous Apple generations: mini 50%, Studio 50%, Ultra 55%, before fees. Recommendation: a measured review gate at year three, refresh by default in year four; the machines are containers for better free models, not disposable. 2. Buy the first fleet unless a written lease quote beats it. Apple's education financing advertises rates as low as 0% for up to four years and a use-only structure; at 0% the 26-mini package ($60,424 with care) is $1,678 a month to own or $874 a month to use and return — but the 0% floor is an advertisement, not a rate UVU can budget, and Utah policy reports equipment payment arrangements over 12 months. 3. Who pays: $0 from student fees. Utah's fee rule (USHE R516) bars general fees for instruction, academic support, and administration — which is what this service is — and UVU's $334.08 semester fee has no technology line. Central IT funds the shared base; colleges and grants buy marginal nodes; showback first, chargeback only for reserved or unusually heavy use; add a college only after a signed three-year commitment covers its hardware and half its marginal staff time. 4. A student operations team under a professional owner. At $18 an hour (about $43,000 per student full-time equivalent with turnover reserve), the three staffing bands become 0.35 professional + 0.20 student ($59,052, +17% — more coverage, not less cost), 0.75 + 0.50 ($129,616, −28%), and 2.5 + 1.5 ($424,879, −26%); UVU already has credit-bearing routes (INFO 4810R/4890R, IT 4810R, TECH 2810R). The simulator and configurator now carry this as a switch. 5. Grants beyond the ones already in the plan: the NSF State and Regional AI Infrastructure Hubs solicitation (deadline November 4, 2026; $4–12 million over five years; one award per state — a Utah consortium, not a solo request); NSF IUSE for a teaching study (January 20 and July 21, 2027; up to $2 million); the Utah AI Workforce Accelerator ($3 million ongoing; request for proposals expected but not yet public); Apple's Community Education Initiative as a partnership target; and a philanthropy package shaped like the ones that worked elsewhere — named student-operations fellowships, named leadership, vendor equipment, and a written commitment to recurring operations. 6. Electricity checks out. UVU's own rate is not public; at the Utah commercial average ($0.1099 per kWh) and a facility overhead of 20%, the plan's power lines ($485 per mini, $503 per Studio, $936 per Ultra over three years) reconcile to the cent; the industry-average overhead of 52% would add about $300 per sustained kilowatt-year. est

Five-year cost

Three-year residual values

Observed September 3, 2026 fact. These are gross marketplace prices, not guaranteed university proceeds.

ClassPublic market evidencePlanning residual
MiniAn M2 Pro mini with 32 GB and 1 TB recorded a completed price of $924.95 fact against a $1,899 fact original configuration price: 48.7% fact. eBay completed listing, price history. Apple currently asks $1,989 fact, or 90.5% fact of its reference price, for another M2 Pro configuration; that is a retail asking price, not resale proceeds. Apple Refurbished50% est
StudioCompleted M2 Max base-model prices ranged from $979 to $1,249 fact against $1,999 fact at launch: 49.0%–62.5% fact. The highest-volume observed listing was close to 50% fact. Apple launch, eBay $979, eBay $999, eBay $1,120, eBay $1,24950% est
UltraCompleted M2 Ultra base-model prices ranged from $2,049 to $2,369 fact against $3,999 fact at launch: 51.2%–59.2% fact. Older M1 Ultra records ranged from 40.0% to 55.0% fact. eBay $2,049, eBay $2,308, eBay $2,36955% est

Back Market’s exact M2 Pro and M2 Ultra comparison pages had no available price unknown. Apple Refurbished prices are useful ceiling checks, but include Apple’s warranty, preparation, and retail margin.

The recommended residuals are before seller fees, shipping, damage, missing accessories, or a bulk-sale discount. Net proceeds are therefore lower by an unknown amount.

Five-year mini-fleet model

Basis:

  • 26 active minis fact scenario.
  • Education hardware price: $2,139 each fact, plan input, plus $90 each fact for the production-plan network option: $57,954 total fact arithmetic.
  • Plan AppleCare allowance: $120 per device for three years est, or $3,120 per fleet est arithmetic.
  • Shared equipment: five kits est × $1,650 est = $8,250 est.
  • Plan power: $485 per device per three years est, extended to $21,017 over five years est arithmetic for the fleet.
  • Year-four residual: 40% est, applying a 20% est haircut to the three-year mini residual.
  • Future replacement prices are held flat est.
  • Taxes, sale costs, support after AppleCare, and terminal value of the final fleet are excluded unknown.
Five-year pathNet hardware cashCare, shared kit and powerFive-year cash cost
Keep the first fleet for five years$57,954 est$32,387 est$90,341 est
Refresh after year threeTwo purchases less $28,977 est resale = $86,931 est$35,507 est$122,438 est
Refresh after year fourTwo purchases less $23,182 est resale = $92,726 est$35,507 est$128,233 est

These are cash costs, net of modeled resale where a refresh occurs. They do not credit the value of the fleet still owned at the end of year five because that fleet is a different age under each path. Its terminal value is unknown.

The table also excludes the plan’s 15% spare policy est. If “26” means active machines rather than total machines, four spares are required est arithmetic. Those add $8,916 est of hardware at each purchase cycle, plus care and power.

Refresh recommendation

Make year three a formal review gate and year four the default refresh est recommendation.

The model-progress review projects that stronger open models will keep appearing for the same memory footprint, with a central open-versus-closed gap of roughly three to six months through 2028 est. That makes the machines reusable containers for better models; it does not make every machine obsolete after three years. Large frontier models will still exceed local memory and belong in the metered cloud pool. See findings-08-model-progress-regression.md and findings-36-mac-generations-side-by-side.md.

Refresh in year three only if measured workloads fail the agreed capability, wait-time, security, or support gate. A year-four refresh has the main risk that the original three-year care period leaves a support gap; extension price and availability are unknown.

Lease vs buy

Public terms

Apple Financial Services currently advertises education financing, level payments, refresh planning, Easy Return, and financing of hardware, software, accessories, and some third-party items. Its education flyer says rates may be as low as 0% fact offer floor, payments may follow a semester or school-year schedule for up to four years fact, payment deferral may be offered, and Guaranteed Buyback may apply. Every offer remains subject to credit approval and signed documents. Apple Financial Services, education flyer

Apple’s published fair-market-value structure lets the institution return the equipment or buy it for then-current market value. The institution does not automatically own it during the use-only term. Apple FMV structure

UVU can purchase through Apple’s Utah NASPO contract 23003 fact identifier, Utah addendum PA4282 fact identifier, or PEPPM contract 576202 fact identifier. NASPO and PEPPM allow financing where the participating contract permits it, but no public UVU rate, guaranteed residual, fee schedule, or detailed lease option was found unknown. Apple Utah contracts, NASPO Apple record, PEPPM terms

Utah policy calls for reporting periodic-payment equipment arrangements longer than 12 months fact. Its exact application to UVU’s chosen structure needs procurement review unknown. Utah finance policy

Side-by-side: mini fleet

This financing comparison uses the current public package of 26 devices fact scenario, $57,954 hardware fact arithmetic, and $2,470 current public AppleCare pricing fact arithmetic: $60,424 financed fact arithmetic.

The 6% rate est sensitivity is not an Apple quote.

PathCash and total costOwnershipFlexibility
Buy$60,424 now fact scenario. Modeled hardware resale after three years: $28,977 est. Net capital and care before selling costs: $31,447 est.UVU owns immediately.Best reuse and resale freedom; UVU carries obsolescence and disposal work.
Plan-to-ownAt 0% est illustration: $1,678.44/month est, $60,424 total est. At 6% est sensitivity: $1,838.22/month est, $66,175.75 total est.Ownership timing depends on the signed schedule unknown.Smooth cash flow, but costs more if the rate is above zero.
FMV/use-onlyWith a 50% residual est, at 0% est illustration: $873.53/month est, $31,447 total est, then return. At 6% est sensitivity: $1,101.56/month est, $39,656.29 total est, then return. Buying at modeled term-end value would raise the 6% case est to $68,633.29 est.Apple or the financier owns during the term.Cleanest scheduled refresh and lowest modeled payments; UVU gives up resale upside and must meet return conditions.

External racks, switches, UPS units, storage, cables, spares, deployment, and staff are excluded unknown.

Side-by-side: $120,000 mixed fleet

The mix of minis, Studios, and Ultras is unknown, so the model uses the recommended 50%–55% residual range est.

PathThree-year cash and total costOwnership and flexibility
Buy$120,000 now fact scenario; $60,000–$66,000 residual est; $54,000–$60,000 net hardware consumption est, before disposal costs.Own immediately; greatest reuse and sale freedom.
Plan-to-ownAt 0% est illustration: $3,333.33/month est, $120,000 total est. At 6% est sensitivity: $3,650.63/month est, $131,422.77 total est.Smooth payments; eventual ownership depends on the signed form unknown.
FMV/use-onlyAt 0% est illustration: $1,500–$1,666.67/month est, $54,000–$60,000 total est, then return. At 6% est sensitivity: $1,972.78–$2,125.32/month est, $71,020.25–$76,511.38 total est, then return.Best scheduled refresh; no resale upside and return-condition risk.

Recommendation: buy the first fleet unless Apple supplies a written use-only offer with a favorable guaranteed return value and UVU values budget smoothing more than ownership. The advertised 0% floor fact is not a rate UVU can budget until it has a quote.

Who pays

Student-fee boundary

USHE policy R516 allows general fees for approved activities, programs, services, and non-instructional facilities that broadly benefit students. It bars general-fee funding for instruction, academic support, general administration, and expenses reasonably covered by tuition or state appropriations. It also requires separate accounting, annual review, a student-majority committee, a public hearing, trustee action, and Board approval. USHE R516

UVU Policy 511 follows this process. UVU’s public page says student fees may support technology, but not academic-program development, one-time funding needs, replacement of budget cuts, or replacement of grants and donations. UVU Policy 511, UVU student fees

UVU’s published semester schedule for 12 or more credits fact totals $334.08 fact:

  • Building bonds: $87.00 fact
  • Athletics: $84.26 fact
  • Student programs: $56.84 fact
  • Student center: $37.85 fact
  • Campus recreation: $33.50 fact
  • Student Life and Wellness Center: $25.84 fact
  • Transportation: $6.54 fact
  • Arts: $2.25 fact
  • Health: $0.00 fact

There is no separate technology line fact. UVU fee schedule

A separate tuition table reports $334.55 fact at 10 or more credits fact. The $0.47 difference unknown is not explained publicly. UVU tuition and fees

Conclusion: the base AI platform should receive $0 from general student fees est recommendation. As scoped, it includes academic support and central administration, so it does not fit the general-fee rules. A later, separately defined non-instructional service could have different eligibility unknown.

Peer cost-sharing patterns

  • UC Berkeley gives faculty a common allocation while allowing faculty, grants, deans, or chairs to purchase added capacity. Campus supplies the shared infrastructure and administration. Berkeley Savio
  • Yale provides standard compute without direct charges and sells optional priority capacity at $0.0049 per service-unit hour in FY2027 fact. Yale priority tier
  • Princeton’s research-software partnership normally splits eligible staffing 50% fact program term with a research partner for one to three years fact program term. Princeton partnership guide
  • USC’s condo model has research groups buy nodes, cables, and a five-year warranty fact, while the university supplies racks, network, power, cooling, room, and administration. USC condo model

Recommended funding model

  • Central recurring IT funds the shared fleet, identity, security, network, monitoring, warranty and refresh reserve, baseline frontier pool, and accountable professional owner.
  • Colleges and grants fund added nodes, specialist software, reserved capacity, and workload-driven staff growth.
  • Use showback and project limits first. Use formal chargeback only for reserved, priority, burst, or unusually heavy use after actual usage and cost data exist.
  • Do not charge ordinary teaching users per request.
  • Budget $0 from the general student fee est recommendation for the base platform.

One-line college rule: “Add a college only after a signed three-year est commitment covers all marginal hardware and 50% est of marginal operating labor; central IT retains the shared fabric, security, baseline service, and accountable owner.”

Student staffing

Pay and structure

Public Utah evidence gives these anchors:

  • A UVU skilled student project-lead posting paid $15–$16/hour fact, April 2026. UVU posting
  • Weber State lists an $11.75/hour minimum fact as of January 2025, a 20-hour weekly cap fact for regular student jobs, and 10–28 hours weekly fact for internships. Weber State
  • Utah Tech pays agency students $12/hour fact and charges departments a 25% markup fact, producing $15/hour fact arithmetic. Utah Tech
  • UVU federal work-study positions publicly range from $12.00 to $21.65/hour fact for academic year 2026–27. UVU work-study

Use $18/hour est for technical student operators. The exact UVU technical wage schedule and payroll burden are unknown.

Students may handle intake, documented health checks, basic diagnostics, inventory, staging, evaluation runs, documentation, and workshop support. Professionals retain production access, security incidents, architecture, policy, purchasing, releases, and service accountability.

Credit-bearing option

UVU already has possible routes:

  • INFO 4810R offers one to three credits fact, while INFO 4890R offers one to four credits fact for mentored research. UVU INFO courses
  • IT 4810R offers one to three credits fact for related employment with a department coordinator. UVU IT courses
  • TECH 2810R offers one to three credits fact for supervised professional experience. UVU TECH courses

Department approval remains required fact. Credit should recognize learning; it should not replace pay for scheduled operational work est recommendation.

Supervision and turnover

Northwestern uses paid graduate students for research-computing tickets, consultations, documents, and workshops, with at least 10 hours weekly fact and weekly staff mentoring. MIT has used students in an HPC, AI, and machine-learning help desk with staff escalation. UVA publicly identified 12 current and nine former student workers fact, observed September 2026. Northwestern, MIT, UVA

Planning allowance:

  • 0.10 professional FTE per four active students est.
  • 15% student onboarding and turnover reserve est.
  • $20.70 effective student hour est arithmetic, using the $18 wage est plus the reserve.
  • $43,056 per student FTE-equivalent est arithmetic, using 2,080 hours est planning convention.

The true UVU turnover rate and supervision load are unknown and should be measured during the pilot.

Effect on the plan’s staffing bands

The plan’s professional line is $144,118 per FTE annually est plan input. The proposed professional share includes supervision.

Plan bandCurrent professional-only costRecommended accountable mixRecommended annual costChange
0.35 FTE est$50,441.30 est0.35 professional + 0.20 student FTE-equivalent est$59,052.50 est+$8,611.20, or +17.1% est
1.25 FTE est$180,147.50 est0.75 professional + 0.50 student FTE-equivalent est$129,616.50 est−$50,531, or −28.1% est
4.0 FTE est$576,472 est2.50 professional + 1.50 student FTE-equivalent est$424,879 est−$151,593, or −26.3% est

At the smallest band, students add coverage rather than replace the accountable professional. At the larger bands, they absorb repeatable work and lower professional staffing needs.

At 10–15 hours per week for 30 teaching weeks est schedule, likely staffing is:

  • One to two students est for the smallest band.
  • Three to four students est for the middle band.
  • Seven to 11 students est for the largest band.

Paid summer coverage, overlapping handoffs, paired access, runbooks, and a professional on-call path are required. Actual student availability is unknown.

Grants and partnerships

The plan already covers HERFP, AI Moonshot, and NAIRR. The following are additional paths.

StatusProgramFinding and fit
OpenNSF State and Regional AI Infrastructure Hubs, solicitation 26-513 fact identifierPosted July 31, 2026 fact; deadline November 4, 2026 fact. NSF expects about 10 awards per cycle fact, with typical requests of $4 million–$12 million over five years fact and about $100 million available fact. Only one award per state or multi-state region fact is planned. This is the strongest infrastructure fit, but UVU would need a Utah or regional consortium with industry, government, philanthropy, and other universities. NSF AI Infrastructure Hubs
OpenNSF IUSE: EDU, solicitation 23-510 fact identifierNext deadlines are January 20, 2027 fact and July 21, 2027 fact. Awards range from up to $400,000 to $2 million fact, with no voluntary cost share fact. Fit: evidence-based AI-supported undergraduate STEM teaching, faculty practice, or learning outcomes—not routine platform operations. NSF IUSE
OpenDepartment of Education IES research training and methodsDeadline October 1, 2026 fact. Maximums are $800,000 fact for research training and $900,000 fact for research methodology. Fit: student training and evaluation research, not fleet purchase. Department of Education available grants
Expected but not locatedUtah AI Workforce AcceleratorUtah appropriated $3 million ongoing fact. A June 2026 fact presentation said an RFP would issue in August 2026 fact for curriculum modernization, embedded AI credentials, and expanded programs. The public RFP, award size, deadline, match, and applicant rules remain unknown. Utah update, Utah budget summary
ForecastNSF Major Research InstrumentationNSF expects a new solicitation during federal fiscal year 2026 fact, but has posted no current deadline unknown. The prior program allowed $100,000–$4 million fact for shared research instruments. A research-only shared AI instrument may fit; a general teaching service does not. NSF MRI
Not currently openNSF Campus CyberinfrastructureThe last program allowed up to $700,000 fact for campus compute and $1.4 million fact for regional compute. It is awaiting a new solicitation fact. Keep it on the watch list but budget $0 est from it now. NSF CC*
Selected partnership; application unknownApple Community Education InitiativeApple says it can provide hardware, scholarships, financial support, curriculum, and expert help. Apple reported support for more than 200 education and community partners fact as of October 2024, but no public application, award range, deadline, or UVU eligibility was found unknown. Treat it as a partnership target, not forecast cash. Apple CEI
Invitation onlyApple Scholars in AIMLSupports nominated doctoral students for two years fact with research and travel support, mentoring, and internship access. UVU invitation status and public dollar value are unknown. Apple Scholars

Philanthropy patterns

  • The University of Florida assembled an $85 million package fact: a $25 million alumnus gift fact, $25 million from NVIDIA fact, $15 million from the university fact, and $20 million in recurring state support fact. Pattern: donor, vendor, university, and state each fund a different layer. University of Florida
  • The University of South Florida received a $40 million naming gift fact for an AI and cybersecurity college, paired with a dollar-for-dollar challenge of up to $5 million fact. Pattern: named anchor plus matching campaign. USF
  • RIT received a $24 million commitment fact supporting an endowed AI institute director, scholarships, and faculty endowments. Pattern: donors favor named people and durable student programs. RIT
  • Cornell received a $10.5 million gift over five years fact to support researchers using shared Empire AI infrastructure. Pattern: fund the people and research program beside publicly backed compute. Cornell

Recommended advancement package: named student-operations fellowships, named applied-AI leadership, vendor-provided equipment or support, and a written university commitment to recurring operations. Electricity is a poor stand-alone donor proposition.

Electricity

UVU rate

UVU’s October 2024 fact emergency plan says its campuses receive power from Rocky Mountain Power and that most of the Orem main campus uses UVU’s north substation. UVU is also named in Rocky Mountain Power’s Schedule 34 clean-energy arrangement. That schedule adds contract-specific terms to the customer’s normal tariff; it does not publish UVU’s actual blended rate. UVU electricity plan, Rocky Mountain Power announcement, Schedule 34

UVU’s actual tariff, demand peaks, delivery voltage, clean-energy adders, fixed charges, and blended rate are unknown.

Rocky Mountain Power Schedule 8 currently lists:

  • $76 per month fact fixed charge.
  • $5.15 per kW fact facilities demand charge.
  • $14.90–$16.84 per kW fact on-peak demand charge.
  • 2.8070–6.2395 cents per kWh fact energy charges, before adjustments.

Public evidence does not establish that UVU’s relevant meter is billed on Schedule 8 unknown. Rocky Mountain Power Schedule 8

The best public proxy is the EIA’s June 2026 fact preliminary Utah commercial average of $0.1099/kWh fact, released August 26, 2026 fact. EIA

Check against the plan

The plan uses 26,280 hours over three years est, $0.1099/kWh fact proxy, and PUE 1.2 est.

DeviceIT drawFacility energy over three yearsCalculated costPlan line
Mini140 W est4,415.04 kWh est$485.21 est$485 est
Studio145 W est4,572.72 kWh est$502.54 est$503 est
Ultra270 W est8,514.72 kWh est$935.77 est$936 est

The plan’s power lines reconcile to the public commercial proxy after rounding. They remain estimates because utilization, actual machine draw, demand charges, and UVU’s contract rate are unknown.

Cooling and facility overhead

PUE includes cooling, UPS and distribution loss, lighting, and other support loads; it is not cooling alone.

At $0.1099/kWh fact proxy:

  • Plan PUE 1.2 est adds 1,752 kWh per sustained IT kW-year est arithmetic, costing $192.54 annually est.
  • The Uptime Institute’s 2026 fact industry-average PUE of 1.52 fact adds 4,555.2 kWh per sustained IT kW-year est arithmetic, costing $500.62 annually est. Uptime Institute
  • UVU’s cooling-only cost per sustained IT kW is unknown.

Budget wording: “Facility overhead is estimated at $193 per sustained IT kW-year est under the plan, with a $501 est industry-average sensitivity. UVU cooling-only cost is UNKNOWN.”

What the plan should change

  • 1. Use a year-four default refresh with a year-three measured review gate. This preserves one extra year of use without assuming local hardware must track every frontier release.
  • 2. Add the five-year cash cases. Show $90,341 est with no refresh, $122,438 est with a year-three refresh, and $128,233 est with a year-four refresh, plus the support and terminal-value gaps.
  • 3. Carry three-year residuals of 50% mini, 50% Studio, and 55% Ultra est. Add a separate, currently unknown, disposition-cost haircut before treating resale as spendable cash.
  • 4. Buy the first fleet unless a written lease quote beats ownership on UVU’s real terms. Require the rate, fees, guaranteed residual, return standards, early-exit terms, and ownership option before approval.
  • 5. Fund the shared base centrally. Colleges and grants pay for marginal capacity; charge only for optional reserved or heavy use. Use $0 from general student fees est recommendation.
  • 6. Keep a professional owner and add paid students underneath. The smallest band gains coverage; the middle and large bands can reduce annual staffing cost by about 28.1% and 26.3% est, subject to a measured pilot.
  • 7. Pursue the NSF AI Infrastructure Hub as a consortium, not a solo hardware request. Pair that with an IUSE teaching study, the expected Utah workforce opportunity, and a named philanthropy package.
  • 8. Replace “UVU electricity rate” with “public Utah proxy.” Keep $0.1099/kWh fact proxy until Facilities supplies a bill, demand history, tariff, and measured PUE.

Source log

NOT_RUN

  • UVU-specific Apple financing or Guaranteed Buyback quote.
  • UVU procurement, legal, finance, student-fee, facilities, payroll, or HR review.
  • Vendor, utility, university, donor, or grant-office contact.
  • Review of a live UVU utility bill, demand history, Apple lease schedule, or internal wage table.
  • Hardware inspection, seller-fee calculation, tax analysis, or bulk-disposition quote.
  • Student staffing pilot or time-and-motion study.
  • Grant application, eligibility ruling, account access, purchase, deployment, publication, or external action.
  • The excluded University of Utah figure was not used.

JSON grades: residual percentages est; lease rate unknown; student wage est; electricity rate fact public proxy, actual UVU rate UNKNOWN; fee eligibility no for the base service as scoped.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/42-money-angles.md in the research pack.

V

Assessment redesign and academic integrity

What happened to integrity after 2022, what detectors get wrong, the assessment designs that hold up at scale, policy patterns, a redesign for each college, and how to roll it out

Appendix V in five lines

  1. QuestionWhat assessment rules protect learning and honesty without letting flawed artificial-intelligence detectors punish students?
  2. AnswerLabel each key task “Secure” or “Open,” add a small human skill check when needed, and never treat a detector score as proof.
  3. Deciding numbersaverage 61.3% false-flag rate for 91 essays by non-native English writersfact59 of 63 unedited artificial-intelligence answers, or 93.7%, drew no concernfact$20,000–$40,000 pilotest
  4. What the plan doesReplace three templates with one task standard, label key work before class opens, run one redesign in each college, and measure time, fairness, and trust.
  5. Still unknownThe unsourced “91% of faculty concerned” figure has been removed from the plan; campus integrity records are UNKNOWN, and local detector and access tests are NOT_RUN.

Faculty's leading fear about campus AI is cheating, and the plan's answer had been three template policies and the phrase "lead with assessment redesign." The owner asked for depth. This appendix is the evidence-based playbook: what actually happened to integrity after 2022, what AI detectors get wrong and for whom, the assessment designs that hold up at 300-student scale and what they cost faculty, the policy patterns Utah's peers use, an example redesign for each of UVU's seven colleges, and what it takes to roll out. Research cutoff September 3, 2026; 160-plus dated sources. fact measured · est judged · unknown not settled.

What this changes in the plan

1. One assessment-assurance standard replaces "three templates." Every material assessment gets an assignment-level Secure/Open label, the allowed and prohibited AI functions, a disclosure rule, and an individual proof point where the learning outcome demands it — set at the course-shell moment, four to six weeks before term, with a report of unlabeled high-stakes tasks. 2. Detector output alone is never evidence — it cannot by itself cause a grade change, referral, interview, or sanction. The independent tests show detectors unstable across models, easy to evade, and uneven against non-native English and neurodivergent writers; several universities have switched them off. 3. The data do not show an integrity collapse. Self-reports, detector flags, referrals, and proven cases measure different things and must never be blended into one "cheating rate"; where institutions published totals, cases did not jump the way the fear predicts. The plan's "91% of faculty concerned" figure could not be traced to a survey and is removed until it can be. 4. Secure assessment stays sparse. A few oral, written, coding, or practical checks where UVU must prove individual skill; everything else open, applied, staged, and explicit about allowed help. Blanket redesign would burden honest students, adjuncts, and online learners most. 5. Each catalyst's paid deliverable becomes one redesigned signature assessment, run, measured for workload and student trust, and published as a design card — Engineering a code defense, Business a changed-fact decision task, Humanities staged research plus close reading, Education microteaching, Science a lab defense, Health a synthetic practical, Arts a provenance-rich portfolio with live critique. 6. A fair follow-up protocol: the student sees the rule and the evidence, may bring drafts and explain assistive tools, is asked questions aligned to the original task, gets written reasons, and keeps an appeal. 7. Success is measured by substantiated cases, overturns, validity, workload, access, and trust — never by detector flags or surveillance events. The $20,000–$40,000 first stage becomes a measured gate before any $150,000–$250,000 expansion. est

Research date: September 3, 2026 MDT. Evidence key: fact = dated source; est = calculation or planning assumption with its basis; unknown = evidence was unavailable, conflicting, or too weak.

Executive verdict

  • 1. Generative AI use in assessed work is now common, but the evidence does not show one clean post-2022 jump in cheating.
  • 2. The plan’s “91% concerned” figure has unknown provenance: the supplied files name no survey, date, sample, or question.
  • 3. Student self-reports, detector flags, referrals, and proven cases measure different things and must never be combined into one “cheating rate.”
  • 4. AI detectors are too unstable, too easy to evade, and too uneven across writers and subjects to prove misconduct.
  • 5. UVU should adopt the rule: detector output alone is never evidence and must never trigger a penalty.
  • 6. Assessment redesign should not mean replacing every assignment with an exam.
  • 7. UVU should use a few secure oral, written, coding, or practical checks where it must prove individual skill.
  • 8. Other work should remain open, applied, staged, and explicit about allowed AI help and disclosure.
  • 9. The course-shell build period is the right control point; every material assessment should receive an AI-use label before students enter the course.
  • 10. Start with one paid, measured redesign in each college, publish the results, and expand only where validity, workload, access, and student trust improve.

What actually happened to integrity, 2023–2026

The central finding

unknown — There is no credible national college series that measures the same misconduct behaviors, at the same institutions, with the same rules and adjudication process, before and after generative AI.

The available evidence shows three different things:

  • Students use AI much more.
  • Some students report submitting AI-written material as their own.
  • Recorded misconduct depends heavily on policy wording, detector use, reporting effort, and when cases are counted.

It does not support “AI increased college cheating by X%.”

Self-report data

EvidenceResultWhat it means
Challenge Success high-school cohortsFACT—2019 and 2023 samples: prior-month dishonest behavior ranged from 59.9% to 64.4% in three 2023 schools. Comparable 2019 school samples ranged from 61.3% to 82.7%. The researchers found no increase attributable to ChatGPT. StudyBroad high-school misconduct, not college AI cheating. Cohorts and schools differed.
Challenge Success follow-upFACT—February–May 2024: 72.06% of 4,354 students at six high schools reported at least one dishonest behavior. 24.27% reported unauthorized “AI or digital device” use, but the question cannot isolate AI. StudyHigh background misconduct persisted. It is not a clean year-to-year comparison.
Australian university surveyFACT—2024: unacknowledged verbatim GenAI copying was reported by 13.9% of 310 Western Sydney University respondents and 14.2% of 1,867 students at five other universities. Detector-using and non-detector institutions did not differ significantly, fact p=.097. StudyDirect self-report evidence of misuse; no evidence that detectors reduced it.
HEPI UK seriesFACT—2024: 53% of 1,250 undergraduates used GenAI for assessed work; 5% inserted unedited AI text. FACT—2025: 88% of 1,041 used it for assessment and 18% included AI text directly, including edited text. FACT—December 2025 survey published 2026: 94% of 1,054 used it for assessed work and 12% directly included AI text, versus 8% in 2025 and 3% in 2024 under the comparable narrow measure. 2024, 2025, 2026Assessed use rose sharply. Much of it—explanations, summaries, ideas, and editing—may be permitted.
Large US public-university surveyFACT—spring 2024: 95,513 students across 20 public research universities responded; about 37% used GenAI at least monthly. About 9% of GenAI users reported submitting generated work as their own. Science studyStrong cross-campus evidence, but still self-report and not a before/after cheating series.
ICAIFACT—public page checked September 3, 2026: ICAI reports an updated survey pilot with 840 students in 2020, but publishes no comparable 2023–2026 national AI-era rate. ICAIOld McCabe-era percentages must not be presented as current AI-era results.

Detector-flag data

Turnitin reported that, through March 21, 2024, it had screened more than 200 million submissions; more than 22 million, about 11%, scored at least 20% likely AI writing, while more than 6 million, about 3%, scored at least 80%. These are FACT—vendor-reported flag counts, not findings of misconduct. Submissions are not unique students; authorized AI use is included; repeat submissions and the population mix are undisclosed. Turnitin

The plan must never convert those figures into “11% cheated” or “3% were caught.”

Adjudicated and institutional cases

InstitutionRecorded resultInterpretation
University of HullFACT—2022/23: 85 AI-related referrals among 802 total cases; 56 were guilty and 12 remained open. FACT—2023/24 through April 25, 2024: 84 AI referrals among 389 total, with 39 guilty and 28 open. Hull responsePartial years and unresolved cases prevent a clean trend.
Deakin UniversityFACT—2024: 2,601 breach instances involving 1,535 students; 687 instances, or 26%, involved GenAI misuse. Deakin reportInstances are not unique students. Deakin warns that recording and detection changes affect the total.
UNSWFACT—2023: 166 serious GenAI referrals were reported and substantiated, while total recorded cases fell 16% from 2022. FACT—2024 closures: 1,983 of 2,154 cases were substantiated or partly substantiated; the detailed table contains 640 AI-related outcomes, while the narrative reports 530 AI cases. 2023 report, 2024 reportThe report does not reconcile the two 2024 AI totals. New categories and closure timing materially affect the result.
University of ManchesterFACT—2023/24: schools reported 457 academic-misconduct cases, versus 477 in 2022/23. Only 15, or 3%, were recorded in the new AI category. Manchester reportA new category and formal-case-only reporting make undercounting likely. It does not show a general rise.

How much is detection artefact?

A large share is unknown, but at least five artefacts are documented:

  • A new AI category creates an apparent rise from a zero category that did not previously exist.
  • Institutions count different units: referrals, assignments, breach instances, cases, students, closures, or sanctions.
  • More faculty training and detector availability produce more referrals even if behavior stays flat.
  • Permitted use can generate a detector flag without misconduct.
  • Open cases move between years; one student or submission may create several breach records.

UVU should publish separate rates for:

  • Student self-report.
  • Instructor concerns or referrals.
  • Detector-assisted referrals, if any.
  • Cases opened.
  • Cases substantiated.
  • Cases overturned or returned on appeal.

The denominator should be students or eligible submissions, not just raw cases.

AI detectors

Measured performance

  • FACT—2023 controlled test: seven detectors incorrectly flagged an average 61.3% of 91 human TOEFL essays written by non-native English writers. All seven flagged 18 of 91, or 19.8%, and at least one flagged 89 of 91, or 97.8%. The control was 88 US eighth-grade essays. Liang et al.
  • FACT—2023 controlled test: 14 tools, 54 documents, and 756 tests produced less than 80% accuracy for every tool. Only five exceeded 70%. False-positive rates ranged from 0% to 50% and false-negative rates from 8% to 100% across tools and conditions. Paraphrasing sharply reduced detection. Weber-Wulff et al.
  • FACT—2024 experiment: Turnitin identified 47 of 50 raw AI medical articles but only 15 of 50 after paraphrasing. Human reviewers caught more paraphrased articles but each mislabeled 12% of 50 human articles. Liu et al.
  • FACT—2024 live university test: researchers submitted 63 unedited AI answers through 33 false student accounts at the University of Reading. 59 of 63, or 93.7%, received no AI concern. PLOS ONE
  • FACT—2025 study: among 250 known-human neurosurgery articles, one detector flagged 30.4%, another 16.0%, and another 0% above the study threshold. Erol et al.
  • FACT—2026 study: in a balanced 192-text corpus, Originality achieved macro accuracy .69 and Turnitin .61; both had macro F1 below .55. Turnitin accuracy ranged from .86 in humanities to .51 in science. This study found no significant Turnitin difference between its EFL and professional-writing samples, fact p=.50. Hadra et al.

The evidence is contradictory in detail but consistent in decision relevance: performance changes with tool, model, editing, genre, length, language, and test design.

Vendor claims versus independent tests

Turnitin initially claimed FACT—February 13, 2023 vendor claim 97% detection with a false-positive rate below 1 in 100, without publishing the test denominator in its release. Release

It later reported a less-than-1% document false-positive rate above its 20% score threshold, based on 800,000 pre-ChatGPT documents, and about 4% at sentence level. Those are FACT—vendor tests, not independent field validation. Turnitin also says false positives cannot be eliminated and the score must not be the sole basis for adverse action. Update, current guidance

Bias and access

The non-native-English evidence is strong enough to establish a material risk, but not to say every current detector is always biased. The 2026 study above found poor overall and genre-dependent performance without a significant EFL difference in its small sample.

Neurodivergent-writing evidence is thinner:

  • FACT—July 2026 preprint: a corpus of 59,947 Reddit posts had flag rates of 1.9% for a “likely autistic” group and 1.5% for general Reddit; the modeled odds were 25% higher before length normalization and 50% higher in a fixed-length analysis. It tested an older GPT-2 detector, inferred autism from subreddit participation, and was not peer-reviewed at the cutoff. Preprint
  • unknown — No controlled peer-reviewed study located by September 3, 2026 estimates current commercial-detector false-positive rates for autistic, ADHD, dyslexic, or broadly neurodivergent UVU students.

That gap supports caution. It does not support inventing a bias percentage.

Universities that disabled or rejected detection

  • Vanderbilt disabled Turnitin’s detector on August 16, 2023, citing opacity, false positives, non-native-writer bias, and privacy. Vanderbilt
  • Bournemouth declined activation on April 13, 2023 because it lacked time to test validity, fairness, policy, and communications. Bournemouth
  • Macquarie disabled it for its second 2023 session. Macquarie
  • Curtin announced on September 4, 2025 that it would disable AI detection from January 1, 2026. Curtin
  • University of Queensland disabled it in mid-2025. UQ
  • UCL, Leeds, Glasgow, Cambridge, Oxford, and the University of the Arts London publish similar warnings against detector-based decisions.

Legal and process exposure

UVU Policy 541 presumes the student not responsible and requires the university to prove every element by a preponderance of evidence—more likely than not. It also requires a reliable, impartial investigation and access to appeal. UVU Policy 541

A detector probability does not prove:

  • What tool produced the text.
  • Who operated it.
  • Whether use was permitted.
  • Which passages were contributed.
  • Whether an assistive tool explains the result.
  • Whether the student crossed the stated assignment boundary.

The UK higher-education ombudsman published FACT—July 2025 six selected AI cases: two justified, one partly justified, two not justified, and one settled. Complaints succeeded where institutions withheld evidence, did not examine drafts or the writing process, or failed to explain detector reliability. In one autistic student’s case, reconsideration ended with no misconduct finding. Case set, autistic-student case

In Harris v. Adams, a US federal court declined to block discipline involving AI-assisted work. But the record included copied material, fabricated citations, revision history, instructions, admissions, and an opportunity to respond—not a detector alone. Decision

unknown — No reported US judgment located by the cutoff establishes civil liability or damages specifically for a detector-only university accusation. The risks are still real: an unsupported finding can violate UVU’s own evidentiary process and create disability, national-origin, privacy, contract, and reputational disputes. That is risk analysis, not a legal opinion.

UVU’s rule

Detector output alone is never evidence of misconduct. UVU will not penalize, lower a grade, or require a student to disprove a detector score. A concern must be tied to a clear assignment rule and supported by independently reviewable evidence.

Independently reviewable evidence may include the assignment rule, inaccurate or fabricated sources, the submitted artifact, voluntarily supplied drafts, task-created process records, a fair student explanation, or a short follow-up demonstration aligned to the original learning outcome.

A student’s silence, disability status, accent, writing style, or inability to recall unrelated details is not proof.

Assessment designs that hold up

“Validity” means the task measures the knowledge or skill it claims to measure. “Secure” means outside help can be controlled or the student must directly demonstrate individual skill.

DesignValidity and integrity evidenceFaculty workloadScale to 300Student reception and named adoption
Oral defense or vivaStrong when prompts and scoring are structured and several questions sample the intended skill. A Melbourne model achieved reliability of fact α=.82 and .77 in cohorts of 319 and 342. StudyHigh for long one-to-one exams. EST—arithmetic: 300 students × 5 minutes ÷ 60 = 25 contact hours before transitions, training, and accommodations. A 20-minute defense is est 100 contact hours.Yes, with short structured checks, trained TAs, labs, or sampling.DCU piloted 322 participants in 2020–2021 and later reported use across 18 modules and more than 2,000 students. FACT—survey n=140: 82% said it encouraged integrity and 77% said it discouraged cheating. Anxiety usually needs practice. DCU
In-class or supervised writingSecures conditions, but timed recall, handwriting, typing speed, language fluency, and anxiety can displace the intended construct. In a randomized crossover, computer essays scored fact about 6 percentage points higher mainly because students wrote more. StudyRoom, proctor, device, accommodation, and make-up costs are unknown for UVU.Yes operationally; validity depends on the outcome.UiT’s 2025 secure redesign produced fact 33 failures among 179 students, versus 4 among 187 in the prior take-home year, but AI, time pressure, closed-book conditions, and anxiety cannot be separated. UiT
Process portfolio and version historyA portfolio can validly sample work over time. Version history records actions, not thought or authorship; content can be pasted, reconstructed, or produced offline.High when faculty inspect every draft. A Sydney portfolio used fact 372 raters for 257 students; rater and task effects were large. StudyPartial. Portfolio scoring reached fact n=1,208 at Idaho, but manual version-history audits at 300 are unknown.Students often value reflection and feedback but dislike paperwork and surveillance. Version history should be limited to records naturally created by the assignment.
Two-lane assessmentStrong policy logic: secure checks assure individual outcomes; open tasks teach responsible tool use. No published causal evaluation yet proves improved validity or lower misconduct.Secure rooms, staff, alternative sittings, and accessible formats can be expensive. UVU cost is unknown.Yes operationally. Sydney adopted it institution-wide from 2025; Bath from 2026/27; UQ from its second 2026 semester.Student trust and workload outcomes remain unknown. Policy reach is not effectiveness evidence.
Authentic and applied tasksStrong for relevance and transfer when the task mirrors real work. Weak as a stand-alone authorship control: a study of fact 419 contract-cheating tasks found outsourcing at every measured authenticity level. StudyLow to medium when one shared scenario and rubric are reused; high when every student needs a different client or dataset.Yes.Students often prefer meaningful work, but comparative reception evidence is thin. Pair with an oral, observed, or supervised sample when individual competence matters.
Staged assignments and checkpointsGood for feedback and revision. Weak proof of authorship unless at least one stage is observed or defended.One engineering comparison found fact 23% more assessment workload. EST—arithmetic: 3 checkpoints × 3 minutes × 300 students ÷ 60 = 45 staff hours. Peer feedback can reduce the bottleneck.Partial. PeerStudio served fact more than 3,600 learners, but secure individual assurance at that scale remains unproved.Rapid feedback supports revision. Repeated compliance uploads can feel like busywork.
Reflective componentUseful for explaining choices when attached to an artifact or observed event. Not secure alone: in a Monash dataset, AI reflections outscored student reflections and human authorship classification was only fact .5489 and .6800. StudyLow for short rubric-based notes; high for rich narrative feedback.Yes as a component, not as sole proof.Reception depends on relevance and feedback. Generic “what did you learn?” writing is easy to outsource.
Code review and pair-programming assessmentPair programming can support learning but cannot prove both partners’ skill. Individual live explanation, modification, debugging, or code review is stronger. In a 2026 pilot of fact n=541, oral project scores correlated r=.45 with a proctored midterm, versus r=.15 for code correctness. ReportUIC used three fact 10–20-minute interviews during existing labs, with each interviewer handling no more than 8 students. UICYes with a TA-rich lab structure. UIC reports courses up to 300.The first interview was stressful; repeated experience reduced anxiety. Pair-programming learning results are mixed.
Lab practical or OSPEStrong when the outcome is equipment use, safety, observation, technique, or real-time interpretation. It directly samples performance but can undersample broad knowledge.High, concentrated staffing and equipment. A Brunel model used an individual fact 20-minute, double-marked station; authors considered total work comparable with extended report marking. StudyPartial. Modern cohorts were 142 and 138; a contemporary single-section 300-student implementation was not found.Karolinska students generally judged an OSPE valid, but only fact 99 of 198 answered the survey. Stress and communication accommodations matter. Study

Design rule

No single design solves the problem:

  • Authenticity does not prove authorship.
  • Version history does not prove thought.
  • Reflection does not prove introspection.
  • Pair work does not prove individual skill.
  • Supervision does not automatically produce a valid task.
  • One oral answer does not sample an entire course.

The safest pattern is triangulation: an open product plus one short, structured individual check and a clear rule set before submission.

Policy patterns

Scales, traffic lights, and two lanes

The AI Assessment Scale now uses five levels: No AI, AI Planning, AI Collaboration, Full AI, and AI Exploration. Its authors warn against attaching labels to unchanged tasks or using unenforceable “No AI” rules. Framework

British University Vietnam reported FACT—January 2023 112 penalties among 1,722 submissions, or 6.50%; FACT—October 2023 0 among 3,996; and FACT—January 2024 4 among 4,159, or 0.10%, after AIAS implementation. The study also reported a 5.9% mean-grade increase and 33.3% higher module pass rate. These are FACT—reported observational outcomes, not causal effects: the institution, rules, communication, checking, and assessment designs all changed, and there was no control group. Study

UCL uses three categories—cannot use, assistive use, and integral use—and says a cannot-use task should normally be secure. It explicitly calls these categories guidance rather than formal policy. UCL

Sydney rejected fine-grained restrictions in unsecured work and adopted open and secure lanes. Bath and Leeds are also moving away from older traffic-light systems toward clearer outcome-based rules. Sydney, Bath, Leeds

Recommended UVU pattern

Reuse UVU’s existing levels, but place them inside two enforceable lanes:

  • Secure: AI-free except named assistive technology or approved tools.
  • Open—Assisted: editing, explanation, brainstorming, translation, or feedback only.
  • Open—Guided: named AI work is required within instructor-set steps.
  • Open—Collaborative: AI may contribute content, code, analysis, or media; students must verify and disclose it.
  • Open—Integral: skilled AI use is itself a learning outcome.

Do not rely on a course-wide label alone. Every high-stakes task needs its own label.

Best wording

AI use for this assessment: [Secure / Open—Assisted / Open—Guided / Open—Collaborative / Open—Integral].
You may use AI for: [named functions].
You may not use AI for: [named learning work].
If you use it, disclose the tool, purpose, material contribution, and what you checked or changed. You remain responsible for every submitted claim, source, calculation, and artifact.
In a secure task, no AI or outside help is allowed except [approved tools and accommodations].
Ask before submission if the boundary is unclear.
A detector score is never evidence of misconduct.

Do not require complete prompt histories by default. They may contain private data, unrelated work, or inaccessible tooling. A short material-contribution disclosure is more useful.

UVU’s current policy position

  • FACT—live manual checked September 3, 2026: no adopted UVU policy titled “Artificial Intelligence” appears in the manual. Manual
  • A proposed Policy 441 executive summary dated September 25, 2025 was a proposal to begin drafting and its file now returns 404. Current Policy 441 is the computing-facilities policy. UVU should not reuse the number for a second subject.
  • Policy 541 controls misconduct definitions, the preponderance standard, investigation, sanctions, and appeal. Policy 541
  • Policy 445 governs confidential and restricted data, including student records. Policy 445
  • Policy 447 governs university-hosted and third-party systems, access, encryption, logs, backup, and production authorization. Policy 447
  • Policy 452 requires accessible electronic technology and equally effective alternatives. Policy 452
  • UVU’s 2026 library guidance says there is no universal course AI rule and students must ask each instructor. Library guidance
  • UVU’s Writing Center already publishes assignment-use levels and disclosure guidance. Faculty guide
  • CHSS already publishes a tiered faculty guide and warns about detector use. CHSS guide

Where the assessment rule should sit

Use three linked layers:

  • 1. Academic Affairs assessment standard: require every course to state its default and every material task to carry a Secure/Open label.
  • 2. Policy 541 amendment or binding interpretation: define undisclosed or prohibited AI use as unauthorized assistance, preserve the preponderance standard, and bar detector-only findings.
  • 3. Implementation guide: reusable assignment text, disclosure form, examples, accessibility checks, and fair follow-up procedures.

Policies 445, 447, and 452 should govern any associated tools or records; they should not contain the academic assessment rule itself.

Utah and peer practice

  • FACT—December 2, 2025: the Utah Board of Higher Education called for responsible AI, literacy, workforce readiness, and preservation of human learning, but created no statewide assessment scale. USHE
  • BYU’s central policy remains tool-neutral and makes faculty responsible for communicating expectations; its Honors Program adopted a stricter written AI rule in 2024. BYU
  • USU’s 2026–27 catalog prohibits unauthorized assistance; teaching guidance says permissions should be clear and reasonable doubt may weigh against filing a case. USU
  • Weber’s code is tool-neutral. Its teaching center offers optional prohibit, limited-use, and full-use examples. Weber
  • SLCC expressly includes unauthorized AI on exams and unacknowledged AI words or ideas in its Student Code. SLCC
  • SUU requires every syllabus to explain the allowed extent of AI use. SUU
  • Snow College publishes prohibited, guided, and cited-use templates and says detector results are only a starting point. Snow
  • Utah Tech treats AI use as misconduct when the instructor prohibited it in writing and treats uncredited AI output as plagiarism. Utah Tech

The strongest Utah pattern is not a universal ban. It is clear written permission, task-level boundaries, disclosure, and ordinary due process.

Per-college patterns for UVU

These are proposed designs, not verified current college rules.

CollegeBest-fit patternExample redesign
Engineering & TechnologyOpen—Collaborative build plus secure individual code review, debugging, calculation, or design defense.Teams build a working system with AI allowed and disclosed. Each student receives a small defect or requirement change and must diagnose, modify, test, and explain it live. Grade the shared product and the individual check separately.
BusinessOpen—Collaborative analysis plus secure decision memo or board-style questioning.Students use AI to analyze a supplied market packet, disclose material contributions, verify sources, and present a recommendation. A short closed follow-up gives a changed fact and asks each student to revise the decision.
Humanities & Social SciencesOpen—Assisted or Guided staged research plus secure close reading or oral defense.Require a question, annotated sources, evidence map, draft, and revision note. Permit brainstorming and feedback but not fabricated sources. Add a short in-class analysis of a new passage or a structured defense of two major choices.
EducationOpen—Guided lesson design plus secure microteaching and defense.Students may use AI to generate alternative lesson ideas, then critique them against standards and learner needs. They teach part of the lesson, respond to a learner misconception, and explain the chosen adaptation.
ScienceOpen—Collaborative analysis plus observed lab skill and method defense.Students may use AI for code or interpretation after collecting data. They must maintain the normal lab record, perform one assigned technique, explain controls and uncertainty, and identify an intentionally flawed result.
Health & Public ServiceSecure simulation or practical for safety-critical skill; open guided critique for documentation and improvement.Use synthetic cases. The student completes an observed response, handoff, interview, or procedure, then may use AI to critique documentation. AI output is never a clinical authority.
ArtsOpen—Integral or Collaborative portfolio plus live critique, performance, or technique demonstration.Students disclose generated or transformed media, source rights, consent, and material AI contribution. The final review includes process artifacts and a live explanation or performance showing decisions and craft.

The common design is product plus proof. The product may use modern tools. The proof samples the human learning outcome.

Making it happen

Cost and workload

UVU’s exact redesign cost is unknown until the university measures paid faculty and support hours.

Useful published reference points are:

  • Cornell offers up to fact 20 fellows per academic year with a $5,000 stipend per fellow. Cornell
  • Duke’s 2024–25 pilot offered fact $2,000–$7,500 per project plus a $1,000 faculty stipend for up to six courses. Duke
  • Missouri proposed fact $64,590 for 20 fellows at $3,000 each plus program costs; larger proposed packages were $913,813 and $1,901,151. These were requests, not verified spending. Missouri
  • Michigan reported fact 318 sign-ups for its 2024–25 self-paced Teaching with GenAI resource. Sign-up is not completion or redesign. Michigan
  • Sydney reports fact about 30,000 assessments and more than 2.2 million submissions in its broader assessment framework. Central system changes, grants, workshops, consultations, and a community of practice supported the work, but AI-specific cost and outcomes are not isolated. Sydney framework

The current UVU plan’s $20,000–$40,000 one-college pilot is EST—plan basis: 30 faculty × $500 = $15,000, leaving est $5,000–$25,000 for leads, design support, administration, and evaluation. The $150,000–$250,000 campus figure is also est; no detailed staffing or workload model is supplied.

Keep the first range, but do not approve the campus range until the pilot measures:

  • Faculty redesign hours.
  • Review and calibration hours.
  • Staff time per secure check.
  • Accommodation and make-up time.
  • Student completion and appeal burden.
  • Reusable-template savings in the next term.

Minimum viable rollout

Before course-shell copy

  • Academic Affairs approves the shared lane and label language.
  • Student Conduct approves the evidence and follow-up procedure.
  • Accessibility reviews secure alternatives.
  • Canvas receives the label block and disclosure field.
  • Colleges select one signature assignment, not an entire course.

At course-shell copy

The supplied plan places this moment FACT—plan timing about 4–6 weeks before term. Every catalyst should:

  • Label each material assessment.
  • State allowed and prohibited functions.
  • Add one individual proof point where needed.
  • Add an accessible alternative.
  • Remove any detector-dependent wording.
  • Estimate staff minutes per student.

During the term

  • Give students a practice oral, practical, or secure mini-task before the graded version.
  • Use a shared rubric and calibration examples.
  • Record only data needed for assessment and review.
  • Pay adjunct faculty for redesign and required training.
  • Let catalysts hold short discipline-specific clinics.

After the term

Each catalyst publishes a de-identified design card containing the old task, new task, AI-use lane, learning outcome, workload, student response, integrity cases, limitations, and recommendation.

What to measure

Use measures that test the plan, not students’ private behavior.

Integrity

  • Referrals per est reporting unit 1,000 eligible submissions.
  • Substantiated findings per the same denominator.
  • Evidence types used.
  • Case resolution time.
  • Appeal, remand, and overturn rate.
  • Policy-clarity errors.

Assessment validity

  • Agreement between open-product scores and secure individual checks.
  • Rater agreement on sampled tasks.
  • Performance on the next course, placement, licensure, or capstone outcome where available.
  • Failure and withdrawal changes, with access and accommodation review.
  • Whether the assessment still covers the stated learning outcome.

Student trust

Use an anonymous survey covering clarity, fairness, fear of false accusation, ability to ask questions, equal tool access, accommodation, and whether the work felt meaningful.

Faculty feasibility

  • Redesign hours.
  • Delivery and marking minutes per student.
  • Make-up and appeal hours.
  • What could be reused next term.
  • What faculty stopped doing to make room.

Do not measure keystrokes, continuous screen activity, private prompts, browser histories, writing “style,” device telemetry, or unvalidated detector flags.

Cross-domain learnings

Medicine

Medical licensing combines knowledge with direct, structured clinical performance. Central standards, trained assessors, common cases, checklists, and review protect validity. The lesson for UVU is to sample actual performance at important assurance points, not turn every class into a practical exam.

Aviation

FAA certification combines a knowledge test, oral questioning, scenario judgment, and demonstrated performance. FACT—current standards effective May 31, 2024 integrate knowledge, risk management, and skill. FAA standards

The lesson is “artifact plus explanation plus action.” Access to modern cockpit tools does not remove the need to demonstrate judgment.

Apprenticeships

England’s end-point assessments normally combine at least two methods such as observation, demonstration, test, interview, or viva. UK guidance

The useful transfer is independence of evidence: a workplace product is meaningful, but an assessor also watches or questions the learner.

Software security

NIST’s zero-trust model says trust should not arise merely from network location or ownership. NIST

The assessment analogy is limited but useful: do not assume authorship because a file came through Canvas. Verify the particular learning outcome at a proportionate assurance point.

Quality assurance

Regulated fields do not rely on one noisy signal. They combine records, direct observation, structured questions, sampling, review, and appeal. UVU should do the same. A detector is closer to an unverified alert than to proof.

Devil’s advocate

The strongest case against this recommendation is that the cure may burden honest students more than dishonest ones.

Secure exams can reward speed and anxiety control instead of deep learning. Oral work can penalize speech disabilities, language learners, trauma, and cultural communication differences. Practical exams need rooms, staff, equipment, make-ups, and calibration. Version history can become surveillance. Frequent checkpoints increase adjunct labor. Students with jobs, care duties, long commutes, or online enrollment may bear the largest cost.

The institutional data also do not prove an integrity collapse. Manchester’s total cases fell; Challenge Success did not observe the feared post-ChatGPT jump; Australian self-report found no significant misconduct difference between detector and non-detector institutions. A campus-wide redesign mandate could spend scarce faculty time on a problem whose true UVU size is still unknown.

This argument defeats blanket redesign. It does not defeat targeted assurance. The proportional response is a small number of secure checks tied to degree-critical outcomes, with open assessment elsewhere and measured equity effects.

What would change the recommendation

The recommendation should change if any of these occur:

  • An independent detector demonstrates stable, externally replicated performance on current models, hybrid writing, UVU disciplines, non-native English writing, and disability-relevant samples, with transparent thresholds and reviewable evidence.
  • UVU’s baseline shows that existing assessments already provide strong individual assurance with low workload and high student trust.
  • A UVU pilot shows that oral or practical checks add little validity while materially worsening access, course completion, or trust.
  • Accreditors or state rules require a different form of secure assessment.
  • Reliable provenance standards make authorship contributions independently verifiable without collecting private writing behavior.
  • UVU’s adjudicated-case data show the main problem is policy confusion, fabricated sources, contract cheating, or another cause better addressed by a narrower intervention.

What the plan should change

  • 1. Replace “three policy templates” with one assessment-assurance standard. Require an assignment-level Secure/Open label, allowed functions, prohibited functions, disclosure, and an individual proof point where the learning outcome demands it.
  • 2. Make the detector rule binding. State that detector output is neither evidence nor a case metric and cannot by itself cause a grade change, referral, interview, or sanction.
  • 3. Use the course-shell moment as the control point. Add the label and disclosure block to Canvas before courses open, with a report showing unlabeled high-stakes tasks.
  • 4. Redefine the catalyst deliverable. Each paid catalyst redesigns and runs one signature assessment, measures workload and student trust, and publishes a de-identified design card.
  • 5. Pilot all seven colleges, but use different designs. Engineering should test code defense; Business a changed-fact decision task; Humanities staged research plus close reading; Education microteaching; Science a lab defense; Health a synthetic practical; Arts a provenance-rich portfolio plus live critique.
  • 6. Keep secure assessment sparse. Programs should identify a few assurance points where UVU must prove individual competence. Do not require a secure component in every assignment.
  • 7. Add a fair follow-up protocol. Give the student the rule and evidence, permit drafts and assistive-tool explanations, use questions aligned to the original task, give written reasons, and preserve appeal.
  • 8. Turn the current budget into a measured stage gate. Keep the est $20,000–$40,000 pilot, collect real hours and accommodation costs, and withhold the est $150,000–$250,000 expansion decision until results exist.
  • 9. Change the success measures. Count substantiated cases, overturns, validity, workload, access, and trust. Never count detector flags or surveillance events.
  • 10. Treat the “91%” claim as unverified. Either attach the actual survey and question or remove the number. The case for action is already strong without an unsupported statistic.

Source log

NOT_RUN

  • Output-file write: not runThe complete result is provided here.
  • UVU contact, faculty interviews, student interviews, vendor contact, logins, forms, and external messages: not run.
  • UVU course-level inventory, Canvas inspection, case-file review, and local student survey: not run.
  • Independent current-version detector testing on UVU writing: not run.
  • Accessibility testing of proposed oral, secure, coding, or practical tasks: not run.
  • Legal opinion on disability, discrimination, FERPA, contract, or civil-liability exposure: not run.
  • Causal evaluation of recent Sydney, Bath, UQ, or other two-lane policies: not run; published outcome evidence was not found.
  • Campus expansion budget approval: not run; the present figures remain estimates.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/43-assessment-integrity.md in the research pack.

W

Does it improve learning? The evidence and the measurement

Randomized trials and meta-analyses 2023–2026, when AI helps and hurts, the tutor-mode question, institutional outcomes, a pre-registered UVU study, and the learning-outcomes contract

Appendix W in five lines

  1. QuestionDoes the service help students learn, not just finish work faster?
  2. AnswerA course-built, hint-first tutor can help, but general chat access is not proven and can hurt later test scores.
  3. Deciding numbers+48% assisted practicefact−17% on the next unaided examfact420 enrolled students for the pilotest
  4. What the plan doesAsk why the student is using the tool, make hint-first learning the class default, register the study plan before it starts, and expand only if later tests without the tool show no harm.
  5. Still unknownNo study proves a campus-wide learning gain; local course data are UNKNOWN, and the university research review and sample-size study are NOT_RUN.

The plan measured adoption and never asked the provost's question: do students learn more, less, or differently? This appendix gathers the randomized trials and meta-analyses from 2023–2026 with their effect sizes, explains when AI helps and when it hurts, settles what the evidence says about tutor modes that refuse to give answers, checks whether any campus deployment moved retention or completion, and lays out a pre-registered study UVU could run in year one — with a learning-outcomes contract that says what result would stop the rollout. Research cutoff September 3, 2026; 110-plus dated sources. fact trial results · est thresholds and designs · unknown where evidence is thin.

What this changes in the plan

1. Access is not an intervention. Generative AI can improve learning, but handing students a general chatbot does not reliably do so. The strongest positive trials used course-aligned material, structured practice, feedback, teacher support, and active student work; the strongest warning trial found unrestricted AI raised assisted practice scores and lowered the next unaided exam. Refusing premature final answers removed that harm but did not by itself improve unaided learning. Immediate performance, retained learning, and transfer are different outcomes and can move in opposite directions; evidence beyond a few weeks is thin. 2. No credible study shows a campus-wide deployment improved retention, completion, time to degree, or institution-wide DFW rates. "Verified adopters," conversations, and satisfaction stay adoption measures, never learning measures — separate dashboards. 3. The router becomes purpose-first: is the student here to learn, produce, check, or decide? Learning-first tutoring (attempts, progressive hints, self-explanation, retrieval, transfer) is the course-practice default; a clearly labeled productivity mode exists where producing the product is the stated outcome; attempt gates must be removable without shame or disability disclosure. Course-owned tutor packs carry goals, approved sources, common errors, worked examples, faculty rules, and independent test boundaries. 4. A pre-registered, section-randomized pilot with delayed rollout, masked grading, delayed unaided assessments plus new problems the tutor never rehearsed, intention-to-treat analysis, and published confidence intervals — funded as part of the service, with UVU's IRB and the telemetry charter. 5. The learning-outcomes contract binds expansion: proceed to a larger trial only with no credible harm on delayed unaided learning; expand after the first credible year only if the delayed-learning estimate is positive with its lower 95% bound above −0.10 SD, transfer is not negative, DFW is no more than 2 points worse, and no subgroup shows repeated harm near 0.20 SD; reshape if assisted scores rise while delayed or transfer scores do not, or students mainly request answers; stop a mode if its upper confidence bound sits below no effect or repeated evidence shows delayed harm. No adoption target overrides these rules. est

Evidence cut: fact — 2026-09-03 MDT.

Executive verdict

  • Generative AI can improve learning, but access to a general chatbot does not reliably do so.
  • The strongest positive trials used course-aligned material, structured practice, feedback, teacher support, and active student work.
  • The strongest warning trial found that unrestricted AI improved assisted practice while reducing the next unaided exam score.
  • Refusing premature final answers removed that measured harm, but did not by itself improve unaided learning.
  • Immediate performance, retained learning, and transfer to new problems are different outcomes and can move in opposite directions.
  • Evidence for retention lasting more than several weeks is thin; evidence for far transfer is thinner.
  • No credible study located shows that a campus-wide generative-AI deployment improved retention, completion, time to degree, or institution-wide DFW rates.
  • UVU should treat the service as an intervention to test, not a benefit already proved.
  • “Verified adopters,” conversations, satisfaction, and generated work must remain adoption measures, not learning measures.
  • Recommendation: launch a learning-first tutor in selected courses, preregister the evaluation, test later without AI, and bind expansion to a public learning-outcomes contract.

Evidence terms

  • Assisted performance: Work completed while AI is available. It may show productivity rather than learning.
  • Immediate unaided learning: A test taken after AI is removed, usually during the same session.
  • Retention: An unaided test after a delay.
  • Transfer: Applying the learning to a different problem or setting.
  • DFW: A final grade of D or F, or withdrawal.
  • fact: A number reported by a dated source.
  • est: A planning estimate whose assumptions are shown.
  • unknown: The required evidence was not found.

What the research says, by design quality

Strong randomized evidence

StudyDesign and populationResultOutcomeWhat it does not prove
Harvard physics tutorfact — 2025 Crossover randomized trial; fact 194 eligible undergraduatesfact +0.63 SD adjusted effect; ceiling-adjusted estimates fact +0.73 to +1.30 SD; confidence interval UNKNOWN/not reportedImmediate unaided learningNo delayed retention or transfer. The tutor used instructor-written solutions, sequencing, videos, and question-specific scaffolds; it was not general ChatGPT.
Turkey high-school mathematicsfact — 2025 Classroom-randomized field trial with nearly fact 1,000 students over fact four 90-minute sessionsDuring practice, ordinary GPT raised scores fact 48%; the guarded tutor raised them fact 127%. On the next unaided exam, ordinary GPT reduced scores fact 17%; the guarded tutor effect was fact −0.004, not significant. Confidence intervals UNKNOWN/not reported; clustered standard errors were reported.Assisted performance and immediate unaided learningNo long-term retention or transfer. The guarded tutor prevented measured harm but did not beat control on unaided learning.
Nigeria World Bank studyfact — 2025 Student-randomized field trial in fact nine public schools; fact 1,328 randomized volunteers; fact 654 in the main assessmentCombined assessment fact +0.310 SD, standard error fact 0.068; English fact +0.238 SD; regular school exam fact +0.206 SD, standard error fact 0.067. Confidence intervals UNKNOWN/not reported.Immediate unaided learning and near transferThe intervention bundled fact twelve 90-minute after-school sessions, extra time, pairs, teacher support, structured prompts, and AI. Attrition was high and there was no time-matched active control.
Khan Academy/Khanmigofact — 2026 Two-year cluster trial in fact 18 Tennessee middle schools and fact 53 grade-within-school clustersPooled achievement fact +0.040 SD, standard error fact 0.019; first year fact +0.020 SD, standard error fact 0.035; second year fact +0.084 SD, standard error fact 0.041. Confidence intervals UNKNOWN/not reported.Within-term and annual mathematics achievementIt did not isolate Khanmigo from the larger Khan Academy practice system. The authors found gains similar to non-AI Khan Academy practice.
Productive AI tutoring experimentfact — 2026 Factorial randomized trial with more than fact 6,000 middle-school studentsAfter errors, guarded AI increased next-attempt correctness fact 0.085, standard error fact 0.009. At fact one week, practiced-item effect was fact +0.032, standard error fact 0.017; unpracticed-item effect was fact +0.002, standard error fact 0.017.Assisted remediation, short retention, transferThe delayed test contained only fact four items. The best signal was small, marginal, and restricted to practiced content.
Tutor CoPilotfact — 2025 version Tutor-level randomized trial; fact 783 tutors and fact 4,136 analyzed sessionsExit-ticket pass rate fact +4.0 percentage points, standard error fact 1.5 points, over a fact 61.7% control meanImmediate session masteryAI advised human tutors; it was not an autonomous student tutor. No measured annual retention or transfer.
Middlebury proctored experimentfact — 2026 preprint fact 211 undergraduates; fact 204 returned about one week laterImmediate unaided effect fact +0.266 SD, standard error fact 0.125; delayed effect fact +0.268 SD, standard error fact 0.120Immediate learning and one-week retentionSelective institution, unfamiliar topics, fixed fact 35-minute task, and preprint status. Explanatory use looked better than answer generation, but use style was not randomized.
LearnLM/Eedifact — 2025 vendor technical report Two-level randomized trial; fact 165 students in fact five UK schoolsNext-topic success: LearnLM fact 66.2%, fact 95% credible interval [61.1%, 71.2%]; human tutors fact 60.7%, interval [55.8%, 65.4%]. Difference fact +5.5 points, interval [−1.4, 12.4].Near transferThe AI-versus-human interval included no difference, every AI draft was reviewed by a human, and the report was vendor-authored.
AI-generated mathematics hintsfact — 2024 Randomized fact 3-by-4 experiment; fact 274 adultsUnaided gains: no hints fact 1.85%; human hints fact 11.62%; ChatGPT hints fact 17.00%. GPT versus no hints fact p=0.011; GPT versus human fact p=0.416. Confidence intervals UNKNOWN/not reported.Immediate unaided learningA nonsignificant GPT-human difference does not establish formal equivalence. fact 32% of raw AI help initially failed quality checks and required filtering.
GPT homework tutorfact — 2024 preprint Stratified randomized trial; fact 76 Italian high-school students over fact eight weeksOverall fact d=0.251, p=0.314; grammar-focused classes fact d=0.603, p=0.087; essay-focused classes fact d=−0.004, p=0.991Course learningSmall and underpowered. The tutor sometimes revealed answers despite instructions not to do so.
Delayed writing/reference tasksfact — 2025 preprint Two undergraduate randomized experiments with unaided follow-up after fact 14 daysEssay task: fact d=0.053, fact 95% CI [−1.6, 1.1] in raw-score units. Reference task: AI fact 5.98 versus control fact 7.54; AI effect fact d=−1.301RetentionSmall, single-university experiments with very different tasks and confusing contrast signs in the paper.

Randomized evidence with important selection or measurement limits

The Stanford Code in Place experiment randomized fact 5,831 active learners from fact 146 countries to receive access and advertising for a guarded coding assistant. Only fact 14.2% used it. Access reduced optional-exam participation by fact 4.3 percentage points, fact 95% CI [−6.9, −1.7], and homework completion by fact 4.6 points, fact 95% CI [−7.2, −1.9]. The intent-to-treat exam effect was fact +0.67 points, fact 90% CI [−0.37, 1.72], or no reliable difference. An estimated fact +6.86-point effect for adopters did not survive multiple-testing adjustment and depends on an untestable assumption. This is powerful evidence that offering a tutor can reduce engagement even when some users benefit. Nie et al., revised 2025

Two preregistered coding experiments found no significant overall unaided learning gain from ChatGPT. Allowing copying increased solution requests from fact 42.2% to fact 64.2% and reduced explanation requests from fact 31.3% to fact 18.4%. Students covered more material but did not learn more overall. Lehmann, Cornelius, and Sting, revised 2025

A fact — 2025 randomized retention study assigned fact 120 undergraduates to unrestricted ChatGPT or traditional study. At a surprise test fact 45 days later, usable observations implied by the reported test statistic were fact 85; ChatGPT students scored fact 57.5% versus fact 68.5%, fact t(83)=−3.19, p=0.002, d=0.68. Attrition and the limited public methodological record require caution, but it is one of the few genuinely delayed tests. Barcaui, 2025

Writing evidence

A randomized writing study with fact 117 university students compared ChatGPT, a human expert, writing analytics, and no extra support. ChatGPT improved the essay product, but knowledge gain and transfer did not differ. The authors used “metacognitive laziness” to describe altered self-regulation and dependence. It is a proposed mechanism, not a diagnosis or proof of permanent decline. Fan et al., 2025

A fact — Spring 2025 Carnegie Mellon quasi-experiment covered fact 424 students in fact 28 first-year writing sections. Multiple rounds of constrained AI feedback outperformed peer feedback on draft improvement, fact t=2.74, p=0.01, but draft revision is neither delayed retention nor far transfer. The AI was prohibited from generating replacement prose. CMU GAITAR@Scale

Seven experiments comparing LLM summaries or chat with web search found faster, shallower information gathering and less original downstream advice in several conditions. In the first experiment, LLM users searched for fact 585.41 seconds versus fact 742.81 seconds; reported learning depth was fact 3.43 versus fact 3.86, both fact p<0.001. These were immediate depth-of-processing tasks, not course retention tests. Melumad and Yun, 2025

Coding evidence

A fact — 2026 preprint meta-analysis covering fact 23 studies and fact 27 effects found assisted productivity fact g=0.33, 95% CI [0.09, 0.58], but learning fact g=0.14, 95% CI [−0.18, 0.47]. The learning result was not significant and mixed experimental with nonrandomized evidence. Maier et al., 2026

A Chinese university quasi-experiment with fact 82 programming students found AI-group performance fact 84.11 versus fact 78.36, but fact t(77)=1.28, p=0.204. Students frequently copied AI code, ran it, and returned errors to the model. Sun et al., 2024

A controlled introductory-programming experiment with fact 56 students found no significant learning improvement and observed reduced use of other learning resources after students began using ChatGPT. Chen et al., 2024

Older intelligent-tutoring baseline

These are structured, domain-specific systems. A general language model does not automatically inherit their effects.

  • fact — 2011 Human tutoring effect fact d=0.79; step-based intelligent tutoring fact d=0.76; answer-based systems fact d=0.31. Confidence intervals UNKNOWN/not reported. VanLehn
  • fact — 2014 Across fact 107 effects and fact 14,321 learners, intelligent tutors produced fact g=0.42 versus large-group teaching, fact g=0.57 versus other computer instruction, and fact g=0.35 versus textbooks or workbooks. They did not significantly beat individual human tutoring. Confidence intervals were not reported in the abstract. Ma et al.
  • fact — 2014 College intelligent-tutoring studies produced random-effects fact g=0.35, 95% CI [0.24, 0.46]. Steenbergen-Hu and Cooper
  • fact — 2016 Across fact 50 controlled evaluations, the median effect was fact 0.66 SD. Locally aligned tests averaged fact 0.73 SD; standardized tests averaged only fact 0.13 SD. This is a warning against testing a tutor with questions too close to its own material. Kulik and Fletcher

Evidence synthesis

The sign is not “AI versus no AI.” It is:

tool design × student behavior × course design × outcome timing.

The best-supported conclusions are:

  • Ordinary answer generation is excellent evidence of immediate productivity, not evidence of learning.
  • Guardrails prevent the clearest measured harm, but guardrails alone do not guarantee learning.
  • Positive trials usually combine AI with additional instructional ingredients.
  • Effects tend to shrink when assessments are delayed, unaided, standardized, or meaningfully different.
  • The causal evidence remains concentrated in mathematics, introductory science, coding, and short writing tasks.
  • Semester-long, multisubject, broad-access university evidence remains unknown.

Mechanisms

When AI helps

Attempt before assistance. Formative-feedback research recommends making the learner attempt the task before seeing the answer. A meta-analysis of problem-solving before instruction found fact g=0.36 for conceptual knowledge and transfer, but fact g=−0.03 for procedures. Novices still need an escape route. Sinha and Kapur, 2021

Progressive hints. The Turkey experiment directly shows that teacher-authored hints and answer resistance can change student behavior and prevent the harm observed with an unrestricted assistant.

Self-explanation. Asking the learner why a step works can support near and far transfer. A randomized worked-example study reported near-transfer effects up to fact f=0.42 and far-transfer effects up to fact f=0.37 for prompting, although the study predates generative AI. Atkinson, Renkl, and Merrill, 2003

Fading. Support should move from a worked model, to a missing final step, to more missing steps, and then to independent work.

Retrieval practice. Repeated study can look better after fact five minutes, while active retrieval wins after fact two days and fact one week. The tutor should therefore ask students to reconstruct an answer later without assistance. Roediger and Karpicke, 2006

Specific feedback. Feedback should address the task, explain what and why, arrive in manageable pieces, and follow an attempt. Too much information can become another answer dump. Shute, 2008

Course alignment plus independent testing. Course content helps the tutor avoid irrelevant answers. Independent or standardized assessments prevent that same alignment from inflating the measured effect.

Human support. Nigeria and Tutor CoPilot suggest that AI can amplify teacher or tutor attention. They do not show that the human layer can safely be removed.

When AI hurts

Answer-giving. Students can copy a correct solution without constructing the knowledge needed to reproduce it.

Cognitive offloading. The learner delegates the same mental operation the course is trying to teach.

Fluency mistaken for mastery. A fast, polished result feels like learning even when later unaided performance is unchanged.

Reduced help-seeking quality. Students often send bare answers, click suggested prompts, or request a solution instead of explaining their reasoning.

Lower engagement. The Code in Place trial found that simply advertising the tool reduced participation in other course activities.

Automation bias. Correct AI can help, but wrong AI can pull a previously correct human decision toward error. Explanations are not a reliable shield.

Mode confusion. A student may believe a productivity assistant is operating as a tutor, or that a tutor’s answer proves mastery.

The MIT EEG essay study and its limits

fact — 2025 preprint The MIT Media Lab study assigned fact 54 people, fact 18 per group, to LLM, search, or unaided essay writing for the first fact three sessions. Only fact 18 returned for the optional crossover. The calendar spanned fact four months; this was not continuous controlled exposure for four months.

In the first session, fact 15 of 18 LLM users said they could not quote their essay, versus fact 2 of 18 in each comparison group. None of the fact 18 LLM users produced a correct quotation. By the third session, the gap had narrowed.

The EEG analysis reported different connectivity patterns and less extensive directed connectivity in several LLM comparisons. It did not measure IQ, permanent brain change, semester learning, or far transfer. Connection counts are not standardized learning effects. Some detailed EEG comparisons also ran in the opposite direction.

A fact — 2025-12-29 methodological critique identified limited power, unclear multiple-comparison handling, missing standardized effects, reporting inconsistencies, and analyses sometimes based on only fact 2–4 essays per group.

Verdict: useful hypothesis-generating evidence about attention, ownership, and immediate source memory during AI-assisted composition. It is not proof that ChatGPT damages the brain or causes lasting intellectual decline. MIT study, independent critique

Equity: who benefits and who may fall behind

Evidence is contradictory.

  • Nigeria’s exploratory analysis found larger gains for girls, but the intervention was a supported after-school program and attrition was high.
  • Tutor CoPilot raised exit-ticket passing by fact 9 percentage points, from fact 56% to fact 65%, for students served by lower-rated tutors. This suggests AI may raise the floor when it improves human tutoring practice.
  • In Code in Place, access increased exam participation by fact 14.8 percentage points among students from lower-HDI countries while reducing engagement overall.
  • General productivity experiments often find larger gains for lower initial performers. That does not establish larger retained-learning gains.
  • The Turkey result warns that students who most need instruction can also be most exposed to harm when the system makes correct-looking answers easy to copy.
  • Effects for disabled students, multilingual UVU students, part-time students, working students, and different racial or ethnic groups remain unknown because most direct trials were too small or did not report suitable subgroup tests.

UVU should not assume that free access closes a learning gap. It should test access, substantive use, delayed learning, and harms separately.

The tutoring-mode question

What is actually proved?

The Turkey trial supplies the cleanest direct comparison using the same underlying model:

ModeAssisted workNext unaided exam
General assistantfact +48%fact −17%
Guarded tutorfact +127%fact approximately no difference
No AIBaselineBaseline

The guarded tutor:

  • used teacher-written solutions and common errors;
  • delivered hints;
  • resisted direct final answers;
  • checked student work;
  • encouraged attempts.

Students using the general assistant often asked for or copied complete solutions. Guarded-tutor students attempted answers and asked for help more often.

This proves that interaction design can flip a result from measured harm to no measured harm. It does not prove that refusal itself causes positive learning.

Why “never give the answer” is insufficient

The fact — 2026 Khanmigo trial is the warning. Although fact 96% tried the tutor, the median student messaged it on only about one-third of practice days and in only fact 17% of sessions containing a mistake. Only fact 14.5% of messages contained a real mathematics question or reasoning step.

A tutor can be pedagogically careful and still be ignored.

The service therefore needs graduated help:

  • first ask for the student’s attempt or plan;
  • diagnose the blocking point;
  • give one useful hint;
  • ask the student to act on it;
  • if still blocked, show one worked step;
  • if a complete worked example becomes necessary, follow it with a fresh, comparable problem completed without help.

An accommodation or urgent-completion path must be able to bypass effort gates without forcing a student to disclose a disability to the chatbot.

Implications for the UVU router and templates

The present plan’s router is designed mainly around workload, speed, context size, and cloud escalation. Add a learning-policy layer before model selection:

Student intent
├── Learn or practice
│   └── Learning-first tutor
│       ├── attempt
│       ├── diagnose
│       ├── progressive hint
│       ├── self-explanation
│       ├── fresh problem
│       └── delayed unaided check
├── Produce or execute
│   └── General assistant, with course-use and disclosure rules
├── Check or critique
│   └── Independent answer first, then AI comparison and source verification
└── High-stakes or accommodation
    └── Human-approved route with an accessible completion option

Required router behavior:

  • Ask whether the user is trying to learn, finish, check, or make a high-stakes decision.
  • Make learning-first tutor mode the default inside course practice.
  • Keep productivity mode visibly separate; never label its outputs as evidence of learning.
  • Cache the course’s learning goals, approved references, common errors, and faculty rules with the existing course prefix.
  • Do not let a faster model or cheaper machine silently weaken the tutor policy.
  • Log the mode, prompt-policy version, hint stage, and completion state without retaining student content beyond the approved study rules.
  • Detect repeated requests for a final answer and move to diagnosis or a worked-example-plus-transfer pattern.
  • Build no-tool retrieval prompts into the workflow.
  • Preserve a human escalation path.
  • Test the tutor against adversarial answer-seeking prompts before course use.
  • Measure actual substantive dialogue; activation alone is not tutoring.

Institutional outcomes

What campus deployments have established

unknown: No located campus-wide generative-AI rollout has published credible causal evidence of improved institution-wide retention, completion, time to degree, or DFW rates.

California State University, Arizona State University, Oxford, Virginia Tech, and other named deployments have reported access, activation, projects, training, use cases, or opinions. Those are implementation signals. They are not learning or completion effects.

A fact — 2026 University of Michigan preprint used administrative data covering fact 137,807 unique students and fact 46,485 offerings in its balanced sample. More AI-susceptible courses did not show significant post-ChatGPT differences in grades, withdrawals, or failures. The conservative estimates were:

  • grade fact +0.045 points, standard error fact 0.026, below conventional significance;
  • withdrawal fact +0.010, standard error fact 0.006, not significant;
  • failure fact −0.002, standard error fact 0.002, not significant.

Pre-trends failed for grades, so even the null should not be read as a clean causal effect. The study measured AI availability and course susceptibility, not verified student use or learning. Dumlao et al., 2026

Useful older, non-generative comparisons

Georgia State’s Pounce system was an administrative chatbot, not a tutor. In a randomized trial with fact 7,489 admitted students, outreach increased timely Georgia State enrollment by fact 3.3 percentage points, standard error fact 1.6 points, among the fact 1,948 students already committed to attend. It did not test subject learning or degree completion. Page and Gehlbach, 2017/2018

Georgia State later randomized proactive, course-specific, non-generative chatbot messages in government and economics. Across courses, the system increased the probability of an A or B by about fact 4 percentage points. In economics, women were fact 10 points less likely to DFW and earned grades fact 7 points higher. These effects came from reminders, course information, and support—not answer generation. Page et al., 2024/2026

Georgia State’s generative TEACH ME study began randomized trials in fact 2024, plans to cover more than fact 20,000 students, and runs through fact 2027. Final results remain unknown. GSU, 2024-10-08

UVU’s public baseline

UVU’s current public dashboard reports:

  • fact — Fall 2025 48,669 enrolled students;
  • fact — Fall 2025 40% first-generation students;
  • fact — 2017/18 cohort measured in 2025 48% completion;
  • fact — Fall 2025 72% annual retention.

UVU Data, accessed 2026-09-03

The Student Right-to-Know report dated fact — 2026-06-12 reports:

  • fact 47% overall graduation for the fact 2019 cohort;
  • fact 72% first-time bachelor’s retention for full-time students in the fact 2024 cohort;
  • fact 50% first-time bachelor’s retention for part-time students in the fact 2024 cohort.

UVU disclosure

These measures use different cohorts and definitions and should not be blended.

Current public UVU gateway-course DFW rates were UNKNOWN/not found. UVU should obtain course-section baselines from Institutional Research before power calculations or target-setting.

What success can reasonably mean

For the pilot, success should mean better delayed, unaided learning in a named course—not a detectable change in university retention.

For the first year:

  • Primary: delayed, unaided course learning.
  • Secondary: near and far transfer, common exams, course pass, DFW, withdrawal, next-course performance, and time to mastery.
  • Exploratory: next-term enrollment and annual retention.
  • Longer-term: completion and time to degree only after several cohorts.

Retention and completion are affected by advising, finances, employment, health, scheduling, course availability, and many other causes. A small tutoring pilot cannot identify its contribution to those outcomes.

How UVU should measure it

Recommended design

Use a preregistered, section-randomized delayed-rollout trial in gateway courses.

Pilot design

  • Select est two gateway courses with common outcomes and at least est seven sections per condition.
  • Randomize sections within course, instructor where possible, modality, and meeting time.
  • Treatment sections receive the learning-first tutor.
  • Control sections receive normal course resources and the same human-help routes, then receive the tutor after the evaluation period.
  • Do not attempt to forbid consumer AI. Record self-reported outside use and treat it as contamination.
  • Separate consent to use the service from consent for identifiable research data.
  • Analyze everyone by assigned condition, whether they use the tutor or not.
  • Run an additional usage-based estimate only as secondary analysis.
  • If section randomization is impossible, use a randomized staggered invitation.
  • If neither is feasible, use opt-in participation with matching on prior grades, course load, attendance, demographics, and baseline knowledge. Label the result quasi-experimental because matching cannot remove unmeasured motivation.

A three-way comparison with an unrestricted assistant would answer the tutoring-mode question directly, but the Turkey harm result weakens ethical equipoise. UVU should include that arm only for ungraded, low-risk practice if faculty and the IRB conclude that both modes are acceptable. A safer alternative is to randomize progressive-hint designs after students have made an initial attempt.

Outcomes and instruments

Primary outcome

est design choice A common, unaided assessment administered est four weeks after the relevant unit. Use external or faculty-written items not shown to the tutor. Score against a preregistered rubric by graders masked to assignment.

Secondary outcomes

  • Immediate unaided test within est 48 hours.
  • Near-transfer problems using different surface details.
  • Far-transfer problem requiring the same principle in a different setting.
  • Common course exam.
  • DFW and withdrawal.
  • Next-course grade or concept test.
  • Time to correct independent solution.
  • Error type and misconception persistence.
  • Student confidence calibration: confidence minus actual correctness.
  • Substantive tutor behavior: attempts, explanation requests, hints used, answer requests, verification, and completion of retrieval checks.
  • Accessibility, privacy, accuracy, and academic-integrity incidents.
  • Student and faculty time.
  • Cash and staff cost per learner achieving the preregistered mastery threshold.

Satisfaction and self-reported learning remain descriptive. They must not replace the learning test.

Sample sizes and detectable effects

These are planning estimates pending UVU’s real section sizes, baseline variance, attrition, and within-section correlation.

Assumptions:

  • est 30 analyzed students per section;
  • est 0.05 within-section correlation;
  • est 15% missing follow-up;
  • est 80% power;
  • est two-sided 5% error rate;
  • equal treatment and control allocation.

Under those assumptions:

StageProposed sampleApproximate detectable standardized effect
Pilotest 14 sections / 420 enrolled studentsest about 0.33 SD
First yearest 40 sections / 1,200 enrolled studentsest about 0.19 SD

Basis: the standard two-group approximation gives total individual-equivalent sample est 15.68/d²; the cluster design effect is est 1 + (30−1)×0.05 = 2.45; projected retention is est 85%.

The pilot is therefore a feasibility, safety, behavior, and large-effect screen. It cannot establish that a small effect is absent.

A DFW calculation is unknown until UVU supplies course-specific baselines. Detecting a small percentage-point change could require several thousand students.

Before preregistration, a statistician should replace these approximations with simulation using UVU’s actual section counts, unequal class sizes, baseline scores, and expected missingness.

Analysis

Preregister:

  • one primary outcome and time point;
  • treatment assignment and exclusions;
  • section-level clustering;
  • course and randomization-block effects;
  • baseline-score adjustment;
  • missing-data and attrition sensitivity analyses;
  • handling of outside AI use;
  • correction for multiple secondary tests;
  • subgroup definitions;
  • stopping rules;
  • prompt, router, model, and course-content version;
  • treatment-change rules if the model fails.

Report intention-to-treat first. Report point estimates and confidence intervals even when results are null.

Predeclare equity analyses for first-generation status, Pell eligibility where authorized, prior preparation, part-time/full-time status, modality, gender, race or ethnicity, multilingual status, and disability accommodation. Suppress unsafe small cells and describe underpowered subgroup findings as exploratory.

IRB and privacy path

UVU states that human-subjects research, including pilot and feasibility studies intended for public dissemination, must be submitted to its IRB. Investigators cannot determine exemption for themselves. The application must describe randomization, controls, recruitment, risks, study materials, and analysis. CITI training is required before the project begins. UVU IRB FAQ, application process

Educational practice research may qualify for an exempt determination if it is minimal risk and does not harm students’ opportunity to learn. The IRB makes that determination. UVU exempt categories

The sequence should be:

  • appoint a UVU faculty principal investigator;
  • include Institutional Research, teaching-center staff, accessibility, privacy, students, and participating faculty;
  • finish CITI training;
  • freeze the protocol, assessments, consent language, data map, and deletion schedule;
  • obtain the IRB determination before recruitment or research data collection;
  • execute any required FERPA studies agreement;
  • preregister before viewing outcomes;
  • begin the pilot only after those gates pass.

FERPA’s studies exception can permit limited education-record disclosure for research conducted for the institution, but it requires a written agreement specifying purpose, use, security, and destruction. Use deidentified data whenever possible. U.S. Department of Education

Cost of a credible evaluation

All figures below are estimates of full economic cost, including existing staff time.

Pilot: est $75,000–$110,000.

Illustrative midpoint basis:

  • evaluation lead: est 0.20 FTE × $140,000 loaded annual cost = $28,000;
  • analyst/data support: est 0.15 FTE × $120,000 = $18,000;
  • faculty participation: est 14 sections × $1,000 = $14,000;
  • retained-assessment completion support: est 300 students × $20 = $6,000;
  • assessment, accessibility, data, and privacy work: est $17,000;
  • contingency: est about $12,000;
  • total illustrative midpoint: est about $95,000.

First credible year: est $210,000–$280,000.

Illustrative midpoint basis:

  • evaluation lead: est 0.50 FTE × $140,000 = $70,000;
  • analyst/data engineering: est 0.40 FTE × $120,000 = $48,000;
  • faculty participation: est 40 sections × $1,000 = $40,000;
  • delayed-assessment completion: est 1,000 students × $20 = $20,000;
  • accessibility, data governance, independent assessment review, and publication: est $50,000;
  • contingency: est about $34,000;
  • total illustrative midpoint: est about $262,000.

Actual UVU salary, workload, incentive, and assessment costs are unknown.

What UVU should publish

Publish:

  • the dated protocol and registration;
  • intervention and control descriptions;
  • model and prompt-policy versions;
  • assessment instruments where test security permits;
  • section and participant flow;
  • assignment balance and attrition;
  • all primary and secondary results with confidence intervals;
  • null and harmful results;
  • subgroup estimates with privacy protections;
  • incident counts and model changes;
  • cost per assigned student, active user, and learner reaching mastery;
  • analysis code, data dictionary, and the most deidentified data the IRB permits;
  • deviations from the protocol.

Do not publish identifiable chats or create a misconduct dataset from research logs.

Learning-outcomes contract

These are proposed governance thresholds, not established scientific constants.

Proceed from pilot to a larger trial only if:

  • no credible harm appears on delayed unaided learning;
  • assessments and data collection work as designed;
  • treatment separation is real;
  • accessibility and privacy checks pass;
  • substantive learning-mode use is high enough to test the intervention;
  • faculty and students can use the system without coercion.

Expand after the first credible year only if:

  • est policy threshold the delayed-learning point estimate is positive and the lower est 95% confidence bound is above est −0.10 SD;
  • transfer is not negative;
  • est policy threshold DFW is not more than est 2 percentage points worse;
  • no important subgroup shows a repeated adverse signal around est 0.20 SD or larger;
  • serious privacy, accessibility, or answer-leak incidents are resolved;
  • costs are acceptable relative to observed learning, not usage.

Reshape the intervention if:

  • assisted performance rises but delayed or transfer performance does not;
  • students mainly request answers or paste outputs;
  • the tutor is rarely used after mistakes;
  • tutor mode reduces engagement with class, instructors, or human tutoring;
  • benefits occur only under intensive human support that the scale plan cannot sustain.

Stop the affected mode if:

  • the upper confidence bound for the primary learning effect is below no effect;
  • repeated evidence shows meaningful delayed harm;
  • the system bypasses faculty assessment rules;
  • serious privacy or accessibility harm cannot be fixed promptly.

No adoption target overrides these rules.

Cross-domain learnings

Calculators

A fact — 1986 meta-analysis of fact 79 reports found that calculators used with normal instruction generally maintained or improved paper-and-pencil computation and problem solving, with a grade-four exception.

A fact — 2003 synthesis of fact 54 studies found strong gains when calculators were part of both instruction and tests: operational fact g=0.38, computational fact g=0.43, conceptual fact g=0.44, and problem-solving fact g=0.33. When calculators were withheld on tests, most effects were near zero; operational skill was fact g=0.17. Only fact three studies assessed retention after fact 2–12 weeks.

Learning: do not ban a useful tool. Teach with it, preserve no-tool practice for skills students must own, and measure both tool-enabled and independent work.

Spell-checkers

A fact — 2017 randomized study with fact 88 university second-language learners found that spell-check choice and dictionary use supported correction after aid removal and a fact one-day delay. A red underline alone did not.

A fact — 2006 study with fact 65 university learners found better surface correction without less content revision, but did not test later unaided spelling.

Learning: correction can free attention for ideas, but passive flags do not teach much. Ask the learner to select, explain, and later reproduce the correction.

GPS and spatial memory

A fact — 2008 experiment found that turn-by-turn GPS users learned routes less well than direct-experience or paper-map users.

A fact — 2020 study of fact 50 regular drivers linked greater GPS use with poorer self-guided spatial memory. Only fact 13 returned about fact three years later, so its longitudinal causal claim is weak.

Landmark-rich guidance and auditory beacons that leave route choice to the traveler performed better on later navigation measures than ordinary turn-by-turn directions in smaller experiments.

Learning: preserve decisions and structural cues. A tutor should teach the map of the subject, not issue isolated turns.

Clinical decision support

A fact — 2023 experiment with fact 457 clinicians across fact 13 states found that standard AI advice raised diagnostic accuracy fact 2.9 percentage points; biased advice lowered it fact 11.3 points. Adding an explanation did not reliably protect clinicians from biased advice.

In another study, fact 41.73% of non-radiologists and fact 27.54% of radiologists accepted both incorrect recommendations presented to them.

A fact — 2025 observational colonoscopy study found unaided adenoma detection fell from fact 28.4% before AI exposure to fact 22.4% afterward. A larger fact — 2026 prospective study covering fact 5,013 colonoscopies did not find significant post-removal decline, contradicting a simple de-skilling conclusion.

Learning: get an independent judgment before showing AI advice. Explanations are not enough; require verification and test performance after assistance is removed.

Aviation automation

A fact — 2014 simulator study with fact 16 active pilots found that basic control and instrument-scan skills held up better than manual navigation and abnormal-event diagnosis. Only fact 1 of 16 completed every manual-navigation phase without a listed error.

The fact — 2013 FAA automation review found substantial safety and workload benefits alongside mode confusion, weak monitoring, and erosion of some manual and cognitive skills.

Learning: keep modes visible, train graceful failure, and schedule periodic no-tool drills for essential skills.

Shared lesson

Other fields stopped asking whether assistance is “good” or “bad.” They ask:

  • Which skill must remain human?
  • Which part may be safely offloaded?
  • Does the tool preserve active decisions?
  • What happens when the tool is wrong or absent?
  • Is independent skill practiced and checked?

UVU should use the same frame.

Devil’s advocate

The strongest case against the recommendation is that an answer-refusing campus tutor may solve the wrong problem.

Students already have unrestricted consumer assistants. A slower institutional tutor could frustrate them and push use into less governed tools. Modern work increasingly rewards delegation, critique, synthesis, and verification rather than unaided production. Protecting every old no-tool skill may resemble requiring manual arithmetic after calculators became normal.

The strongest positive studies are narrow, bundled, and often conducted with younger learners. The strongest harm study lasted only a few sessions. Models and interfaces change faster than normal educational research. A costly trial could measure an obsolete implementation by publication time. Faculty may get more value from redesigning assessments and teaching AI verification than from restricting answers.

That case is serious. It supports a dual-mode service, not an unrestricted default:

  • tutoring mode for learning goals;
  • productivity mode where tool-enabled work is the course goal;
  • faculty control over which mode applies;
  • no-tool checks only for skills the course says students must retain;
  • rapid versioned experiments instead of one frozen, multiyear product.

What would change the recommendation

The recommendation would move toward a general-assistant default if independent, multi-course university trials showed all of the following:

  • ordinary assistants outperform guarded tutors on delayed unaided learning;
  • the advantage persists across semesters and transfers to new problems;
  • unrestricted assistance does not widen gaps by prior preparation;
  • answer refusal materially reduces engagement without a compensating learning benefit;
  • transparent productivity-mode teaching produces equal or better independent verification skill.

The recommendation would become more restrictive if:

  • UVU or replicated trials find delayed harm around est policy threshold 0.20 SD or more;
  • students systematically bypass tutor controls and perform worse later;
  • course DFW or withdrawal increases;
  • harmful effects concentrate among students with weaker prior preparation;
  • faculty cannot keep assessments independent of tutor training material.

The recommendation would pause entirely if UVU cannot obtain an IRB determination, a lawful data path, accessible alternatives, common assessments, or enough sections for a credible comparison.

What the plan should change

  • 1. Add the learning-outcomes contract before buying or scaling. Adoption may unlock evaluation, but it cannot authorize expansion by itself.
  • 2. Change the router from workload-first to purpose-first. Ask whether the student intends to learn, produce, check, or make a high-stakes decision before selecting a model or machine.
  • 3. Make learning-first tutoring the course-practice default. Require attempts, progressive hints, self-explanation, retrieval, and transfer.
  • 4. Keep a clearly labeled productivity mode. Use it when generating the product is allowed and tool-enabled performance is the stated outcome.
  • 5. Add delayed, unaided measurement. Every pilot course needs a common assessment outside the tutor, plus new problems the tutor did not rehearse.
  • 6. Run a preregistered section-randomized pilot. Use delayed rollout, intention-to-treat analysis, masked grading, and published confidence intervals.
  • 7. Measure student behavior without treating surveillance as learning science. Record content-minimized events such as attempt, hint stage, answer request, verification, and retrieval completion under the telemetry charter.
  • 8. Add an accessibility escape. Attempt gates must be removable without shame, penalty, or unnecessary disability disclosure.
  • 9. Require course-owned tutor packs. Each pack should contain learning goals, approved sources, common errors, worked examples, faculty rules, and independent test boundaries.
  • 10. Separate dashboards. Report adoption, service quality, assisted productivity, immediate learning, delayed retention, transfer, DFW, equity, safety, and cost as different measures.
  • 11. Do not call campus access an intervention. The evaluated intervention is the full combination of tutor policy, course integration, faculty support, student behavior, and assessment.
  • 12. Fund evaluation as part of the service. A rollout without credible outcome measurement leaves the provost’s central question unanswered.

Source log

NOT_RUN

  • UVU contact: not run, as required.
  • Vendor contact: not run, as required.
  • Private UVU gateway-course DFW extraction: not run; public current rates were not found.
  • Student-level UVU power simulation: not run; real section sizes, outcome variance, missingness, and within-section correlation were unavailable.
  • IRB submission or determination: not run; only UVU’s public process was reviewed.
  • Independent full-text risk-of-bias scoring for every source: not run.
  • New confidence intervals from published standard errors: not run; intervals are given only where sources reported them.
  • Meta-analysis combining generative-AI studies: not run; interventions and outcomes were too different for a defensible pooled estimate.
  • Deployment, data collection, student tracking, messages, purchases, and public release: not run.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/44-learning-outcomes-evidence.md in the research pack.

X

Trust, sources, and error handling: the interface

How wrong and how often, trust calibration, showing sources and uncertainty, repairs and refusals, conversation design for tutoring, and the pilot's usability tests

Appendix X in five lines

  1. QuestionHow can students know whether an answer can be trusted, see what supports it, and get mistakes fixed?
  2. AnswerUse one trust rule that shows exact source passages and limits, says not found when needed, and tracks each mistake until it is fixed.
  3. Deciding numbers2,000 students served by Wilsonfact10-second first-word targetestat least 95% right-source rate for cited claimsest
  4. What the plan doesKeep one student front door, add exact-passage links, track repair cases, use hint-first tutoring, provide a real person handoff, and enforce pilot gates.
  5. Still unknownLive Wilson and LibreChat behavior, support routes, and the exact model-and-document error rate are UNKNOWN; signed-in tests and benchmarks are NOT_RUN.

Trust is won or lost at the moment a student sees a source, a doubt, a mistake, or a refusal — and the plan had chosen a front door and a router without designing that moment. This appendix measures how wrong each model class is by task, what grounding on documents fixes and what it does not, what the research says about over-trust and under-trust, which citation and uncertainty designs measurably help, how to handle errors and refusals, what conversation design does for tutoring, and how to test all of it inside the 14-week pilot. Research cutoff September 3, 2026; 110-plus dated sources. fact measured · est designs and thresholds · unknown untested at UVU.

What this changes in the plan

1. One campus trust contract across Wilson (UVU's existing course assistant), LibreChat, and every routed model: the same source states, uncertainty states, refusals, report route, and human route. The review's sharpest point: do not launch a second, competing student front door until UVU decides Wilson's role — the fleet can serve behind Wilson or inside approved course tools first. 2. A citation badge is not enough. Every sourceable claim opens the exact supporting passage; every answer shows whether it came from approved UVU material, an uploaded file, the open web, or model memory; an evidence-sufficiency gate lets the assistant say "not found" or "sources conflict" instead of filling the gap. Raw confidence percentages are not shown unless UVU validates them for that model, task, and course — the studies show confident tone and unverified numbers miscalibrate trust in both directions. 3. Errors become repair cases with a receipt, an owner, a status, a visible correction notice, and correction memory — not a thumbs-down counter. Verifiable tools (calculators, code tests) are used wherever the task permits. 4. Tutoring is student-first and hint-first, enforced outside the model prompt where sequence matters: the student attempts the work before advice; steps before final answers. 5. Identity and route stay visible — "UVU AI assistant — not a human" and the active route on screen throughout, which is also Utah's safe harbor. 6. Response time: keep the 10-second first-word target, add truthful status and restartable streaming. 7. Helpful refusals and a real, acknowledged human route before broad student use — and if a staffed route cannot be funded, remove the features that imply one; never fake it. 8. Test it in the pilot: named tasks, participants, calibration and error-recovery metrics, and launch gates; publish results by task and route, never one blended "accuracy" or "trust" score. est

Research date: fact — September 3, 2026, MDT.

Evidence labels: fact means a dated source reports it; fact-derived is arithmetic from reported data; est is a planning choice with its basis stated; unknown means the evidence does not support a number. Benchmark failure is not called hallucination unless the study defines it that way. Model names, statute numbers, weeks, and source-list numbers are identifiers rather than measurements.

Executive verdict

  • 1. Keep LibreChat as an engine candidate, but do not make it a second, conflicting campus identity beside Wilson.
  • 2. Give Wilson and LibreChat one shared trust contract: the same sources, uncertainty states, refusals, reports, and human routes.
  • 3. A citation badge is not enough; every sourceable claim should open the exact supporting passage.
  • 4. Show whether an answer came from approved UVU material, an uploaded file, the open web, or model memory.
  • 5. If the available evidence does not answer the question, the assistant should say so instead of completing the gap.
  • 6. Do not show a raw confidence percentage unless UVU validates it for that model, task, and course.
  • 7. Make the student attempt the work before showing tutoring advice; use hints and steps before final answers.
  • 8. Treat feedback as a repair case with a receipt, owner, status, and correction—not as a thumbs-down counter.
  • 9. Keep “UVU AI assistant — not a human” visible throughout the interaction.
  • 10. Do not launch broad student use until source fidelity, error recovery, crisis routing, and accessibility pass the pilot tests.

How wrong, how often

There is no defensible universal “hallucination rate.” The answer changes with the task, model settings, retrieval quality, tools, prompt, grading method, and whether abstention is allowed.

No public evaluation tests the plan’s current local, heavy, and frontier models on the same factual, citation, math, code, and uploaded-document tasks. That combined comparison is unknown and should be produced during the pilot.

Published model evidence

Model class and dated sourceFactual questionsMathCodeDocuments and citations
Current local class: Qwen3.8-27B, official card, fact — August 14, 2026GPQA accuracy fact 89.2%; raw miss fact-derived 10.8 percentage points. This is scientific reasoning, not a conversational hallucination test.A comparable text-math rate is unknown. Visual MathVision accuracy was fact 90.0% without code execution and fact 94.6% with it.SWE-bench Pro resolved fact 61.7%; unresolved fact-derived 38.3 points. LiveCodeBench scored fact 90.3%, under a different harness.True factual-hallucination, citation, and uploaded-summary rates are unknown. Qwen model card
Older local comparator: Qwen3-32B thinking, technical report, fact — May 14, 2025GPQA accuracy fact 68.4% across fact 198 questions sampled 10 times, or fact-derived 1,980 generations.MATH-500 accuracy fact 97.2%; raw miss fact-derived 2.8%.LiveCodeBench accuracy fact 65.7%; raw miss fact-derived 34.3 points.No matching citation or document-summary rate. Qwen3 report
Local-class caution: Gemma 3 27B IT, official card, fact — March 2025SimpleQA accuracy fact 10.0% over fact 4,326 questions. The remaining fact-derived 90.0 points include abstentions and errors; they are not all hallucinations.MATH accuracy fact 89.0% over the fact 5,000-problem test set.HumanEval pass rate fact 87.8% over fact 164 problems.FACTS Grounding score fact 74.9; its deficit is not a literal count of bad responses. Gemma 3 model card
Heavy class: gpt-oss-120b, official card, fact — August 5, 2025On SimpleQA without browsing: accuracy fact 16.8%, hallucination fact 78.2%, and abstention/other fact-derived 5.0%, over fact 4,326 questions. PersonQA accuracy was fact 29.8% and hallucination fact 49.1%.AIME 2025 with tools scored fact 97.9% over fact 30 problems; raw deficit fact-derived 2.1 points.SWE-bench Verified resolved fact 62.4%, or fact 312 of 500 tasks; fact 188 of 500 remained unresolved.Exact citation and current private-document rates are unknown. gpt-oss evaluation
Heavy alternative: Mistral Small 4, official card, fact — March 16, 2026GPQA accuracy fact 71.2% over the standard fact 198 prompts; raw miss fact-derived 28.8 points.Auditable absolute results for the plan’s use are unknown.Auditable absolute results in accessible text are unknown.Hallucination, citation, and document-QA rates are unknown. Mistral Small 4
Frontier comparator: GPT-5.6, official release, fact — July 9, 2026GPQA accuracy fact 94.6%; raw miss fact-derived 5.4 points. This is not a campus factuality rate.FrontierMath scores were fact 89% on lower tiers and fact 83% on the hardest tier; deficits fact-derived 11 and 17 points.SWE-bench Pro resolved fact 64.6%; unresolved fact-derived 35.4 points.A current absolute uploaded-document or citation error rate is unknown. GPT-5.6
Frontier browsing evidence: GPT-5 system card, fact — August 7, 2025On representative factual traffic with browsing, incorrect claims were fact 4.5% and responses containing a major factual error were fact 4.8%. Sample size was unknown. Without web on SimpleQA, accuracy was fact 55%, hallucination fact 40%, and abstention/other fact-derived 5%.Not comparable here.Not comparable here.Browsing helps materially, but errors remain and the production grader was itself a model. GPT-5 system card

Citations and uploaded documents

Grounding means giving the model retrieved evidence. It changes the failure pattern; it does not make the answer safe by itself.

  • RAGTruth contains fact 17,790 retrieval-grounded responses. At least one annotated hallucination appeared in fact 1,724 of 5,934 QA responses, or FACT-derived 29.05%, and fact 1,686 of 5,658 summaries, or FACT-derived 29.80%. The models date mainly from 2023, and later work suggests subtle errors were undercounted. These are warning rates, not forecasts for UVU. RAGTruth, ACL 2024
  • FACTS Grounding supplied the full source document for fact 1,719 prompts. Scores were fact 78.8 for GPT-4o, fact 83.6 for Gemini 2.0 Flash Experimental, and fact 79.4 for Claude 3.5 Sonnet. These scores combine judges and splits; their deficits are not literal error percentages. factS Grounding, January 6, 2025(https://arxiv.org/abs/2501.03200)
  • In the ALCE citation benchmark, GPT-4 answering ELI5 questions with fact 20 retrieved passages achieved citation recall fact 48.5 and citation precision fact 53.4. About half of the tested long answers were not fully supported by their citations. On ASQA, citation recall was fact 73.0 and precision fact 76.5. ALCE, December 6, 2023
  • A study of fact 636 generated citations found fabrication in fact 55% of GPT-3.5 citations and fact 18% of GPT-4 citations. Among citations that existed, substantive errors remained in fact 43% and fact 24%, respectively. These were closed-book 2023 systems, but they prove that bibliography-shaped text is not evidence. Walters and Wilder, September 7, 2023
  • “Sufficient Context” found that irrelevant retrieval can reduce abstention and increase wrong answers. For one Gemma condition, incorrect answers rose from fact 10.2% without context to fact 66.1% with insufficient context. A sufficiency-aware answer gate improved correct-among-answered by fact 2–10 percentage points. ICLR 2025 paper
  • A citation can support a sentence without having caused the model to produce it. In an adversarial test, a retrieval model attached apparently supporting citations to claims driven by uncited or random material in as many as fact 57% of one condition. Correctness is not Faithfulness, December 23, 2024

What remains wrong after grounding:

  • The right passage was never retrieved.
  • An old or unauthorized document outranks the current source.
  • The passage is relevant but does not support the claim.
  • The model merges facts from different people, dates, policies, or courses.
  • A correct answer is surrounded by invented explanation.
  • A summary omits an important exception or qualifier.
  • A citation exists but points to the wrong section.
  • Conflicting sources are silently collapsed into one answer.
  • The model answers even though the evidence is incomplete.
  • The automatic citation checker is also wrong. MiniCheck found only fact 75.3% balanced accuracy for GPT-4 across fact 10 grounding datasets. MiniCheck, November 12, 2024

Therefore, UVU should measure source retrieval, claim support, answer correctness, and abstention separately.

Trust calibration

Trust calibration means using the assistant when it is right and checking or rejecting it when it is wrong. High trust is not the goal. Appropriate trust is.

What the studies say

  • Confident length can masquerade as quality. In controlled experiments with fact 60, 60, and 59 participants, longer explanations increased confidence without improving discrimination between right and wrong answers. Human discrimination was fact AUC 0.54 for long explanations and fact AUC 0.57 for uncertainty-only answers. Language tied to measured model uncertainty improved calibration, though it did not improve users’ underlying subject knowledge. Nature Machine Intelligence, January 21, 2025
  • Citations can raise trust before they are checked. A randomized study reports fact 303 participants and fact 1,976 citation-bearing answers. Only fact 193 citations, or FACT-derived 9.77%, were opened or hovered over. fact 83 of 197 citation-condition participants, or FACT-derived 42.1%, checked at least one. Random citations still increased trust unless people inspected them. The paper has an internal count conflict: its printed condition totals sum to fact-derived 305, not the reported fact 303. AAAI, April 11, 2025
  • Uncertainty wording can reduce blind agreement. In a preregistered medical-question study with fact 404 participants and an intentionally fact 50%-accurate assistant, agreement fell from fact 80.9% with unhedged answers to fact 74.8% with first-person uncertainty, while answer accuracy rose from fact 63.9% to 72.8%. It did not produce a significant increase in source checking. Kim et al., FAccT, June 5, 2024
  • Confidence displays work only when the confidence is calibrated. A study with fact 72 participants found that a calibrated confidence display helped people distinguish when to follow an assistant, but did not improve combined human-AI accuracy. Zhang, Liao, and Bellamy, January 27, 2020
  • Forcing a small pause reduces overreliance. In a study with fact 199 participants, asking for an independent judgment or delaying advice reduced wrong-AI overreliance from fact 0.64 to 0.48 on one decision measure. Participants liked the more demanding interfaces less. Buçinca et al., April 2021
  • Visible errors usually hurt more than correct answers help. Algorithm-aversion experiments found that people abandoned an algorithm after observing its mistakes even when it still outperformed a human forecaster. Across the reported studies, fact 610 of 741 participants, or FACT-derived 82.3%, saw the algorithm outperform the human. Dietvorst, Simmons, and Massey, 2015
  • But the first error does not always cause permanent distrust. A legal-advice experiment with fact 208 participants and fact 14 cases found an immediate trust drop after a planted error, followed by relatively rapid behavioral recovery. This contradicts a universal “one error destroys trust” claim. Kahr et al., April 5, 2024
  • Honest advance warning can soften later failure. A randomized chatbot experiment with fact 558 users found that candid low-performance messaging reduced discontinuance after a response failure; inflated claims did not. Weiler, Matt, and Hess, December 22, 2021

Interface conclusion

Use evidence-based uncertainty, not personality-based uncertainty.

Good:

  • “The current course policy says …”
  • “I found this in the syllabus and assignment page.”
  • “The sources conflict: the syllabus says X, while the newer announcement says Y.”
  • “I could not find this in the materials I searched.”
  • “This is an inference, not a statement in the source.”
  • “I can offer a general explanation, but it is not verified against UVU material.”

Avoid:

  • “I’m definitely right.”
  • “I’m 92% confident” when the number has not been validated.
  • A long rationale that adds no evidence.
  • Green checkmarks based only on model self-confidence.
  • Citations that open only a home page or document title.

Showing sources and uncertainty

Recommended answer anatomy

For institutional, course, policy, research, and uploaded-file answers:

  • 1. Give the short answer.
  • 2. Highlight each checkable claim.
  • 3. Attach an inline citation to the exact passage.
  • 4. On selection, show the passage, title, owner, effective date, page or timestamp, and access boundary.
  • 5. Provide “Show what I used,” listing all retrieved sources—not just those the model cited.
  • 6. Mark conflicts, missing evidence, and expired sources.
  • 7. Provide “Report this answer” beside the evidence, not buried in settings.

The source panel should distinguish:

  • Approved UVU source
  • Course source
  • Your uploaded file
  • Open-web source
  • General model knowledge — not independently verified

A document title alone is not enough. The user should be able to inspect the supporting sentence without searching the entire file.

What campus systems document today

SystemPublished behaviorImportant gap
UVU Wilsonfact — current page checked September 3, 2026: uses UVU resources on the website and available course material in selected Canvas courses. It cannot access grades, private files, or tests, and students acknowledge that it is AI and may hallucinate. A Qualtrics feedback link exists. Wilson AIPublic documentation does not show inline claim citations, uncertainty states, a feedback receipt, correction status, response time, or crisis flow. Actual authenticated behavior is unknown.
Wilson course deploymentfact — February 11, 2026 interview: UVU’s CIO reported use in fact 44 courses serving 2,000 students, after a fact fall 2023 biology pilot. Wilson can link to the relevant point in a lecture recording and is intended to coach rather than give answers. EdTech interviewThis is a named executive interview, not an independent outcome evaluation. Source-checking behavior is unknown.
TritonGPTfact — August 13, 2024 release: “See Context” exposes source documents, webpages, relevance scores, and direct interaction with selected sources. UC San Diego releaseCurrent live behavior and student source-opening rates are unknown. A relevance score is not proof that a passage supports a claim.
TritonGPT course tutorsfact — 2025–26 published results: instructors can include or remove Canvas sources and select Socratic or directive behavior. Among fact 68 survey responses, fact 81% said it helped explain concepts, fact 86% found it easy, and fact 67% wanted it in future courses. Instructional programResponse rate and course mix are unknown. The survey does not measure citation checking or factual accuracy.
ZotGPT ClassChatfact — page updated May 20, 2026: faculty can upload curriculum, set instructions and guardrails, restrict access, and view activity. UCI ClassChatPublic evidence for per-claim citations, uncertainty labels, crisis language, or closed-loop corrections is unknown.
TitanGPTfact — page published August 18, 2026: available to students and employees; users are told they remain responsible for mistakes and must follow instructor policy. CSUF TitanGPTPublic source, uncertainty, correction, and crisis details are unknown.
Khanmigofact — documentation updated through August 2026: uses moderation, adult notification, limitations messaging, feedback categories, and optional severe-content administrator alerts. Safety, feedbackIt does not publish a correction service level or exact student-facing crisis script. These are vendor descriptions, not independent outcome evidence.

Measured checking behavior by students inside Wilson, ZotGPT, TritonGPT, TitanGPT, or Khanmigo is unknown. The best direct citation study above was not campus-specific. A separate undergraduate study with fact 66 students found a mean of only fact 1.76 correct classifications out of 4 after a short reference-verification lesson. Franzoni Velázquez et al., 2024

This supports two actions: make checking much easier, and teach lateral reading—opening another source or site to verify ownership and claims. In a college intervention with fact 230 students, the share that both used lateral reading and correctly judged at least one source rose from fact 7.0% at baseline to fact 61.0% in the intervention group at post-test. Brodsky et al., 2021

Error handling and feedback loops

“Report this response”

LibreChat can display thumbs-up and thumbs-down controls, but its published behavior is rating capture, not case management. LibreChat interface documentation, checked September 3, 2026

The pilot needs this flow:

  • 1. User selects Report this response.
  • 2. User chooses: wrong answer, wrong/outdated source, unsupported claim, harmful or biased, bad refusal, privacy concern, tutoring problem, or other.
  • 3. The interface previews what will be sent and lets the user omit their prompt where policy permits.
  • 4. The system records the response, model and router versions, retrieved-source identifiers, and course/source version.
  • 5. The user receives an immediate case number and status.
  • 6. A named owner reviews it.
  • 7. An approved correction updates the authoritative source or correction layer.
  • 8. The affected answer shows “Corrected” with date and explanation.
  • 9. The reporter receives closure when contact is permitted.

Planning service levels:

  • Receipt: est within 2 seconds. Basis: the plan’s existing upload-acknowledgment target.
  • Human ownership: est within 1 business day. Basis: a campus support target, not published evidence.
  • Ordinary correction or reasoned rejection: est within 5 business days. Basis: a proposed pilot operating target.
  • Urgent privacy, safety, or cross-user leakage: follow the incident process, not the ordinary queue.

Correction memory

Do not store corrections as the model’s personal “memory.”

LibreChat memory is a user-controlled key/value store, not a governed institutional correction record. LibreChat memory documentation

Use a separate correction record containing:

  • Claim or source affected.
  • Approved replacement.
  • Source owner and reviewer.
  • Effective and expiry dates.
  • Course and audience scope.
  • Reason for the change.
  • Retrieval re-index status.
  • Tests rerun.
  • Earlier answer identifiers that need a correction notice.

A student correction suggestion is evidence for review, not automatic truth.

Escalation to a person

The interface should name the destination before sending anything:

  • Course meaning or grading policy → instructor or named course staff.
  • Tutoring help → academic tutoring.
  • Account or service failure → service desk.
  • Accessibility barrier → accessibility support.
  • Privacy or data concern → privacy/security queue.
  • Immediate safety concern → emergency and crisis services.

Never say “I alerted someone” until an acknowledged route confirms receipt.

Response-time perception

The program’s current service target is est p95 first token within 10 seconds for tutoring/chat, with an est p95 queue budget within 2 seconds. Service-engineering findings

A 2026 experiment with fact 240 participants, fact 3 tasks per participant, and first-token delays of fact 2, 9, or 20 seconds found that the fact 2-second condition could be perceived as less thoughtful or useful than fact 9- or 20-second conditions, although long waits also caused frustration. Response latency study, 2026

Keep the est 10-second operational target, but do not fake thinking. Show real status:

  • “Searching approved course sources.”
  • “Checking whether the passages support the answer.”
  • “Using the calculator.”
  • “The local service is busy; you can wait or use the approved cloud route.”
  • “The answer stopped. Retry from the last completed step.”

Streaming improves perceived responsiveness and recovery. It is not evidence of correctness.

Refusals

A useful refusal has three parts:

  • 1. State the boundary.
  • 2. Give a short reason.
  • 3. Offer the closest safe action.

Examples:

  • Academic work: “I can’t complete this graded answer for you. Show me your first step, and I’ll give one hint.”
  • Grades: “I can’t assign or predict your grade. I can help compare your draft with the published rubric, one criterion at a time.”
  • Missing evidence: “I couldn’t find that in the approved course materials. I can show what I searched or help you ask the instructor.”
  • Regulated advice: “I can give general information, but I can’t provide personal legal, medical, mental-health, or financial advice. Here is the appropriate UVU service.”

In a refusal study with fact 480 participants and fact 3,840 comparisons, partial compliance—safe general help without dangerous detail—reduced negative reactions by more than fact 50% compared with a flat refusal. “Let Them Down Easy,” November 2025

Crisis message

Use the crisis pattern already developed in the duty-of-care findings:

I’m sorry you’re dealing with this. I’m an AI tutor, not a crisis service. If you or someone else may be in immediate danger, call 911 or UVU Police now. For suicide or emotional crisis, call or text 988. You do not need to repeat the details here.

This is an est design based on UVU, NIMH, and federal crisis guidance—not a claim that a bot can assess risk. Duty-of-care findings, UVU crisis services, NIMH guidance revised 2024

The assistant should stop ordinary tutoring, avoid long disclaimers, use the smallest necessary question, and never promise monitoring, confidentiality, dispatch, or response time unless those are real. Publicly validated false-positive and false-negative rates for named campus crisis bots remain unknown.

AI identity and Utah law

Utah’s current disclosure chapter took effect fact May 7, 2025. It requires disclosure on a clear request in covered consumer transactions and advance disclosure for regulated occupations’ high-risk AI interactions. Its safe harbor is broader: clearly identify the system at the outset and throughout as generative AI, not human, or an AI assistant. Whether a no-charge UVU assistant is a covered consumer transaction is unknown. Utah Code Title 13, Chapter 77

Use:

UVU AI assistant — not a human

Place it above the first prompt and keep it visible in the header, shared chats, errors, mobile views, and routed model changes.

Disclosure’s trust effect is mixed:

  • A randomized field study involving fact 11,000 truck drivers found initial voice-bot disclosure reduced response probability by about fact 11%. Management Science, August 26, 2024
  • In an online experiment with fact 194 participants, only fact 24% in the disclosed condition correctly recalled the chatbot introduction. AI & Society, January 5, 2024

These settings do not predict campus behavior. They do show that one small opening label is insufficient. Identity disclosure is a transparency and legal control, not a substitute for evidence or quality.

Conversation design for tutoring

Recommended default

The assistant should:

  • 1. Ask what the student is trying to learn.
  • 2. Ask for the student’s current work or first attempt.
  • 3. Identify one misconception or missing step.
  • 4. Give one hint.
  • 5. Wait for a response.
  • 6. Increase help gradually: prompt → hint → partial step → analogous worked example → full explanation when course policy allows.
  • 7. Ask the student to summarize or apply the idea without assistance.
  • 8. Offer a direct-explanation mode for accessibility, review, or explicit instructor-approved use.

“Socratic” should mean guided questioning, not endless questions. Students must be able to request a direct explanation where that is pedagogically and academically permitted.

Evidence

  • A preregistered field experiment with nearly fact 1,000 high-school mathematics students found that unrestricted GPT access improved assisted-practice grades by fact 48% but reduced later unaided grades by fact 17%. A safeguarded tutor improved assisted practice by fact 127% without the later loss. Its guardrails used teacher material, hints, and limits on giving answers. Bastani et al., June 25, 2025
  • A Harvard physics trial with fact 194 eligible students found better immediate post-test results from an expert-built AI tutor than from an active-learning class. The system prompt alone was not reliable enough to maintain the lesson sequence; the researchers added a custom structure that moved students through each problem part and supplied expert solutions. Scientific Reports, June 3, 2025
  • Tutor CoPilot’s preregistered field trial involved fact 900 tutors, fact 1,800 students, and more than fact 550,000 messages. Access increased lesson-topic mastery by fact 4 percentage points, with a fact 9-point gain for students of lower-rated tutors. It increased guiding questions and reduced answer-giving, but no significant end-of-year math-test gain was reported. Tutor CoPilot, revised January 26, 2025

These results support structured tutoring, not a generic “be Socratic” prompt.

“Explain your reasoning” prompt families

Ask the student to expose their reasoning:

  • “What have you tried?”
  • “Which rule do you think applies?”
  • “Where did your result stop matching the example?”
  • “Explain this step in your own words.”
  • “What evidence supports that claim?”
  • “How would you check the answer without the assistant?”

Ask the model for:

  • A concise rationale.
  • The evidence and tools used.
  • A checkable intermediate result.
  • The next instructional move.
  • A different worked example.

Do not present the model’s hidden chain of thought as reliable proof. Visible reasoning can rationalize a biased answer; one NeurIPS study found accuracy reductions as large as fact 36% across 13 tasks while explanations failed to mention the biasing prompt feature. Turpin et al., NeurIPS 2023

What LibreChat and the router can enforce

LibreChat can configure:

  • Persistent agent names and descriptions.
  • Welcome text and terms.
  • Course-specific system instructions.
  • Role permissions.
  • File search and file citations.
  • Model and endpoint availability.
  • Feedback controls.
  • Uploaded text.
  • Streaming behavior.

LibreChat model specifications, agent citation configuration, upload behavior

The router should enforce outside the model:

  • Data classification and allowed destination.
  • Course and user access to each source.
  • Local, free-cloud, or metered-frontier routing.
  • Calculator, code runner, or other verified tool use.
  • Evidence-required mode.
  • No-sufficient-source abstention.
  • Safety classification and crisis path.
  • Pinned model, prompt, retrieval, and policy versions.
  • Grade and private-file exclusion.

A custom service or interface layer is needed for:

  • Claim-to-passage highlighting.
  • “Show what I used.”
  • Source freshness, conflict, and authorization state.
  • Structured reports and correction status.
  • Human acknowledgement and escalation.
  • Governed correction memory.
  • Calibrated uncertainty state.
  • Consistent identity disclosure across all routes.
  • Sequential tutoring state that a prompt alone cannot reliably maintain.

Evaluation

Participants

  • Students: est 48 unique participants. Basis: est 4 rounds or strata of 12, covering first-generation and commuter students, writing/research work, quantitative/code work, and accessibility or English-language needs. Participants may overlap characteristics.
  • Instructors: est 12 unique participants. Basis: course-policy, source-control, tutoring, and correction-workflow review.
  • Support and safety staff: est 6 unique participants. Basis: service, accessibility, privacy, and synthetic crisis tabletop roles.
  • Total: est 66 unique participants.

This is a formative usability sample, not a powered learning-effect trial. Faulkner found that samples of fact 5 users could uncover anywhere from fact 55% to 99% of observed usability problems, which supports repeated diverse rounds rather than one tiny test. Faulkner, 2003

A formal outcome claim needs a separate power analysis after UVU obtains baseline variance; required enrollment is unknown.

Pilot tasks

Each participant receives only role-appropriate tasks:

  • Find a course rule and inspect the exact passage.
  • Ask a question the approved sources do not answer.
  • Resolve two sources that conflict by date.
  • Summarize an uploaded document containing a planted unsupported claim.
  • Judge an answer with a valid citation and one with a wrong citation.
  • Solve a math problem using the calculator and verify the result.
  • Repair code and run a provided test.
  • Complete a tutoring problem without receiving the answer first.
  • Report an incorrect answer, follow its status, and inspect the correction.
  • Encounter an academic-integrity or regulated-advice refusal.
  • Recover from delay, interrupted streaming, or route failure.
  • Staff only: run synthetic crisis and privacy scenarios.

Metrics

Measure behavior, not only opinion:

  • Task success against an instructor-approved answer key.
  • Time on task.
  • Correct acceptance of correct advice.
  • Correct rejection of wrong advice.
  • Overreliance and underreliance.
  • Confidence-to-correctness calibration and Brier score.
  • Source opens, exact-passage views, and successful source verification.
  • Claim-level citation precision and citation coverage.
  • Correct abstention when evidence is absent.
  • Trust before an error, immediately after it, and after repair.
  • Error-recovery success.
  • Report receipt, ownership, correction, and closure times.
  • First-token and complete-answer latency by route.
  • Abandonment and retry behavior.
  • System Usability Scale, a fact 10-item standardized questionnaire, plus open comments. Brooke, 1996
  • Keyboard, screen-reader, zoom, contrast, focus, and streaming-announcement task success.

Proposed launch gates

These are planning requirements, not published norms:

  • Cross-user or unauthorized-source disclosure: est 0 observed cases.
  • Claim-level citation precision: est at least 95% on the pilot gold set.
  • Citation coverage for sourceable claims: est at least 90%.
  • Abstention on unsupported institutional questions: est at least 90%.
  • Core-task success: est at least 80% in every tested student group.
  • Wrong-answer overreliance: est no more than 20% on planted-error tasks.
  • Synthetic imminent-risk message and route: est 100% correct, with est 0 unacknowledged claimed handoffs.
  • Accessibility blockers: est 0 unresolved launch-blocking defects.
  • SUS planning threshold: est at least 70, used as a comparison signal rather than proof of safety.

Basis: these are risk-based pilot gates selected to force correction before scale. UVU should revise them after the first baseline round rather than lowering them to excuse a failing interface.

Schedule inside the plan’s 14-week sequence

WeeksWork and proof
est Weeks 1–2Approve task corpus, source versions, issue categories, consent language, and research/IRB determination. Establish baseline answers and planted errors.
est Week 3Expert review of citations, uncertainty states, refusals, keyboard flow, screen-reader announcements, and source authorization.
est Week 4Moderated round with est 12 students. Measure first error, source checking, abstention, and repair discovery.
est Week 5Fix only observed high-severity problems; freeze the next test build and source set.
est Week 6Moderated round with est 12 new students and 6 instructors. Test course policy, tutoring sequence, and instructor source controls.
est Week 7Synthetic safety and incident tabletop with est 6 support staff. Prove acknowledgement and records.
est Weeks 8–10Limited field pilot with est 24 new students and 6 additional instructors. Collect aggregate behavior, not private prompt surveillance.
est Week 11Correct sources and interface defects; freeze model, router, retrieval, and prompt versions.
est Week 12Blinded challenge set covering all three model routes, missing evidence, conflicting sources, math, code, summaries, and citations.
est Week 13Retest with an est 12-person subset from the field pilot; no additional unique student count. Verify repair and accessibility regressions.
est Week 14Publish the evidence packet and decide: stop, extend pilot, or scale. No automatic scale-up.

Cross-domain learnings

  • Clinical decision support: the FDA’s fact January 2026 guidance emphasizes that a professional must be able to independently review the basis of a recommendation. UVU should apply the same idea: show inputs, sources, limitations, and known unknowns rather than asking users to trust a score. FDA guidance
  • Aviation: flight systems show mode, intended action, and limits because hidden mode changes create “automation surprise.” UVU should always show which assistant, route, evidence mode, and tools are active. FAA automation report
  • Safety engineering: NIST separates transparency from accuracy. A transparent system can still be wrong, but opacity prevents review, accountability, and repair. NIST AI RMF, January 26, 2023
  • Information literacy: lateral reading works better than teaching users to judge a page by appearance. Source links should make checking fast, but the pilot must also teach students to leave the answer and inspect independent evidence.
  • Tutoring: good tutors control the help sequence. The strongest education studies added teacher-authored material, question order, hints, and deliberate withholding. They did not rely on a personality prompt alone.

Devil’s advocate

The strongest case against this recommendation is that it creates too much interface and operational work around a system that may still be unreliable.

The evidence is fragmented. Many error studies use older models or laboratory questions. Many trust studies use medicine, forecasting, legal vignettes, or crowdsourced participants rather than UVU students. Citation panels can increase trust without checking. Uncertainty warnings can lower useful reliance. Cognitive forcing adds friction and is often disliked. A custom source, correction, and escalation layer can become a larger project than the local-model service.

UVU also already has Wilson. Adding LibreChat as a visible second front door could split support, source governance, analytics, and student expectations. The simpler safe path may be to use the new model fleet behind Wilson or within narrowly approved course tools, instead of launching another general student chat interface.

That objection is strong. It changes the implementation path, but not the need for source proof, abstention, repair, and human routes.

What would change the recommendation

  • A live UVU evaluation shows that Wilson already provides exact claim-to-passage citations, source conflicts, correction status, and acknowledged escalation.
  • The selected LibreChat release supplies the same controls without a maintained custom fork.
  • A shared UVU benchmark shows one route is reliable enough to simplify the uncertainty design.
  • Student testing shows the source panel materially harms task completion without improving verification, and a simpler passage preview performs better.
  • UVU counsel gives a narrower written interpretation of the identity requirement.
  • Accessibility testing shows the proposed inline evidence interaction is not usable and identifies a better equivalent.
  • Pilot data show that structured tutoring harms learning or produces unacceptable frustration in UVU courses.
  • A staffed feedback or crisis route cannot be funded. In that case, remove the claims and features that imply such support; do not fake the route.

What the plan should change

Ranked:

  • 1. Use one campus trust contract across Wilson, LibreChat, and every routed model.
  • 2. Do not create a competing student front door until UVU decides Wilson’s role.
  • 3. Require exact-passage citations for institutional, course, research, and uploaded-document claims.
  • 4. Add an evidence-sufficiency gate that can return “not found” or “sources conflict.”
  • 5. Build structured reporting, governed corrections, and visible correction notices.
  • 6. Keep the AI identity and active route visible throughout every interaction.
  • 7. Enforce student-first, hint-first tutoring outside the model prompt where sequence matters.
  • 8. Use calculators, code tests, and other verifiable tools for tasks that permit them.
  • 9. Add helpful refusals and a real, acknowledged human route before broad student use.
  • 10. Preserve the existing EST 10-second chat target, but add truthful status and restartable streaming.
  • 11. Run the proposed trust and usability work inside the 14-week pilot before scaling hardware or access.
  • 12. Publish results by task and route; never publish one blended “accuracy” or “trust” score.

Source log

Source-log ordinals are reference labels, not measurements.

NOT_RUN

  • not run: Saving this report to the requested output file.
  • not run: Authenticated testing of Wilson, LibreChat, ZotGPT, TritonGPT, TitanGPT, or Khanmigo.
  • not run: Benchmarking the exact quantized models, prompts, router, retrieval index, and Apple Silicon machines proposed by the plan.
  • not run: Inspecting current campus feedback queues, crisis staffing, source indexes, logs, or correction records.
  • not run: Contacting UVU, any university, or any vendor.
  • not run: A statistical power calculation for learning outcomes; baseline UVU variance and minimum meaningful effect are unknown.
  • not run: Legal advice or a UVU counsel determination about Utah’s “consumer transaction” boundary.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/45-trust-sources-interface.md in the research pack.

Y

Accessibility and language

WCAG 2.1 AA for a streaming chat, the front door's real evidence, voice, neurodivergent users, Spanish and multilingual service, phones and weak Wi-Fi, and the launch-gate checklist

Appendix Y in five lines

  1. QuestionWhat must work so disabled students, Spanish speakers, and phone users can use the service?
  2. AnswerLaunch text chat first, keep LibreChat only if the exact build passes, and test uploads, voice, language, and phones one by one.
  3. Deciding numbersApril 26, 2027 deadlinefact18–24 paid studentsest300 paired English-and-Spanish tasksest
  4. What the plan doesFreeze and test the exact build, pay disabled and bilingual testers, test English and Spanish, check weak networks, and provide accessible Hypertext Markup Language pages first.
  5. Still unknownLibreChat’s fit with access rules and current language, phone, and weak-network needs are UNKNOWN; live screen-reader, speech, Spanish, and network tests are NOT_RUN.

The plan made accessibility a launch gate with a federal deadline and said little else. A chat that streams answers, accepts uploads, reads documents, and may speak is hard to make accessible, and UVU serves many first-generation and Spanish-speaking students. This appendix names the WCAG 2.1 AA criteria that bite on a streaming chat and how to audit them with real assistive technology, tests the chosen front door's actual accessibility evidence, covers on-device voice, neurodivergent design, Spanish-language quality of the candidate models against UVU's demographics, and phones on weak Wi-Fi — with a prioritized backlog, testers, cost, and the launch-gate checklist. Research cutoff September 3, 2026; 130-plus dated sources. fact standards, dates, published measurements · est designs and priorities · unknown what UVU has not published.

What this changes in the plan

1. LibreChat moves from "chosen front door" to "conditional candidate pending exact-build proof." No public evidence shows any current release is WCAG 2.1 AA conformant; the pilot must audit the exact pinned build with screen readers (VoiceOver, NVDA, JAWS), switch access, and magnification on complete conversations, and keep an alternative front door ready. 2. Accessible text chat launches first; uploads, math, diagrams, PDF export, and voice each pass their own gate later. Streaming is treated as a visual effect: announce short states and complete messages, never every token; focus stays under the student's control. 3. Language controls are separate: interface language, content language, and reply language, each chosen by the student; a paired UVU English/Spanish benchmark for every exact deployed model and quantization — only gpt-oss-120B has a maker-published Spanish score, and it does not prove tutoring quality, so Spanish parity is a launch test, not an assumption. 4. Reading and voice support as reversible choices — read-aloud, plain-language, short-answer, and step-by-step modes — never diagnosis-based presets; on-device speech is feasible but accent, Spanish, latency, and error rates are launch tests. 5. Paid disabled and bilingual student testers, with an IRB determination before recruitment; a funded accessibility owner; a release-specific conformance report, a known-issues statement, and a barrier-response route. 6. Phones and weak Wi-Fi: captive-portal, weak-network, interrupted-upload, and data-use tests, because UVU's phone-only share and preferred-language mix are not published. 7. Generated documents ship as accessible HTML first; PDF is an additional audited artifact. The federal deadline for UVU is April 26, 2027; existing equal-access duties apply now. est

Research cutoff: fact — 2026-09-03 MDT Scope: Read-only review of the current plan, named findings, public standards, official product records, UVU pages, and published research. Method: every date, contradiction and evidence grade is kept explicit, and anything that could not be proved is marked unknown rather than smoothed over.

Executive verdict

  • Proceed with a conditional pilot, but launch accessible text chat before voice, complex uploads, diagrams, or PDF export.
  • fact — deadline April 26, 2027: UVU’s web and mobile services generally must conform to WCAG 2.1 Level A and AA; existing ADA effective-communication duties apply now.
  • UNKNOWN: Public evidence does not prove that any current LibreChat release is WCAG 2.1 AA conformant.
  • Treat streaming as a visual effect: announce short states and complete messages, not every generated token.
  • Keep focus under the student’s control and make every action work by keyboard, switch, touch, and screen reader.
  • On-device voice is feasible, but accent, Spanish, latency, and hallucination performance remain launch-test questions.
  • Offer read-aloud, plain-language, short-answer, and step-by-step modes as reversible choices; do not create diagnosis-based presets.
  • FACT: Only gpt-oss-120B has a located maker-published Spanish knowledge score; it does not prove Spanish tutoring quality.
  • UNKNOWN: UVU’s current phone-only share, preferred-language mix, poor-Wi-Fi rate, and exact accessible-chat demand are not published.
  • The plan needs a funded accessibility owner, paid disabled and bilingual testers, a pinned front-door build, and evidence for each release.

WCAG labels, release numbers, dates, source-log ordinals, and arithmetic expressions identify standards or artifacts. Measured quantities and planning quantities are graded fact, est, or unknown.

WCAG 2.1 AA for a streaming chat

The legal scope is the service, not its landing page. It includes authentication, chat history, model selection, streaming, uploads, errors, generated content, mobile layouts, and downloads. Contractor or open-source status does not transfer UVU’s responsibility. Newly generated documents normally do not receive the exception available to some older documents. DOJ rule summary

Criteria that bite

WCAG criterionChat riskHow UVU should meet itRequired proof
1.1.1 Non-text ContentUploaded images, generated diagrams, icons, charts, and formula images may have no text equivalent.Name functional icons; mark decorative images as decorative; require a useful text description for meaningful visual output; never make an image the only version of code, data, or math.Screen-reader inspection of uploaded and generated examples; compare the text equivalent with the visual meaning.
1.3.1 Info and Relationships; 1.3.2 Meaningful SequenceVisual bubbles, tables, headings, citations, and code may become an unstructured stream.Use native headings, lists, links, tables, <pre><code>, message authors, and chronological DOM order.Navigate by headings, landmarks, links, tables, and reading order with VoiceOver, NVDA, and JAWS.
1.3.4 OrientationA phone layout may force portrait mode.Support portrait and landscape unless orientation is essential.Rotate real iOS and Android devices during generation, upload, and dialogs.
1.4.3 Contrast; 1.4.11 Non-text ContrastTokens, placeholder text, model badges, focus rings, errors, and disabled controls may disappear in themes.fact threshold: normal text at least 4.5:1, large text at least 3:1, and essential controls/focus indicators at least 3:1.Measure every theme and state, including forced colors, selected, disabled, error, hover, focus, and streaming.
1.4.4 Resize Text; 1.4.10 Reflow; 1.4.12 Text SpacingLong answers, code, tables, and the composer may clip or overlap.fact thresholds: support 200% text resize and reflow at 320 CSS pixels, equivalent to a 1280-pixel viewport at 400% zoom. Confine necessary two-dimensional scrolling to a named table, code, or math region.Complete a conversation at 400% zoom and with WCAG text-spacing overrides.
2.1.1 Keyboard; 2.1.2 No Keyboard Trap; 2.1.4 Character Key ShortcutsHover-only controls, rich editors, upload widgets, menus, and code panes can trap users.Every action must work without pointer gestures. Provide visible buttons for shortcuts. Permit character-key shortcuts to be disabled or remapped.Keyboard-only and switch-scanning runs across every state; confirm escape and focus recovery from every overlay.
2.2.1 Timing Adjustable; 2.2.2 Pause, Stop, HideAuthentication expiry, long uploads, auto-scroll, streaming, animations, and disappearing notices can cause loss.Warn before expiry, allow extension where security permits, preserve drafts, provide Stop generating, disable auto-scroll when the user moves away, and respect reduced motion.Run timeout, slow-upload, network-loss, stop, resume, and reduced-motion cases.
2.4.1 Bypass Blocks; 2.4.3 Focus Order; 2.4.6 Headings and Labels; 2.4.7 Focus VisibleRepeated navigation and a growing transcript can make the latest answer hard to reach. Focus can jump to new output.Add skip links and “Jump to latest response.” Keep focus in the composer after Send. Move focus only after an explicit user action or for a blocking dialog, then restore it.Record focus before and after every action. Navigate long conversations without a pointer.
2.5.3 Label in NameVisible labels may not match accessible names, breaking speech control.Begin the programmatic name with the visible label.Operate controls by their visible names using Voice Control or another speech-input tool.
3.1.1 Language of Page; 3.1.2 Language of PartsSpanish text may be pronounced with English rules. Code, names, and quoted material may be misidentified.Set the interface language on the page and tag answer passages when language changes. Do not tag programming code as natural language.Read English, Spanish, and code-switched conversations with each screen reader and voice.
3.2.1 On Focus; 3.2.2 On InputFocusing or selecting a model may unexpectedly submit, navigate, or replace the conversation.Do not change context on focus. Tell the student before a selection changes context or clears work.Tab and switch through every control without activating it.
3.3.1 Error Identification; 3.3.2 Labels; 3.3.3 Error Suggestion; 3.3.4 Error PreventionUpload, sign-in, prompt, export, and connection errors may rely on color, disappear, or lose the draft.Identify the field or file in text, explain the correction, keep entered content, focus or link to the error summary, and confirm destructive actions.Test invalid type, oversized file, malware rejection, expired session, rate limit, model failure, and failed export.
4.1.2 Name, Role, ValueCustom buttons, toggles, menus, upload controls, and generation state may be unlabeled.Prefer native HTML. Expose name, state, selection, expansion, and disabled status.Inspect the accessibility tree and spoken output; test JAWS with AI-generated labels disabled.
4.1.3 Status MessagesGeneration, upload progress, search counts, save confirmation, and failure may be visual-only or overwhelmingly verbose.Use a restrained status region. Announce “Generating response,” important progress, failure, and “Response ready” without moving focus.Record exact speech, interruption, duplication, delay, and missing announcements across complete conversations.

Sources: WCAG 2.1, WAI-ARIA 1.2, and status-message guidance.

WCAG 2.2 is not the named federal baseline, but UVU should design now for focus not obscured, non-drag alternatives, consistent help, accessible authentication, and fact minimum 24 × 24 CSS-pixel targets. This is cheaper than retrofitting them later. WCAG 2.2

Recommended streaming behavior

  • Keep one stable semantic transcript. Each complete user or assistant message becomes one navigable item.
  • Put visual streaming inside the pending assistant message, but do not recreate the accessible node for every token.
  • Use a separate polite, atomic status for “Generating,” “Stopped,” “Connection lost,” and “Response ready.”
  • Offer announcement choices: completion only, paragraph chunks, and manual read-aloud. Completion only should be the default.
  • Keep composer focus after Send. Do not move focus into the answer automatically.
  • Provide “Jump to latest response,” “Read response,” “Stop,” and “Copy” as named buttons.
  • If the user navigates upward, stop automatic scrolling. Do not pull them back to the newest token.
  • Reserve assertive alerts for urgent, actionable failures. Repeated alerts can erase or interrupt queued screen-reader speech.
  • Mark the pending response busy only after cross-screen-reader testing; aria-busy support is not a substitute for an explicit completion message.

What mature products show

Slack documents adjustable message verbosity, replay of recent messages, reduced motion, simplified layouts, keyboard access, and support for major screen readers. GitHub publishes a current Copilot Chat ACR, yet still reports partial support for Name, Role, Value. A Gemini screen-reader user reported receiving only a completion notice and then having to hunt for the answer. A preliminary academic comparison of major generative-AI chats found continuing problems with navigation, labeling, feedback, and prompt handling for blind users. The evidence says mature teams manage announcement load and navigation, but maturity does not equal conformance. Slack accessibility, GitHub Copilot Chat ACR, blind-user AI-chat study.

Audit method

Use WCAG-EM with complete task paths:

  • 1. Freeze the exact commit, configuration, themes, translations, identity provider, browser versions, and generated-content renderer.
  • 2. Inventory every view and state, including failures and responsive variants.
  • 3. Run automated checks as a regression smoke test, not as conformance proof.
  • 4. Conduct criterion-by-criterion manual inspection.
  • 5. Run complete conversations with the assistive-technology matrix below.
  • 6. Test with disabled students using their normal devices and settings.
  • 7. Audit generated HTML and every exported artifact separately.
  • 8. Fix each blocking barrier and repeat the affected flow.

est test matrix: macOS Safari with VoiceOver; iOS Safari with VoiceOver; Windows with NVDA in Firefox and Chromium; Windows with JAWS in Chrome or Edge with AI Labeler disabled; Android Chrome with TalkBack; one-switch auto-scan; keyboard only; Windows Magnifier or ZoomText; macOS Zoom; 400% browser zoom; forced colors; reduced motion; speech input.

The conversation fixture must include sign-in, model selection, multiline prompt editing, valid and invalid uploads, a scanned document, upload progress, streaming, stop, retry, network loss, history search, rename, deletion confirmation, citations, headings, lists, links, a table, code, math, an image or diagram, read-aloud, Spanish, code-switching, session expiry, and PDF export.

The chosen front door: LibreChat

Current evidence state

EvidenceFinding
Harvard accessibility workFACT: Harvard states that it partnered with LibreChat to create and review a final Accessibility Conformance Report, or ACR. UNKNOWN: The report’s product version, scope, exceptions, test methods, findings, and remediation state were not publicly retrievable. Harvard Accessibility Summit
LibreChat release claimfact — January 28, 2026: LibreChat v0.8.2 claimed a major accessibility overhaul covering screen readers, keyboard use, focus, and contrast. This is maintainer evidence, not an independent audit. v0.8.2 changelog
Open screen-reader reportfact — opened January 6, 2025 and still open at review: issue 5187 describes hover-only controls, unlabeled controls, missing focus movement, and inaccessible image-paste actions. Some reported areas may have changed later. Issue 5187
Later reportsfact — April 2026: an enterprise user reported remaining missing labels and contrast errors in v0.8.4. fact — August 2026: six axe defects were fixed on the development branch with six bounded Chromium scenarios passing afterward. Neither result is a complete AT audit. Discussion 12532, issue 14978, PR 14979
Upload historyFACT: a closed issue documented an attachment control that did not work with Enter or Space and a native input hidden from assistive technology. The issue page does not prove which stable release fixed it. Issue 4087
Mobile riskFACT: a v0.8.5 report described iOS WebKit killing the page after sidebar and orientation changes. The issue later closed, but it establishes mobile regression risk. Issue 12824
Release driftfact — September 3, 2026: the visible newest release was v0.8.8-rc2 and marked pre-release. It adds adaptive streams, attachment-only turns, artifacts, diagram export, and downloads—new accessibility surfaces. LibreChat releases
Public conformance proofUNKNOWN: No public, downloadable LibreChat VPAT or ACR was located after searches of the official site, documentation, repository, releases, issues, and exact VPAT/ACR terms. This is not proof that none exists.

Verdict: LibreChat remains a reasonable engineering starting point, not an approved accessible front door. UVU should use a pinned fork or pinned stable build and make accessibility evidence a release gate.

What UVU must add or prove

  • Obtain the Harvard/LibreChat ACR and map it to UVU’s exact commit and enabled features.
  • Reproduce the open and historical accessibility reports against the pinned build.
  • Backport or verify the v0.8.2 work and PR 14979.
  • Replace hover-only and drag-only actions with named controls.
  • Implement the stable transcript and restrained status-channel pattern.
  • Add component and full-flow regression tests for keyboard, focus, accessible names, state, errors, and reflow.
  • Publish an accessibility statement with supported combinations, known limits, an access-barrier route, and a response owner.
  • Require the same review when LibreChat, its Markdown renderer, math renderer, upload library, authentication, or theme changes.

Uploaded and generated content

ContentAccessible handling
Upload controlStart with a labeled native file input. Put accepted type and size before the control. Drag-and-drop is only an enhancement. Announce selected file name, count, validation, progress, cancel, removal, failure, and retry. Preserve the prompt if upload fails.
Document ingestionShow name, type, size, page count where available, extraction state, OCR use, and any unsupported or unreadable pages. Never say “document read” when extraction is partial. Provide the original and extracted accessible text.
TablesUse semantic HTML with caption, header cells, and correct scope or headers relationships. Offer CSV. Use a named scrolling region for genuinely wide tables. Never export a table only as an image. W3C tables tutorial
CodeUse selectable <pre><code> text with a language label, line-wrap option, and named Copy button. Announce copy success politely. Avoid focus traps in editors and horizontal scrollers.
MathPrefer native MathML. Keep accessible LaTeX or linear text available. Add spoken explanations for complex expressions. Test current VoiceOver, NVDA with MathCAT, and JAWS rather than assuming equivalent support. MathML Core
Images and diagramsRequire alt text or a structured text explanation. Mermaid or canvas output must have an equivalent list, table, or relationship description.
Generated PDFKeep HTML as the primary accessible version. A PDF must have a descriptive title and filename, document language, tags, headings, lists, logical reading and tab order, real text, figure alternatives, table structure, descriptive links, contrast, bookmarks where useful, and OCR for scans. Test with PAC or Acrobat plus a real screen reader. Read Out Loud is not a screen reader. Section 508 PDF testing

Alternatives if LibreChat falls short

  • 1. Recommended: Build the smallest accessible web shell around UVU’s local OpenAI-compatible endpoints. Use native HTML, one transcript, one composer, one model control, and limited file support. Result: local models remain available and the surface is auditable. Risk: UVU owns the interface. Ease of undo: high because model APIs remain unchanged.
  • 2. Use a procured cloud chat with a current, scoped ACR for the accessible lane. Result: faster front-door proof. Risks: different privacy, retention, cost, model-routing, and FERPA boundaries. Ease of undo: medium.
  • 3. Switch to another open-source chat such as Open WebUI. Result: wider feature set. Risk: public reports show continuing transcript, labeling, and keyboard issues, so this is not a proven lower-risk swap. Ease of undo: medium. Open WebUI accessibility discussion, current transcript issue.

Voice and reading support

On-device speech-to-text

OptionEvidenceLimits and recommendation
Apple SpeechAnalyzer / SpeechTranscriberFACT: Apple describes on-device, low-latency streaming, meeting, distant-speech, and long-form transcription with downloadable assets. WWDC 2025 sessionUNKNOWN: Apple publishes no comparable English/Spanish accent word-error rate or end-to-end latency. Use only after checking locale assets and on-device support at runtime.
Apple SFSpeechRecognizerFACT: APIs expose whether on-device recognition is supported and can require it. Apple warns local recognition may be less accurate.Do not call Apple Speech private merely because it runs on a Mac. Fail closed or disclose when the selected locale cannot stay local.
Whisper large-class modelFACT: Whisper large-v2 reported Spanish word-error rates of 4.2% on Multilingual LibriSpeech, 5.6% on Common Voice, 8.2% on VoxPopuli, and 3.0% on FLEURS. English results on the same datasets were 6.2%, 9.4%, 7.0%, and 4.2%. Whisper paperThese are dataset results, not UVU guarantees. Accented English reached fact 18.6% on the cited VoxPopuli evaluation. Course names, noise, disabilities, and code-switching require local tests.
Smaller Whisper modelsFACT: On Multilingual LibriSpeech, tiny produced 15.7% English and 19.2% Spanish WER versus 6.2% and 4.2% for large-v2.Smaller models may improve speed while materially reducing accuracy. Do not select them by latency alone.
WhisperKitFACT: A published M3 Max test reported encoder latency improving from 612 ms to 218 ms after compression; tested WER rose from 1.93% to 2.25% on LibriSpeech-clean and from 11.55% to 12.85% on Earnings22. Streaming TIMIT mean interim latency was about 0.45 seconds and confirmed output about 1.7 seconds. WhisperKit paperThe evaluation did not prove Spanish parity or the same result on UVU’s target Macs. Interim words can change, so students must review before sending.
whisper.cppFACT: The project supports Metal and Core ML and reports more than 3× encoder speedup from Core ML compared with CPU-only operation. whisper.cppIt is an implementation claim, not an accuracy or end-to-end latency study.

Voice product rules

  • Begin with push-to-talk dictation, not an always-listening conversational assistant.
  • Show live partial text, clearly mark it as changing, and require review before Send.
  • Provide Stop, Cancel, Undo, microphone state, input-language selection, and keyboard equivalents.
  • Default to local processing. If local processing is unavailable, do not silently send audio to a cloud service.
  • Delete transient audio after transcription unless the student explicitly saves it under an approved purpose.
  • Never treat a transcript as authoritative. FACT: a study of Whisper transcripts found hallucinations in 1.7% of clips from participants with aphasia and 1.2% of control clips; fact 38% of hallucinated transcripts fell into harmful categories. The study was American English and does not establish current local-model rates. Careless Whisper
  • est launch test: record at least 40 consenting speakers across English, Spanish, accents, speech disabilities, microphone types, quiet rooms, and ordinary campus noise. The count is a coverage target, not a prevalence sample.
  • Report word-error rate, proper-name error rate, harmful insertions, language switches, interim corrections, time to first partial, time to stable text, and thermal/battery behavior by subgroup.

Text-to-speech and reading modes

Apple AVSpeechSynthesizer and macOS Spoken Content can operate on-device, select voices, adjust rate and pitch, pause, resume, stop, and highlight words or sentences. Additional voices may require downloads. UNKNOWN: Apple publishes no current Spanish-versus-English intelligibility or naturalness benchmark. Apple speech synthesis, macOS Spoken Content.

Provide:

  • Read answer, pause, resume, previous paragraph, next paragraph, stop, speed, and voice controls.
  • Synchronized highlighting that can be disabled.
  • Standard, short answer, plain language, and step-by-step views.
  • Original answer and source text beside every simplified version.
  • A warning when simplification may remove detail.
  • Downloadable text rather than audio-only export.

What helps, and what does not

  • FACT: A meta-analysis of 22 studies with pooled fact n=2,942 found a small positive read-aloud effect, fact d=.35, fact 95% CI .14–.56; publication-bias adjustment reduced it to fact d=.24. Heterogeneity was very high. Read-aloud is worth offering, not promising. Text-to-speech meta-analysis
  • FACT: A study with 170 children with dyslexia found no reading advantage for Dyslexie over Arial; most preferred Arial. A second sample contained fact 102 children with dyslexia and fact 45 controls and again found no benefit. Kuster et al.
  • FACT: In another study, a roughly fact 7% Dyslexie reading-speed advantage disappeared after Arial’s spacing was matched. Adjustable spacing is more defensible than a branded font. Marinus et al.
  • FACT: Extra-large spacing reduced errors and improved speed in a crossover study of fact 74 Italian and French children with dyslexia. This does not justify an extreme default for skilled adult readers. Zorzi et al.
  • FACT: A small Spanish study found better comprehension for shorter words among readers with dyslexia, but automatic simplification studies have not shown reliable objective-comprehension improvement. Preserve the original and make simplification optional. Rello readability study, simplification study
  • Evidence for decorative “Easy Read” pictures is mixed and can include confusion. Use visuals only when they add meaning.

Neurodivergent users

The evidence is stronger for calm, predictable, adjustable design than for diagnosis-specific AI modes. The service must not infer ADHD, autism, dyslexia, or anxiety from behavior.

NeedEvidence-based designEvidence boundary
PacingOne task or choice at a time; Stop generation; visible progress; no forced typing animation; user-selected answer length.W3C cognitive guidance is expert consensus, not a clinical treatment trial.
StructureDescriptive headings, short paragraphs, numbered steps, summary first, stable control positions, and a visible conversation outline.Helpful across conditions; not every user wants reduced detail.
PredictabilityExplain waits and model changes, warn before destructive actions, keep navigation consistent, and avoid surprise context changes.FACT: anxiety is strongly associated with intolerance of uncertainty, but interface predictability has not been proved to treat anxiety.
InterruptionDefault notifications and reminders off; no streaks, guilt, autoplay, or unrelated suggestions; respect reduced motion.FACT: in a randomized undergraduate study with fact n=221, active alerts increased reported inattention and hyperactivity with effects around fact d=.44–.45.
Memory aidsAutosave drafts, pin steps, show “where you left off,” keep a checklist, and allow optional reminders with fading.FACT: external reminders can improve immediate prospective-memory performance but may reduce independent practice effects after removal.
FocusLow-distraction view, hide optional sidebars, keep one primary action, and prevent the page from jumping during generation.An adult ADHD experiment found larger distractor costs, but its laboratory manipulation does not justify making the interface cognitively demanding.
ControlStandard, brief, plain-language, and step-by-step modes; adjustable spacing and read-aloud; easy return to the original.Diagnosis-based presets risk stereotyping and disclosure.
Human exitClear link to tutoring, Accessibility Services, Service Desk, and other appropriate people without forcing more AI dialogue.AI-specific dependency evidence for disabled college students remains thin.

Sources: W3C cognitive and learning-disability guidance, phone-notification experiment, prospective-memory reminder trial, and anxiety meta-analysis.

Dependency and distraction risks

  • Endless follow-up suggestions can turn a short task into an unbounded session.
  • Confident simplification can remove conditions or exceptions.
  • Reminders can become interruption pressure.
  • Personalization can expose or falsely infer disability.
  • A student may rely on the assistant instead of practicing planning, reading, or help-seeking.
  • Saved accessibility preferences can become sensitive profile data.
  • A calm tone can make weak advice seem more trustworthy.

Mitigations include optional session goals, visible elapsed conversation length without pressure, “finish this task” mode, reminders default-off, periodic source checks, easy export to a human helper, and no disability inference.

UVU’s existing practice

FACT: UVU Accessibility Services already offers individualized accommodations, learning specialists, organization and test-taking support, peer mentoring, alternative formats, text-to-speech, speech-to-text, screen readers, enlargement, Braille, captioning, transcription, and distraction-reduced testing. Its Accessible Technology Center asks students to bring the devices they normally use. Accessibility Services, programs, testing services.

The AI service should complement those practices, not require disclosure or replace the interactive accommodation process. FACT: UVU says information provided to Accessibility Services is protected by FERPA. Accommodation rights

Ethical testing with disabled students

  • Ask UVU’s IRB for a written determination before recruitment or recording. UVU states that pilot and feasibility studies can require review and that the IRB decides exemption.
  • Accessibility Services may distribute a neutral invitation, but should not disclose its accommodation roster to the product team.
  • Recruit by access method and functional need, not by demanding a diagnosis list.
  • Make participation voluntary, unrelated to grades, employment, service access, or accommodations.
  • Provide consent in plain English and Spanish, with text, audio, and asynchronous options.
  • Permit a support person, camera-off participation, breaks, and the participant’s own assistive technology.
  • Pay for preparation, testing, and follow-up. Prorate payment if someone stops early.
  • Treat an inaccessible session as a product failure, not a participant failure.
  • Do not publish identifiable quotations, audio, disability details, or chat content without explicit consent.

est coverage plan: 18–24 paid students across screen reader, low vision or magnification, motor or switch access, deaf or hard-of-hearing access, reading or cognitive access, neurodivergent workflows, Spanish or code-switching, and mobile or weak-network use. Overlap is welcome; this is a barrier-discovery sample, not a statistical prevalence sample.

Spanish and multilingual service

UVU population evidence

  • fact — fall 2025: UVU reports 48,669 students, 14% Hispanic, and 40% first-generation on its current Key Indicators page. UVU Key Indicators
  • FACT: The Common Data Set reports Hispanic/Latino counts of 821 of 5,104 degree-seeking first-year students, 4,054 of 28,625 degree-seeking undergraduates, and 6,660 of 47,518 total undergraduates. The calculated shares are fact 16.09%, fact 14.16%, and fact 14.02%. UVU Common Data Set
  • FACT: UVU’s First-Generation Student Success Center reports 41% and uses a stated definition tied to parents or guardians lacking a United States bachelor’s degree. The difference from 40% likely reflects source timing or cohort. Do not average them. First-generation page
  • fact — fall 2019 survey, n=1,028: 6.1% selected Spanish as a native language. fact — separate spring 2019 survey, n=939: 21.4% said they could speak Spanish. These old, self-selected surveys measure different things.
  • UNKNOWN: No current public administrative census of preferred, home, or instructional language was located. Hispanic identity must not be used as a Spanish-preference proxy.

Candidate-model language evidence

ModelPublished resultWhat it establishesVerdict
Qwen3.8-27Bfact — Spanish Arena snapshot September 2, 2026: score 1410 ±44 from 187 votes; rank 113 with a broad interval of 20–191. Spanish ArenaPairwise preference with wide uncertainty, not curricular correctness.unknown Spanish tutoring quality.
Qwen3.6-35B-A3BFACT: maker reports SWE-bench Multilingual 67.2. Model cardSoftware repositories in multiple programming languages; it does not measure Spanish instruction.unknown Spanish quality.
GLM-5.3-FlashNo exact Spanish result was located in its official card. Model cardPublished agent and coding results do not establish natural-language Spanish quality.unknown.
Mistral Small 4FACT: maker says Spanish is supported among dozens of languages but publishes no Spanish score in the card. Model cardCapability claim, not measured parity.unknown.
gpt-oss-120Bfact — MMMLU Spanish: 80.6% at low, 84.6% at medium, and 85.9% at high reasoning. fact — fourteen-language averages: 74.1%, 79.3%, and 81.3%. fact — Spanish Arena: 1365 ±21 from 854 votes, rank 173. Model card, Spanish ArenaMMMLU measures translated multiple-choice knowledge; Arena measures preference. Neither measures UVU tutoring.Best published Spanish evidence, but selection remains conditional.

The apparent Qwen3.8-versus-gpt-oss contradiction is real but not actionable: Qwen3.8 leads on a small, volatile preference sample; gpt-oss has the stronger documented knowledge benchmark. The measures are not comparable.

Quantization is another gap. FACT: gpt-oss uses native post-training with 4.25-bit MXFP4 expert weights, so its official result should not be described as full precision. UNKNOWN: no Spanish evaluation was located for the Apple-quantized builds of the other candidates.

Bilingual tutoring and code-switching

  • FACT: A 2026 meta-analysis of multilingual versus target-language-only teaching included 24 studies, 30 samples, and fact n=2,138; it reported between-group fact d=.49 and within-group fact d=1.63. It was language instruction, not AI tutoring. Zhang and Brown
  • FACT: A peer-tutoring meta-analysis covering 14 studies reported fact g=.58, fact 95% CI .22–.94; samples were mainly Spanish-speaking and school-aged. Higher-education transfer is uncertain. Romero et al.
  • FACT: A randomized online-science study with 50 Spanish-speaking eighth-grade students found benefits from bilingual support on several immediate and delayed measures. The sample was small and not college-level. Clark et al.
  • FACT: A 2026 code-switching benchmark found that inserting non-English material into English consistently reduced comprehension and reasoning accuracy, while inserting English into non-English contexts often helped. Prompt-only mitigation was inconsistent. Lost in the Mix
  • Evidence that AI code-switching improves learning is thin. A 2024 tutoring study used fact 400 simulated Chinese-English and Korean-English dialogues rather than Spanish learners or measured learning outcomes. SIGDIAL study

Language design

Keep three controls separate:

  • Interface language: menus, buttons, errors, consent, privacy, and help.
  • Content language: the source document or course material.
  • Reply language: English, Spanish, both, or follow the current prompt.

Use human-reviewed Spanish for identity, privacy, safety, accessibility, and campus-service text. Keep canonical English course terms visible with optional Spanish definitions. If a student code-switches, preserve the mix or ask a short non-blocking language question; do not force translation. Never infer language from name, ethnicity, or browser locale alone.

Required UVU language evaluation

est benchmark: 300 paired tasks, built as five UVU-relevant domains with 60 tasks per domain: course explanation, writing support, quantitative reasoning, coding or technical help, and campus-service/safety information.

For each exact deployed quantization and prompt:

  • Run matched English, Spanish, and natural code-switched variants.
  • Use est two independent bilingual graders per answer, with adjudication.
  • Score correctness, completeness, instruction following, terminology, source faithfulness, language fidelity, cultural assumptions, unsupported confidence, and safety.
  • Blind graders to model identity.
  • Separate first-answer quality from multi-turn recovery.
  • Treat Spanish-only critical errors as blockers even if the English average is strong.
  • est policy threshold: no more than a five-percentage-point paired correctness gap between English and Spanish for ordinary tutoring, plus no unresolved critical gap in campus, safety, financial, disability, or academic-policy tasks. This is a proposed governance threshold, not a published fact.

Mobile and low-bandwidth access

What is known

  • UNKNOWN: No current UVU-specific public measure of phone-only, primary-phone, or home-broadband access was located.
  • fact — national 2025 survey, n=5,022: 16% of United States adults were smartphone-dependent, including 27% of adults aged 18–29, 28% of Hispanic adults, and 34% of adults with household income below $30,000. These are national, not UVU, figures. Pew
  • fact — one Hispanic-Serving Institution study, n=2,188: optimal smartphone-plus-computer access was 72% among first-generation students and 85% among non-first-generation students; unstable internet was reported by 30% and 28%. This supports risk planning, not a UVU estimate. Digital inequality study
  • fact — multi-campus sample, n=2,913: 7% reported losing internet for inability to pay, 28% hit a mobile-data cap, and 18% lacked a working laptop for at least ten nonconsecutive days during the prior year. The study occurred during the pandemic. PLOS ONE
  • FACT: UVU offers encrypted Eduroam on major desktop and mobile platforms. Its open Wolverine-WiFi does not fully protect device-to-access-point traffic.
  • FACT: Guest Wi-Fi uses a captive portal and grants eight hours before re-registration.
  • FACT: UVU lends laptops and hotspots free to current-semester students on a first-come basis. UNKNOWN: Public inventory and unmet demand were not located.
  • FACT: UVU’s Service Desk has a Spanish chat link, but continuous ticket intake is not continuous staffed support.

Sources: UVU Eduroam, Wolverine-WiFi, guest instructions, equipment checkout, and Service Desk.

Phone and weak-network requirements

  • Complete every core task at 320 CSS pixels in portrait and landscape.
  • Test iOS Safari with VoiceOver and Android Chrome with TalkBack on real devices.
  • Keep all primary actions visible without hover.
  • Use a text-first page and lazy-load optional media.
  • Do not autoplay audio, video, diagrams, or long read-aloud output.
  • Show file size before upload and permit cancel, resume, and retry.
  • Save unsent drafts locally with a privacy warning and clear-delete control.
  • Make request retries idempotent so a reconnect does not submit twice.
  • Expose Offline, Stale, Syncing, Synced, Upload paused, and Generation lost states in text and to assistive technology.
  • Cache the application shell, help, and authorized static material. Do not claim that model answers are available offline when the inference server is unreachable.
  • Preserve the completed portion of an answer after a dropped stream and clearly mark whether generation can resume.
  • Provide accessible HTML or text downloads that are smaller than PDF when possible.
  • UNKNOWN: Data cost per conversation remains unmeasured. Instrument payload bytes without collecting prompt content, publish upload size before transfer, and offer a Wi-Fi-only upload option.

Plan

Prioritized backlog

RankPilot workEvidence required
P0-AFreeze a stable LibreChat commit or select the small accessible shell. Name one accessibility owner.Exact commit, configuration, ownership, and supported-browser record.
P0-BImplement the stable transcript, restrained status region, stable composer focus, Stop, jump/read-latest, and no forced auto-scroll.Complete VoiceOver, NVDA, JAWS, TalkBack, keyboard, and switch runs.
P0-CFix names, states, focus, contrast, reflow, timeouts, errors, reduced motion, and language tagging.Manual WCAG 2.1 A/AA audit plus automated regression checks.
P0-DShip English and human-reviewed Spanish interface, privacy, errors, safety, and help.Bilingual review and language-tag testing.
P0-ERun the exact deployed models through the paired UVU English/Spanish evaluation.Versioned prompts, answers, rubrics, grader agreement, adjudication, and release decision.
P0-FRecruit and pay disabled, bilingual, and low-bandwidth student testers after UVU’s IRB determination.Consent, compensation record, de-identified findings, fixes, and retests.
P1-AAdd accessible upload state and extracted accessible HTML/text.Valid, invalid, oversized, scanned, partial-extraction, retry, and cancellation tests.
P1-BAdd semantic tables, code, citations, and math.AT and 400%-zoom conversation fixture.
P1-CAdd on-device dictation and system read-aloud.Locale, privacy, WER, latency, harmful-insertion, and control tests.
P1-DAdd PDF export only after tagged-PDF proof; otherwise offer HTML, DOCX, or text.PDF structure inspection and screen-reader task completion.
P2-AAdd broader multilingual support only after separate language evaluations and human-reviewed interface text.Per-language evidence record.
P2-BMaintain an ACR and accessibility statement for each supported release.Updated scope, exceptions, methods, known limits, and remediation dates.

Testers and cost

ItemPlanning basisCost
Accessibility remediation and test automationest 600–1,200 staff hours, based on six major workstreams at roughly 100–200 hours each.unknown dollars: UVU loaded labor rates and current LibreChat condition were not supplied.
Independent expert auditManual WCAG and AT audit of production-like flows plus retest.est $20,000–$50,000 placeholder; not a quote. Scope and procurement rates can change it materially.
Disabled and bilingual student testingest 18–24 students × two paid hours × $40/hour.est $1,440–$1,920, plus unknown interpreter, transport, attendant, equipment, and accommodation costs.
Spanish model evaluationest 300 tasks × two graders × ten minutes = 100 grading hours, plus est 40 hours adjudication and setup.est $4,900 at an assumed est $35/hour; practical budget est $5,000–$9,000.
Speech evaluationest 40 participants, recruitment, transcription references, and analysis.est $4,000–$12,000; basis is paid participation plus specialist review, not a vendor quote.
Scale operationsRelease review, regression maintenance, issue triage, student feedback, and annual independent check.est 0.5–1.0 accessibility FTE plus est $15,000–$40,000 external review per year.

Launch-gate checklist

Launch only when all applicable items are PASS:

  • Exact production commit, configuration, themes, translations, renderer, authentication, and model routes are frozen.
  • WCAG 2.1 Level A and AA manual audit covers complete processes.
  • No unresolved defect blocks independent sign-in, prompt entry, response reading, Stop, error recovery, history, or sign-out.
  • VoiceOver, NVDA, JAWS, TalkBack, keyboard, switch, magnification, forced colors, text spacing, and reduced-motion flows pass.
  • Status announcements are neither missing nor duplicated or continuously interrupting.
  • Focus never moves merely because tokens arrive.
  • Every visible action has a matching accessible name and exposed state.
  • The service reflows at 320 CSS pixels and remains operable at 400% zoom.
  • English and Spanish interface content has human review and correct page/part language metadata.
  • The exact deployed model and quantization pass the UVU paired language gate.
  • Upload failure never loses the prompt; extraction limits are disclosed.
  • Tables, code, math, diagrams, citations, and links work in the complete AT conversation.
  • PDF export is disabled unless the produced file passes its independent audit.
  • Dictation is disabled unless locale, local-processing, transcript-review, accuracy, latency, and deletion gates pass.
  • Weak-network retry does not submit twice or hide whether a message was received.
  • Paid disabled-student testing has been completed and blocking barriers have been fixed and retested.
  • Accessibility statement, barrier-report route, support owner, and rollback path are live.
  • Each later release repeats affected checks before deployment.

Cross-domain learnings

  • Air-traffic and operations consoles: status is separated from the operator’s working focus. Chat should announce state without stealing focus.
  • Captioning and speech recognition: partial output is visibly useful but unstable. Final text should be confirmed before it becomes an instruction, record, or submitted prompt.
  • Localization systems: interface locale, source-content language, and output language are separate controls. Human-approved terminology is retained for high-risk text.
  • Mobile banking and resilient forms: autosave, idempotent retry, clear receipts, and explicit pending states prevent duplicate or lost actions.
  • Document publishing: structured source content produces more reliable accessible outputs than repairing PDFs after generation.
  • Participatory disability research: standards inspection and lived-experience testing answer different questions. Disabled users are paid experts, not free final-stage validators.
  • Learning support: optional scaffolds can help, but permanent prompts and reminders may reduce independent practice. Controls should be adjustable and fadeable.
  • Safety-critical alarms: frequent alerts become noise. Screen-reader announcements need priority, acknowledgement, and restraint.

Devil’s advocate

The strongest case against this recommendation is to avoid a custom campus front door altogether.

LibreChat’s public conformance proof is incomplete, dynamic chat is unusually difficult to audit, accessible document conversion is expensive, Spanish benchmark coverage is poor, and voice creates new privacy and accuracy risks. UVU already provides Read&Write, screen readers, alternative formats, accessible testing, learning specialists, and other human support. A custom interface could consume scarce funds while duplicating mature cloud tools and still failing students at the edges.

Under that view, UVU would procure one established cloud chat with a current scoped ACR, restrict it to low-risk content, retain existing accessibility services, and delay local-model access until an accessible interface exists.

That is a credible alternative. Its weakness is that it changes the plan’s privacy, model-control, cost, and local-compute goals, while a vendor ACR still does not prove the complete UVU service.

What would change the recommendation

  • A current, independently reviewed LibreChat ACR covering UVU’s exact release and complete workflows could reduce the need for a separate shell.
  • Reproduction showing that LibreChat’s open screen-reader reports are fixed in a stable release could lower remediation scope.
  • A different open-source front door passing the complete UVU AT audit could replace LibreChat.
  • Current UVU device and language research showing little phone, Spanish, or weak-network need could change priority, though not ADA obligations.
  • Exact Apple-locale tests showing strong private English and Spanish speech performance could move dictation earlier.
  • Exact deployed-model testing showing reliable Spanish parity could permit broader Spanish launch.
  • A federal rule change could alter the compliance date or standard; it would not remove existing equal-access duties.
  • Insufficient staffing or budget should narrow the pilot to accessible text chat, not lower the gate.

What the plan should change

  • 1. Replace “WCAG 2.1 AA by launch” with the complete process, AT matrix, tester, evidence, and release-gate requirements above.
  • 2. Change LibreChat from “chosen front door” to “conditional implementation candidate pending exact-build proof.”
  • 3. Make accessible text chat the first release; gate uploads, math, diagrams, PDF, and voice separately.
  • 4. Add independent interface language, content language, and reply-language controls.
  • 5. Add the paired UVU English/Spanish benchmark for every exact deployed model and quantization.
  • 6. Fund paid disabled and bilingual student participation and obtain UVU’s IRB determination before recruitment.
  • 7. Add mobile, captive-portal, weak-network, interrupted-upload, and data-use tests.
  • 8. Require accessible HTML as the primary generated document; treat PDF as an additional audited artifact.
  • 9. Assign an accessibility owner and maintain a release-specific ACR, known-issues statement, and barrier-response process.
  • 10. Reserve the metered frontier pool as an accessible quality fallback only after its privacy, language, and front-door path passes the same gates.

Source log

Source numbers are identifiers, not measured claims.

NOT_RUN

  • NOT_RUN: No UVU or vendor contact.
  • NOT_RUN: No live UVU or LibreChat deployment inspection; exact build and configuration remain unknown.
  • NOT_RUN: No VoiceOver, NVDA, JAWS, TalkBack, switch, magnification, keyboard, or PDF task run.
  • NOT_RUN: No accessibility defect reproduction against LibreChat.
  • NOT_RUN: No download or review of the Harvard/LibreChat ACR because no public copy was located.
  • NOT_RUN: No Apple-Silicon speech benchmark, local model run, quantization comparison, or Spanish tutoring benchmark.
  • NOT_RUN: No UVU student recruitment, recording, survey, compensation, or IRB submission.
  • NOT_RUN: No campus Wi-Fi coverage, latency, outage-history, hotspot-inventory, phone-share, or data-cost measurement.
  • NOT_RUN: No deployment, publication, procurement, account change, or external send.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/46-accessibility-language.md in the research pack.

Z

Roadmap and timing

Apple's chip cadence, the open-model trajectory by memory tier, the cost of waiting, gated purchase waves, what could break the trajectory, and a 24-month watchlist

Appendix Z in five lines

  1. QuestionShould the university buy machines now or wait for the next chip, and how should later purchases be staged?
  2. AnswerBuy four machines only after one delivered unit passes every test, buy later waves only when use proves the need, and do not commit to 26 now.
  3. Deciding numbersabout 15 months of delayest390 fleet device-monthsfactabout $8,916 all-inest
  4. What the plan doesUse a delivered-machine gate, buy in three waves, retest model fit, compare local and cloud costs, and review replacement in year three.
  5. Still unknownThe new machines have not been tested, the 512-gigabyte price and state compute share are UNKNOWN, and live security, access, and recovery checks are NOT_RUN.

A finance committee will ask "why not wait for the next chip?" and "what if the models outgrow the boxes?" This appendix answers with dated evidence: every Mac mini and Mac Studio chip since 2020 with bandwidth, memory, and launch price; Apple's stated AI direction and the credible reporting on what comes next; the best open model that fit each memory tier, quarter by quarter; the cost of waiting in device-months and stream-months; an options-style purchase plan in gated waves; what could break the trajectory; and a dated 24-month watchlist with the decision each event triggers. Research cutoff September 3, 2026; 150-plus dated sources. fact release facts and arithmetic · est thresholds and wave sizes · unknown rumors, so labeled.

What this changes in the plan

1. "Buy after September 22" becomes a delivered-unit acceptance gate. Shipment is availability, not proof: one delivered machine must load the selected model, complete the 1k/8k/32k-token tests, sustain at least four simultaneous streams at the 10-token-per-second floor, pass batch 1/2/4/8, cold-load, peak-memory, and 24-hour soak tests, plus identity, logging, isolation, accessibility, and recovery checks. 2. Buy in waves, not 26 at once: wave one, four minis (about $8,916 all-in) for a 30-faculty cohort; wave two, up to eight more, only after half the pilot participants return at their next natural task within 30 days, measured demand uses at least 60% of capacity for four teaching weeks, support and incidents stay inside the staffed plan, and no model release has changed the best tier; wave three, up to the planned 26, after spring results, observed queue times, and a refreshed local-versus-cloud comparison. Do not reserve 26 for symmetry. 3. Waiting has a price. The rumored next Pro-class chip implies about 15 months of delay, forgoing 60 validation device-months and 390 fleet device-months; a hypothetical machine 25% faster repays no more than 9.6 months of waiting inside a 48-month service life. Delay only if a must-have requirement fails, Apple officially announces within 90 days a machine with the required memory or 35% more measured throughput for no more than 20% more cost, a dated grant award within 60 days would cover a material share, or accessibility or security testing blocks the service. 4. Memory rule: buy for models that create useful service now; the biggest flagships already exceed 512 GB and belong in the metered pool. 5. The 512 GB line stays unpriced until Apple posts the configuration and at least two approved research workloads need a model that cannot fit 256 GB — and only at a price no higher than about six all-in minis. 6. Year three is a review with real resale quotes; year four a refresh only where measured gains justify it; older machines become development, spare, or low-priority capacity. 7. Dual-runtime acceptance (MLX and llama.cpp interchangeable behind one service contract) and a quarterly open-model checkpoint on one benchmark version. The 24-month calendar of Apple events, model checkpoints, grant deadlines, state milestones, and the April 26, 2027 accessibility deadline is at the end. est

Prepared independently by Utlyze. Not a publication of, endorsed by, or affiliated with Utah Valley University. As of September 3, 2026, MDT.

Evidence labels: fact means a dated source or arithmetic from sourced figures. est means a planning estimate whose basis is stated. unknown means the needed evidence was not public or could not be verified. Source-log numbers are reference identifiers, not quantitative claims.

Executive verdict — ten lines

  • 1. Buy the four-machine validation set after the September 22, 2026 shipment date, provided one delivered machine passes the full service test.
  • 2. Do not buy the full 26-machine fleet at once; preserve the option to stop, expand, or change tiers after real use data.
  • 3. Waiting for the rumored next Pro-class chip would likely mean about 15 months of delay est: named reporting says late 2027.
  • 4. That delay would forgo 60 validation device-months and 390 fleet device-months fact: arithmetic.
  • 5. It would also forgo 240–480 validation stream-months and 1,560–3,120 fleet stream-months est: four-plan versus eight-tested streams per machine.
  • 6. The historical Mac mini interval is 22.4 months at the median fact, but the most recent Pro memory ceiling did not rise beyond 64 GB fact.
  • 7. A hypothetical 25% faster next machine needs about 9.6 months of waiting to repay itself inside a 48-month service horizon est inputs; FACT arithmetic; a 15-month wait is too long.
  • 8. Open models are improving quickly, but the largest flagships already exceed 512 GB; UVU should buy useful service tiers, not try to own every frontier model.
  • 9. Review resale and workload evidence in year three; refresh in year four only where measured service gains justify it.
  • 10. The plan should replace its fixed four-stream assumption with delivered tests and make adoption, accessibility, price, and model releases explicit expansion gates.

# 1. Apple silicon cadence

Mac mini and Mac Studio releases, 2020–2026

Dates below are availability dates. Prices are original US retail starting prices, not UVU’s configured or education prices.

Machine and chipAvailableUnified-memory bandwidthMaximum memoryLaunch priceEvidence
Mac mini, M1Nov. 17, 2020 fact68.25 GB/s fact-secondary; Apple did not publish it on the launch page16 GB fact$699 factApple launch, Apple specifications
Mac Studio, M1 MaxMar. 18, 2022 fact400 GB/s fact64 GB fact$1,999 factApple Studio launch, M1 Max
Mac Studio, M1 UltraMar. 18, 2022 fact800 GB/s fact128 GB fact$3,999 factApple M1 Ultra
Mac mini, M2Jan. 24, 2023 fact100 GB/s fact24 GB fact$599 factApple M2 mini
Mac mini, M2 ProJan. 24, 2023 fact200 GB/s fact32 GB fact$1,299 factApple M2 mini
Mac Studio, M2 MaxJune 13, 2023 fact400 GB/s fact96 GB fact$1,999 factApple Studio launch, M2 Ultra specifications
Mac Studio, M2 UltraJune 13, 2023 fact800 GB/s fact192 GB fact$3,999 factSame Apple sources above
Mac mini, M4Nov. 8, 2024 fact120 GB/s fact32 GB fact$599 factApple M4 mini
Mac mini, M4 ProNov. 8, 2024 fact273 GB/s fact64 GB fact$1,399 factSame Apple source above
Mac Studio, M4 MaxMar. 12, 2025 fact410–546 GB/s fact: chip option128 GB fact$1,999 factApple 2025 Studio
Mac Studio, M3 UltraMar. 12, 2025 fact819 GB/s fact512 GB fact$3,999 factSame Apple source above
Mac mini, M6Sept. 22, 2026 fact: announced availability153–170 GB/s fact: configuration32 GB fact$899 factApple 2026 mini
Mac mini, M5 ProSept. 22, 2026 fact: announced availability307 GB/s fact64 GB fact$1,699 factSame Apple source above
Mac Studio, M5 MaxSept. 22, 2026 fact: announced availability460–614 GB/s fact: chip option128 GB fact$2,499 factApple 2026 Studio
Mac Studio, M5 UltraSept. 22, 2026 fact: announced availability1,200 GB/s fact512 GB fact$5,499 factSame Apple source above

The table covers every released Mac mini generation and every Mac Studio chip option through the announced September 2026 shipment wave. Apple did not sell an M1 Pro Mac mini, an M3-series Mac mini, an M4 Ultra Studio, or an M5 base mini fact.

What the cadence says

Mac mini availability intervals were:

  • M1 to M2: 26.2 months fact: arithmetic
  • M2 to M4: 21.5 months fact: arithmetic
  • M4 to the September 2026 generation: 22.4 months fact: arithmetic
  • Median: 22.4 months; mean: 23.4 months fact: arithmetic

Mac Studio intervals were 14.9, 21.0, and 18.4 months fact: arithmetic, giving an 18.4-month median and 18.1-month mean fact: arithmetic.

The intervals are irregular enough that “wait one more quarter” is not a reliable purchasing strategy.

Maximum memory has risen in large steps over the full history, but it often stays flat for a generation:

  • Pro mini: 32 GB to 64 GB to 64 GB fact
  • Max Studio: 64 GB to 96 GB to 128 GB to 128 GB fact
  • Ultra Studio: 128 GB to 192 GB to 512 GB to 512 GB fact

The next chip therefore cannot be assumed to bring a larger memory tier.

Bandwidth trend

Generation-to-generation maximum-bandwidth changes were:

  • Base mini: +46.5%, +20.0%, then +27.5% using the lower M6 configuration or +41.7% using its highest configuration fact: arithmetic.
  • Pro mini: +36.5%, then +12.5%; median gain 24.5% fact: arithmetic.
  • Max Studio: 0%, +36.5%, then +12.5% using the highest chip options fact: arithmetic.
  • Ultra Studio: 0%, +2.4%, then +46.5% fact: arithmetic.

A 25% next-generation Pro gain is therefore a reasonable central planning case est: recent Pro median rounded. A 12%–37% range better represents the evidence est: bounded by the two observed Pro transitions.

Bandwidth is important because model weights and the growing attention cache must be read from memory during generation. It is not the sole speed measure: software support, prompt length, active parameters, quantization and batching also matter.

Apple’s stated AI direction

Apple’s M5 announcement put a neural accelerator in every GPU core and claimed as much as 3.5 times M4 performance on selected AI workloads fact: Apple claim, not independent testing. The M5 Pro and Max continued that design and Apple claimed up to four times peak GPU AI compute fact: Apple claim. M6 and M5 Ultra continued the same direction while increasing memory bandwidth to as much as 1,200 GB/s fact. See Apple’s M5 direction, M5 Pro and Max and M6/M5 Ultra.

Apple’s “4.8 times” LM Studio statement for the new mini concerns prompt processing fact: Apple claim. It is not proof of 4.8-times text-generation speed or eight simultaneous production sessions.

What reportedly comes next

Named reporting by Bloomberg’s Mark Gurman, summarized by MacRumors on June 25 and July 12, 2026, says:

  • Base M7 machines may arrive in the first half of 2027 [RUMOR].
  • M7 Pro and M7 Max may arrive near the end of 2027 [RUMOR].
  • M7 Ultra may follow in 2028 [RUMOR].
  • Apple may skip M6 Pro and M6 Max [RUMOR].

Sources: June report and July report.

Those reports correctly anticipated parts of the August 2026 product line, which raises their credibility, but the remaining dates and specifications are still rumors. They are not suitable as purchase commitments.

Apple has announced a September 9, 2026 event fact, but it has not promised a Mac announcement for that event unknown. See Apple’s event notice.

Memory-price pressure and Apple’s price changes

TrendForce forecast conventional PC DRAM contract prices rising more than 100% quarter over quarter in the first quarter of 2026 fact: Feb. 2 forecast. In July, it projected another 13%–18% quarter-over-quarter conventional DRAM increase and a 10%–15% NAND flash increase in the third quarter fact: July 3 forecast. Forecasts can be wrong, but the direction matches Apple’s actions. Sources: TrendForce February and TrendForce July.

In March 2026, Apple temporarily removed the 512 GB Studio option and raised the 256 GB upgrade from $1,600 to $2,000 fact: contemporaneous reporting. In June, Apple attributed broader increases to memory and storage costs driven by AI data-center demand fact: reported Apple statement.

Compared with the previous generation’s original starting price, the September line is higher by:

  • Base mini: $300, or 50.0% fact: arithmetic
  • Pro mini: $300, or 21.4% fact: arithmetic
  • Max Studio: $500, or 25.0% fact: arithmetic
  • Ultra Studio: $1,500, or 37.5% fact: arithmetic

The amount of each increase caused specifically by memory is unknown; processors, storage configurations and product positioning also changed.

The 512 GB M5 Ultra price is due in late October 2026 fact: Apple timing, but its exact day and final configured price are unknown.

# 2. Open-model trajectory

Method and warning

The following table asks a narrow question: what was the highest-scoring downloadable open-weight model released by each quarter’s end whose four-bit weights could fit inside 48, 64, 128, 256 or 512 GB?

“Four-bit” means roughly half a gigabyte per billion parameters. Where an official artifact size was unavailable, the fit estimate is:

total parameters × 0.5 GB × 1.10 overhead est.

This is only a weight-fit test. It does not reserve memory for macOS, the model’s working cache, long prompts, simultaneous users or the server. The service plan should retain its 15% operating reserve est from the machine-matrix method.

Scores are the currently reported Artificial Analysis Intelligence Index v4.1.1 scores observed September 3, 2026 fact: provider-reported, sometimes model-estimated. Applying one current score version avoids joining incompatible index versions. The winning-model choice remains est, because no public catalog proves it contains every downloadable model and every four-bit artifact.

Best open model by weight-only tier

Quarter48 GB64 GB128 GB256 GB512 GB
2024 Q1Solar Mini — 6Solar Mini — 6Solar Mini — 6Solar Mini — 6Solar Mini — 6
2024 Q2Qwen2 72B — 6Qwen2 72B — 6Qwen2 72B — 6Qwen2 72B — 6Qwen2 72B — 6
2024 Q3Qwen2.5 72B — 9Qwen2.5 72B — 9Qwen2.5 72B — 9Qwen2.5 72B — 9Qwen2.5 72B — 9
2024 Q4QwQ Preview 32B — 9QwQ Preview 32B — 9QwQ Preview 32B — 9QwQ Preview 32B — 9DeepSeek V3 — 14
2025 Q1QwQ 32B — 13QwQ 32B — 13QwQ 32B — 13QwQ 32B — 13DeepSeek R1 — 19
2025 Q2QwQ 32B — 13QwQ 32B — 13QwQ 32B — 13QwQ 32B — 13DeepSeek R1-0528 — 20
2025 Q3gpt-oss-20b — 15gpt-oss-120b — 24gpt-oss-120b — 24gpt-oss-120b — 24gpt-oss-120b — 24
2025 Q4Qwen3 Next 80B — 17gpt-oss-120b — 24gpt-oss-120b — 24GLM-4.7 — 34GLM-4.7 — 34
2026 Q1Qwen3.5 35B — 30Qwen3.5 35B — 30MiniMax M2.5 — 34Qwen3.5 397B — 34GLM-5 — 41
2026 Q2Qwen3.6 27B — 38Qwen3.6 27B — 38Qwen3.6 27B — 38MiniMax M3 — 45GLM-5.2 — 53
2026 Q3 through Sept. 3Qwen3.8 27B — 52Qwen3.8 27B — 52Qwen3.8 Flash Next — 56GLM-5.3 Flash — 57GLM-5.3 — 60

Every displayed score is fact: current index report; every “best” selection is est: documented method.

Important operating corrections:

  • gpt-oss-120b’s 62.4 GB artifact fits 64 GB by weight only fact, but is not a sound 64 GB production choice after operating reserve and prompt cache.
  • MiniMax M2.5’s estimated 126.5 GB footprint barely fits 128 GB est; it is not a practical 128 GB campus service.
  • The current deployable choices from the machine matrix are Qwen3.8 27B on 48/64 GB, Qwen3.8 Flash Next on 128 GB, GLM-5.3 Flash on 256 GB and GLM-5.3 on 512 GB.
  • Kimi K3’s official four-bit package is about 1.56 TB fact; DeepSeek V4 Pro is about 850 GB fact. Neither fits a 512 GB Mac.
  • UVU does not need every frontier model locally. It needs useful local tiers plus a controlled cloud path for models that exceed them.

Compact-model progress

The strongest approximately 64 GB-or-smaller series in the existing research was:

  • Qwen2.5 72B, September 2024: score 9 fact
  • Qwen3 Next 80B, September 2025: score 17 fact
  • Gemma 4 31B, April 2026: score 30 fact
  • Qwen3.8 27B, August 2026: score 52 fact

A simple four-point regression gives 20.62 index points per year, with an R² of 0.825 fact: arithmetic. It projects 65.12 in September 2027 and 85.78 in September 2028 est.

Those point forecasts should not be used in a budget. The sample has only four models fact; Artificial Analysis changed its index repeatedly; the scale may have a ceiling; and model selection creates survivor bias. A safer September 2027 compact-tier planning range is 58–66 est: regression center with a wide judgment discount.

Open versus closed

Observation dateEvidenceGap
Nov. 4, 2024Epoch broad capability reviewAbout 12 months est from study
Oct. 30, 2025Epoch ECI comparisonAbout 3 months and 7 ECI points est from study
May 29, 2026Epoch updateAbout 4 months; 90% interval 7–11 ECI points; strict crossing about 6 months est from study
May 1, 2026NIST evaluation of DeepSeek V4 ProAbout 8 months on its particular suite est from NIST analysis
Sept. 3, 2026Current best open versus leading closed index result60 versus 66, a 6-point gap fact: current index
Sept. 3, 2026Chatbot Arena comparison1,489±5 versus 1,507±5, an 18-point gap fact: current leaderboard

Epoch’s broad “best open” estimate and NIST’s named DeepSeek result are not mutually exclusive: they measure different model sets and tasks. Together they support a present planning range of roughly 3–8 months est, not one exact lag.

Artificial Analysis changed from index v1 to v2 in February 2025, then through v2.1, v2.2, v3, v4, v4.1 and v4.1.1 by August 2026 fact. Scores from those versions must not be plotted as one uninterrupted historical line.

Release cadence by maker

MakerObserved releasesCadence conclusion
QwenAt least 7 material families or updates from Sept. 2024 through Aug. 2026 factRoughly one material release every four months est, with bursts
Z.aiGLM-4.5, 4.6, 4.7, 5, 5.1, 5.2 and 5.3 from July 2025 through Aug. 2026 factMedian observed gap about 63 days, or 2.1 months fact: arithmetic
DeepSeekV3, R1 and several major revisions through V4 Pro factIrregular; about 3–6 months between major capability steps est
MoonshotK2 through K3 across six named releases from July 2025 through July 2026 factMedian observed gap about 82 days, or 2.7 months fact: arithmetic
OpenAI open weightsOne family with 20B and 120B sizes on Aug. 5, 2025 factFuture cadence unknown
MistralMistral 3 on Dec. 2, 2025 and Small 4 on Mar. 16, 2026, plus specialists factGeneral-model cadence is irregular and unknown
MetaLlama 3, 3.1, 3.3 and 4 from Apr. 2024 through Apr. 2025 fact; Muse Spark remained a private preview in 2026 factNo new downloadable general family for about 17 months by Sept. 2026 fact; restart date unknown

Official sources include Qwen2.5, Qwen3, Z.ai release notes, DeepSeek updates, Moonshot’s release index, gpt-oss, Mistral 3, Mistral Small 4, Llama 4 and Muse Spark.

Plausible next 12–24 months

These are planning judgments, not maker promises:

  • At least two meaningful new models that fit the 128 GB weight tier by September 2027: 75% likelihood est: Qwen, Z.ai and Moonshot cadence.
  • A 48/64 GB model reaching approximately 58–66 on the current index by September 2027: 60% likelihood est: compact-model trend, discounted for benchmark changes.
  • At least one high-profile new flagship still exceeding 512 GB: 80% likelihood est: Kimi K3, DeepSeek V4 Pro and Qwen flagship direction.
  • Tool use, image input and longer contexts becoming normal in compact models: 80% likelihood est: current maker roadmaps and releases.
  • A fully permissive license becoming standard across the strongest models: below 50% likelihood est: current mixture of Apache, MIT and custom licenses.

The planning implication is important: smaller models are likely to make today’s machines more useful, even while the largest models remain too large. Model progress does not automatically make the hardware obsolete.

# 3. Cost of waiting, in numbers

Capacity forgone each month

The current plan assumes four simultaneous streams per 48 GB mini est. Delivered-model research estimates eight streams at at least ten generated tokens per second for the current 27B Qwen model est: not yet tested on delivered M5 Pro hardware.

DelayFour-machine validation set26-machine fleet
One month4 device-months fact; 16–32 stream-months est26 device-months fact; 104–208 stream-months est
15 months to a rumored late-2027 Pro generation60 device-months fact; 240–480 stream-months est390 device-months fact; 1,560–3,120 stream-months est
22.4-month historical mini median89.6 device-months fact; 358.4–716.8 stream-months est582.4 device-months fact; 2,329.6–4,659.2 stream-months est

A stream-month is one model-serving slot available for one month. It is a capacity measure, not proof of demand or student benefit.

The plan prices four machines at $8,916 and 26 machines at $57,954 est: education configuration plus $90 network upgrade per unit. Spread over 48 months, those amounts are $185.75 and $1,207.38 per calendar month est inputs; FACT arithmetic.

Those amounts are not the economic loss from waiting. UVU retains the cash when it waits. The monetary value of missed work is unknown until the pilot measures demand, successful tasks, time saved and outcome value.

Adoption momentum

The behavioral findings show that access does not equal repeat use:

  • CSU reported at least 250,000 activations by spring 2026 fact, but activation is not proof of sustained use.
  • One reported measure found student training at about 0.7% and faculty training at 16% fact: different populations and denominators.
  • CSU Bakersfield recorded 37 completions among 43 participants offered a $500 incentive, or 86% fact: arithmetic, but it had no control group.
  • Virginia Tech reported typical weekly use among 78% of 425 surveyed users fact; this is a selected-user result, not a campus rate.
  • Manchester’s reported 90% measure used a 30-day definition that was not public in enough detail to reproduce fact rate; UNKNOWN definition.

For a 30-faculty cohort, a one-month delay removes 30 faculty-months of scheduled exposure est: arithmetic under a fixed cohort. Missing the roughly four-month spring teaching period removes about 120 faculty-months est. It does not prove 120 lost completions or any exact learning loss.

The sound adoption measure is whether a participant returns at the next natural teaching task within 30 days est recommendation, not whether an account was activated or used every week.

September shipment and course preparation

The machines become available September 22, 2026 fact: Apple announcement.

UVU’s spring 2027 schedule says faculty return January 4, classes start January 11 and the term ends May 5 fact. UVU says instructors receive a Canvas shell for assigned classes fact, but a university-wide public date for creating those shells was not found unknown. New employees may not receive access until a class is within 30 days fact: UVU faculty guidance.

Therefore:

  • September through November is a defensible course-design and recruitment window est.
  • November 30 is a proposed internal content-freeze date est, not an official UVU deadline.
  • Waiting past that window risks turning a spring teaching pilot into a technical demonstration rather than a course-integrated service est.

Sources: UVU spring dates and UVU faculty Canvas resources.

Does a faster future machine repay the wait?

For a 48-month service horizon, a future machine that is proportionally faster by g repays a wait of at most:

maximum wait = 48 × g ÷ (1 + g) fact: algebra.

Future service gainMaximum break-even wait
12% est low case5.1 months fact: arithmetic
25% est central case9.6 months fact: arithmetic
40% est high case13.7 months fact: arithmetic

A roughly 15-month wait for a rumored late-2027 Pro machine misses even the 40% case. A normal 22.4-month mini interval misses it by a wider margin.

This calculation assumes immediate full use, identical price and no resale value. Actual adoption will ramp more slowly, which makes the result less decisive. Conversely, missing a fixed spring cohort makes delay more costly. Those effects should be measured, not guessed.

Plain decision rule

Buy the four-machine validation set after September 22 if all of these conditions hold:

  • 1. One delivered machine loads the selected model and completes the 1,000-, 8,000- and 32,000-token tests est test lengths.
  • 2. It sustains at least four simultaneous service streams at the plan’s ten-token-per-second floor est acceptance rule.
  • 3. It passes batch sizes 1, 2, 4 and 8, cold-load, peak-memory and 24-hour soak tests est acceptance protocol.
  • 4. Identity, logging, network isolation, accessibility and administrator recovery tests pass.
  • 5. Apple has not officially announced, for shipment within 90 days est decision window, a machine that provides required memory or at least 35% more measured service throughput for no more than 20% additional configured cost est thresholds.

Delay only when one of these is true:

  • No current configuration passes a must-have requirement.
  • An official product inside that 90-day window crosses the test above.
  • A dated state, federal or shared-compute award within 60 days est window would cover a material part of the same need.
  • Accessibility or security testing blocks the intended service.

Do not delay merely because another chip will eventually exist.

# 4. Refresh and resale, forward

Observed retained value

Prior machineAge at observationObserved value versus original price
M2 Pro mini, 32 GB/1 TBAbout 3.6 years fact$924.95 versus $1,899; 48.7% fact: transaction and arithmetic
M2 Max StudioAbout 3.3 years fact$979–$1,249 versus $1,999; 49.0%–62.5% fact
M2 Ultra StudioAbout 3.3 years fact$2,049–$2,369 versus $3,999; 51.2%–59.2% fact
M1 Ultra StudioAbout 4.5 years factApproximately 40%–55% fact: market sample
Base M1 miniAbout 5.8 years factSwappa average sold price $402 versus $699; 57.5% fact: Sept. 2026 snapshot and arithmetic

The M1 mini result may be temporarily inflated by new-memory scarcity est. Samples also differ by memory, storage, condition, fees and sale channel. Net resale proceeds after fees, labor and warranty risk are unknown.

Conservative planning values remain:

  • Year-three mini residual: 50% est
  • Year-three Studio residual: 50% est
  • Year-three Ultra residual: 55% est
  • Year-four mini residual: 40% est

These are budget assumptions, not guaranteed sale prices.

The money analysis calculated five-year cash outlay of $90,341 for keeping the fleet, $122,438 for a year-three refresh and $128,233 for a year-four refresh est. Those cases used different terminal fleet values, so they do not prove that one refresh date has the lowest economic cost.

Recommended option-style purchase plan

An option-style plan buys enough evidence to make the next choice without locking in the whole fleet.

Wave 1 — four machines

After September 22, buy four configured M5 Pro 48 GB minis for an estimated $8,916 all-in est, subject to the delivered-unit gate.

Purpose:

  • Validate actual throughput and stability.
  • Operate a 30-faculty cohort est cohort size.
  • Establish demand, support effort and repeat-use evidence.
  • Preserve the right not to buy the remaining 22 machines.

Wave 2 — up to eight more

Release an additional order only after:

  • At least 50% of pilot participants return at their next natural task within 30 days est adoption gate.
  • Measured service demand uses at least 60% of available capacity during meaningful teaching windows for four weeks est operating gate.
  • Support load, accessibility and incident handling remain within the staffed plan.
  • No current model release materially changes the best hardware tier.

Eight additional machines would make 12 total est option size. The exact number should be set from queue and concurrency evidence, not this placeholder.

Wave 3 — up to the planned 26

Release the final order only after spring-course results, observed queue time and a refreshed local-versus-cloud cost comparison. Fourteen additional machines would reach 26 total est option size; FACT arithmetic.

Do not reserve a 26-machine quantity simply to preserve architectural symmetry.

Special triggers

  • Delivered-unit test: A failure changes the configuration or pauses the purchase; it does not lower the gate.
  • 512 GB price: Consider a 512 GB Ultra only if at least two approved research workloads need a model that cannot fit 256 GB est gate, delivered tests confirm the benefit, and the configured price is no more than the cost of approximately six all-in minis—$13,374 est threshold; FACT arithmetic.
  • Next Apple event: Reopen only unplaced orders. Do not idle machines already producing evidence.
  • Model release that changes a tier’s leader: Re-run the narrow model/machine test. A better model does not automatically require new hardware.
  • Year three: Obtain real resale quotes, compare service throughput per dollar and decide which machines merit replacement.
  • Year four: Refresh where the measured gain clears the same waiting rule. Keep useful older machines as development, test, spare or low-priority capacity rather than forcing a uniform replacement.

# 5. What could break the trajectory

The percentages below are subjective 24-month planning priors est, not measured frequencies.

RiskLikelihoodSignal to watchPlan response
US–China policy or geopolitics slows open-weight releasesMedium, 35% estNew BIS restrictions involving model weights or training; PRC export controls; repository removalKeep approved model artifacts and license records; maintain several makers; retain a metered cloud route
Licenses tightenMedium, 40% estNew commercial-use, redistribution, data-use or geographic clausesReview every version before deployment; never auto-upgrade; keep Apache/MIT fallbacks
Apple changes memory prices againHigh, 70% estTrendForce contract-price reports; configurator changes; removal of high-memory optionsBuy the four-machine option now; set price caps for later waves; do not budget the 512 GB machine before its real quote
Serving leadership shifts between MLX and llama.cppHigh, 65% estUnsupported architecture, crash or throughput regression in either stackKeep one common service interface; allow either runtime; test both for each selected model
A frontier lab releases a large permissive modelMedium-low, 30% estDownloadable model above 300B parameters with a usable commercial license est signal thresholdRe-test 256/512 GB value and cloud pricing before a Studio order
Hosted open-model prices collapseMedium, 45% estDelivered cloud cost per successful task stays at least 30% below fully loaded local cost for two quarters est gateMove variable and oversized work to the hosted pool; retain local privacy and continuity capacity

Policy and license evidence

NTIA’s July 2024 open-weight report found benefits and risks but did not find enough evidence then for immediate restrictions; it recommended continued monitoring fact. The July 2025 White House AI Action Plan supported open models fact. BIS later changed advanced-chip and training controls fact. Policy currently supports access, but it is not permanent. Sources: NTIA, AI Action Plan and BIS policy statement.

“Open-weight” does not mean “open-source.” gpt-oss uses Apache 2.0 fact; several DeepSeek and Z.ai releases use permissive licenses fact; Meta Llama and some newer models use custom terms fact. UVU needs a version-specific license record.

Serving-stack evidence

MLX supports affine two- through eight-bit quantization plus newer four-bit formats fact. But an April 2026 issue reported unsupported DeepSeek V4 architecture, and August reports described a Metal-buffer leak and an invisible server-thread failure fact: issue reports, not universal defects. llama.cpp supports Apple Metal and native formats for several families fact.

The response is not to select one permanent winner. The service boundary should allow either runtime while keeping authentication, logs and user behavior stable.

Hosted-price evidence

Together currently lists GLM-5.3 Flash at about $0.15 per million input tokens and $0.50 per million output tokens fact: Sept. 2026 listed price. Its April DeepSeek V4 Pro offer was $2.10 input and $4.40 output per million tokens fact. These prices can fall or rise.

No responsible local-versus-cloud comparison is possible until UVU knows:

  • Prompt and output volume per completed task.
  • Cache-hit rates.
  • Peak demand.
  • Local support and electricity cost.
  • The privacy class of each request.

All are unknown before the pilot.

# 6. Triggers and calendar

DateEventDecision
Sept. 9, 2026 fact event dateApple eventReopen only unplaced orders if an official Mac ships within 90 days and crosses the hardware rule
Sept. 9, 2026 fact meeting dateUtah Board of Higher Education meetingCheck agenda and later minutes for AI compute allocation and task-force action
Sept. 22, 2026 fact announced dateM6/M5 Pro mini and M5 Studio shipment waveAccept and test; do not scale until the delivered gate passes
Sept. 30, 2026 est checkpointQuarterly model reviewRe-run only tiers whose leading model changed
Oct. 1–2, 2026 fact meeting datesUSHE meetingCheck allocation, credential and task-force decisions
Oct. 31, 2026 est operating deadlineLate-October 512 GB price expectedCompare real quote and tested unique workload against six-mini threshold
Nov. 4, 2026 fact deadlineNSF AI Infrastructure Hubs deadlineSubmit only through an approved consortium and normal institutional approval
Nov. 19, 2026 fact meeting dateUSHE meetingCheck state compute and task-force milestones
Nov. 30, 2026 est internal dateProposed spring course-content freezeDecide whether spring launch is course-integrated or a smaller technical pilot
Jan. 4, 2027 factUVU faculty returnFinish faculty orientation and accessibility checks
Jan. 11, 2027 factUVU spring classes beginStart measured pilot only if launch gates pass
Jan. 20, 2027 fact deadlineNSF IUSE deadlineDecide whether measured pilot evidence supports an institutional proposal
Mar. 25–26, 2027 fact meeting datesUSHE meetingCheck state allocation and task-force progress
Mar. 31, 2027 est checkpointSpring Apple/model reviewReopen unplaced hardware options only
Apr. 26, 2027 fact federal dateFederal web/mobile accessibility deadline for covered state entities with populations at least 50,000Treat WCAG 2.1 AA compliance as a launch requirement; UVU counsel should confirm the exact legal classification unknown legal determination
May 13, 2027 fact meeting dateUSHE meetingReview recommendations, allocations and credential implementation
June 30, 2027 est checkpointMidyear model and Apple reviewCompare compact-model improvement with delivered fleet demand
July 21, 2027 fact deadlineSecond NSF IUSE deadlineApply only if evidence and university approvals support it
Sept. 30, 2027 est checkpointOne-year hardware/model reviewDecide whether Wave 3 or a changed tier is justified
Dec. 31, 2027 [RUMOR window]Reported possible M7 Pro/Max windowCompare actual product, if any, with year-one evidence
Mar. 31, June 30 and Sept. 3, 2028 est checkpointsQuarterly reviews through the 24-month horizonReview models, prices, grants and runtime support; avoid calendar-driven replacement

USHE announced a statewide AI task force on May 1, 2026 and named UVU representation fact, but no public deadline for its final recommendations was found unknown. USHE also reported a $15 million one-time state allocation for system AI compute capacity fact; UVU’s share, access rules and award timing are unknown. Sources: USHE task force, legislative update and meeting calendar.

NAIRR provides shared research-compute allocations fact, but the next allocation date and capacity available to this service are unknown. It should be watched as a complement, not used as an assumed fleet replacement. See NAIRR.

The Department of Justice’s April 2026 rule change moved the relevant federal accessibility date to April 26, 2027 for state and local entities serving at least 50,000 people fact. Public universities and vendor-supplied content are within the rule’s described scope fact; UVU should obtain counsel’s classification rather than infer it from student enrollment. See the DOJ compliance guide.

# Cross-domain learnings

Software delivery: canary releases

Site-reliability teams send a change to a small part of the system, compare it with the existing service and stop the rollout when the evidence is bad. Google’s published canary method formalizes that approach fact.

UVU equivalent: four machines are the canary. The other 22 are not a promised follow-on order.

Source: Google SRE canarying.

Clinical trials: preplanned adaptation

Adaptive clinical trials allow enrollment or allocation to change at interim reviews, but the rules are written before investigators see the results fact.

UVU equivalent: approve throughput, adoption, accessibility and price triggers now. Do not invent a favorable success rule after the pilot.

Source: FDA adaptive-design guidance.

Capital programs: stage gates

GAO’s technology guidance separates discovery, development, production and deployment decisions fact. Each gate asks for stronger proof.

UVU equivalent: delivered-unit performance, faculty-pilot behavior and campus expansion are different gates. SSH access or a loaded model is not campus readiness.

Source: GAO technology transition guidance.

Public health: tools do not create behavior by themselves

A bundled hand-hygiene program in Geneva reported compliance rising from 48% to 66% and infections falling from 16.9% to 9.9% fact: uncontrolled before/after evidence. A mandatory surgical checklist in Ontario later showed no significant population-level improvement fact.

The general lesson is not that checklists work or fail. The surrounding training, feedback, local leaders and workflow matter. UVU should fund adoption work, not count machines or activated accounts as success.

# Devil’s advocate

The strongest case against buying even four machines is this:

Hosted open models are already inexpensive, and their prices may fall faster than local hardware costs. The largest open models exceed Apple’s 512 GB ceiling. Apple memory cannot be upgraded after purchase. Current throughput is estimated, not delivered proof. Local service also creates staffing, patching, accessibility, incident-response and physical-continuity costs. USHE’s $15 million allocation, NAIRR or another shared service may cover the same need. A rumored Pro-class Apple update could arrive near the end of 2027.

On that view, UVU should run the first faculty pilot entirely through free or metered hosted tools, gather demand through spring 2027 and make no hardware purchase until state allocations and M7 products are clear.

This is a strong argument against an immediate 26-machine fleet. It is weaker against four machines if UVU’s goals include private local work, predictable service, instruction in local AI operations or protection against provider changes. The four-machine purchase is best understood as the price of learning whether those benefits are real.

# What would change the recommendation

The recommendation should change from “buy four, then stage” to “wait or use cloud” if any of these occurs:

  • The delivered M5 Pro cannot sustain four simultaneous ten-token-per-second sessions under the required context and safety settings.
  • An official machine shipping within 90 days supplies required memory or at least 35% more measured service throughput for no more than a 20% configured-price premium est rule.
  • USHE or NAIRR provides recurring, usable capacity before UVU must commit the order.
  • Hosted service remains at least 30% cheaper per successful task for two quarters after privacy, support and egress are included est rule.
  • Fewer than 40% of pilot participants return at their next natural task within 30 days est stop/rework threshold.
  • Accessibility, identity or incident-recovery requirements cannot be met without a material redesign.
  • A new license prevents the intended campus use.
  • The final 512 GB price is unusually low and at least two approved workloads demonstrate a unique need for the full GLM tier; that would add a Studio option, not justify replacing the mini validation set.

# What the plan should change

  • 1. Replace “buy after September 22” with a delivered-unit acceptance gate. Shipment is availability, not performance proof.
  • 2. Keep the four-machine purchase; remove any implied commitment to all 26. Use four, then up to eight, then up to fourteen more est wave sizes.
  • 3. Replace the fixed four-stream capacity constant. Report four as the acceptance floor and eight as an unverified upside for the current 27B model.
  • 4. Add the waiting equation. A 25% improvement repays no more than 9.6 months of waiting inside a 48-month horizon est input; FACT arithmetic.
  • 5. State the memory rule plainly. Buy for models that create useful service now; use the metered pool when a flagship exceeds local memory.
  • 6. Make adoption a purchase gate. Measure return at the next natural task within 30 days, support burden and completed work—not activation.
  • 7. Move the year-three review and year-four refresh from policy to default options. Refresh only where actual throughput, resale and workload data justify it.
  • 8. Add dual-runtime acceptance. MLX and llama.cpp should be interchangeable behind one campus-facing service contract.
  • 9. Add the dated 24-month watchlist. Include USHE meetings, grant deadlines, Apple events, quarterly model checkpoints and the April 26, 2027 accessibility deadline.
  • 10. Mark the 512 GB line as unpriced. Do not budget or recommend it until Apple posts the final configuration and UVU demonstrates a unique workload.
  • 11. Update the open-model history quarterly using one benchmark version. Preserve old snapshots separately; do not splice changing index versions.
  • 12. Add a per-successful-task local/cloud comparison. Token price alone does not include failed work, support, privacy, peak capacity or staff time.

# Source log

# NOT_RUN

  • not run — Vendor, UVU, USHE, NSF or regulator contact, as prohibited.
  • not run — Purchase, reservation, quote request or configurator checkout.
  • not run — Delivered M5 Pro/M6/M5 Studio testing; machines do not ship until September 22, 2026.
  • not run — 512 GB M5 Ultra price comparison; final price is not yet public.
  • not run — Live UVU identity, network, accessibility, security or recovery testing.
  • not run — Legal determination of UVU’s exact federal accessibility classification.
  • not run — Verification of a public Utah AI workforce RFP; none was located.
  • not run — Exhaustive machine-readable reconstruction of every historical open-model leaderboard. The tier winners are documented estimates using the stated fit rule.
  • not run — Economic-value calculation for missed usage; UVU demand, task value, support cost and successful-task volume are unknown.

The machine-readable data behind this appendix — every row, with its own grade — is not printed here: it is a page of raw structured text, not something a reader can check by eye. Everything it contains that bears on a decision is stated above in words. The file itself is findings/47-roadmap-timing.md in the research pack.

·

Method note

Where these tables come from

Every appendix row traces to a dated public source or a measured result in the Utlyze working papers (forty-eight research lanes, an adversarial verification pass, a three-persona red team, and a live benchmark of the exact serving stack on our own 512GB-class hardware — method in the guide, §13). Claims are graded fact where a cited source states them, est where we computed or judged, unknown where the public record is silent. Nothing in these appendices required contacting UVU — by design, so this document arrives owing nothing.

Continue: the guide · the budget simulator · the campus map. This page prints clean — File → Print.