Back to Selected Work

CASE STUDY · ENTERPRISE WORK

AI Innovation Lab Platform

Developer Platform for LLM Testing & Prompt Evaluation

Redesigning an internal AI experimentation platform used by developers, prompt engineers, and product teams to test LLMs, compare outputs, evaluate document analysis, and scale reusable prompt workflows.

Developer PlatformAI WorkflowsLLM TestingPrompt EvaluationInternal ToolsOutput ComparisonRun HistorySelf-Service Platform
SuccessNeeds reviewFailed
High
Medium
Low
v1v2v3 ← current
AI Innovation Lab — Workspace
AI LAB
Workspace
Models
Prompts
Documents
Run Test
Compare Outputs
Logs
History
Templates
RECENT RUNS
Run #247Success
Run #246Needs review
Run #244Failed
Model
LLM-Alpha-v2
Prompt template
Document Extraction v3v3
Saved versions:v1v2v3
PROMPT EDITOR
1system:
2 You are a document analysis assistant.
3 Extract key fields and validate against schema.
5instructions:
6 1. Parse the uploaded document
7 2. Identify entity type and classification
8 3. Extract required fields with confidence
9 4. Flag anomalies and missing data
11output_format: structured_json
12confidence_threshold: 0.82
13
UPLOAD DOCUMENT
Drop file here or
browse to upload
PDF, DOCX, TXT
sample-document.pdf
124 KB · Uploaded
CONFIDENCE SCORE
88%
High
Success
OUTPUT PREVIEW
entity_type:"contract"
classification:"high_priority"
confidence:0.88
fields_extracted:12 / 12
anomalies:[]
validation:"pass"
LOGS · RUN #247
09:14:03INFOModel LLM-Alpha-v2 initialized
09:14:03INFOPrompt template v3 loaded · 247 tokens
09:14:04INFODocument parsed · 12 fields identified
09:14:05OKExtraction complete · confidence 0.88 · validation PASS

ROLE

Principal UX Designer / Design Lead

USERS

Developers, AI prompt engineers, product owners, and business stakeholders

PLATFORM

Enterprise web platform

FOCUS

LLM testing, prompt workflows, document analysis, output evaluation, self-service tooling

TOOLS

Figma, prototypes, workflow maps, stakeholder reviews, engineering collaboration

STATUS

Redesigned / productized self-service platform direction

CONFIDENTIALITY NOTE

Some enterprise work is recreated, abstracted, or generalized to protect confidential information. Case studies focus on the design challenge, workflow complexity, decision-making, and product outcomes rather than proprietary implementation details.

PLATFORM DESIGN CHALLENGE

This was not just an AI demo tool. It was a platform design challenge: how to make LLM experimentation, prompt iteration, document analysis, output evaluation, and reusable workflows understandable across developers, prompt engineers, product owners, and business stakeholders.

01
SITUATION

Situation

Business context

The AI Innovation Lab began as an internal development tool for testing LLMs and exploring Intelligent Document Processing use cases. As more teams and use cases were added, the tool became harder to navigate, harder to evaluate consistently, and less accessible for non-developer stakeholders.

User context

Developers, prompt engineers, product owners, and business stakeholders needed a clearer way to test models, refine prompts, upload documents, compare outputs, review confidence, and understand whether results met expectations.

System context

The platform needed to support technical experimentation while evolving toward a more productized self-service experience. The design had to make complex AI workflows understandable without hiding the technical details users needed to evaluate results.

02
DESIGN CHALLENGE

Design Challenge

How might we redesign an internal AI experimentation platform so developers, prompt engineers, product teams, and stakeholders can test LLMs, evaluate outputs, and reuse prompt workflows with less friction and more confidence?

The tool had grown organically

As more use cases were added, the experience became convoluted and harder to navigate — making it difficult to understand where to start and what to do next.

Prompt testing lacked clear workflow structure

Users needed a more organized way to select models, manage prompt versions, run tests, review results, and compare outputs — without losing track of what they had already tried.

Evaluation needed to be easier to trust

Users needed visibility into confidence scores, pass/fail validation, logs, run history, and document analysis results to evaluate AI behavior with real confidence.

Developer tool complexityMultiple user typesLLM testingPrompt versioningDocument uploadOutput comparisonConfidence scoresLogs and run historyPass/fail validationSelf-service productizationEngineering constraintsReusable prompt templates
03
MY ROLE

My Role

What I led

Led the UX redesign of the AI Innovation Lab platform — framing the work around reducing complexity and supporting repeatable AI experimentation. Partnered closely with development teams, facilitated stakeholder reviews with product, engineering, prompt engineering, and business partners, and helped shift the platform toward a clearer self-service experience.

What I designed

Platform information architecture, model selection workflows, prompt versioning patterns, test run flows, document upload and analysis experience, side-by-side output comparison, confidence score and pass/fail review patterns, logs, run history, and reusable prompt template concepts for self-service workflows.

Who I partnered with

Developers, AI prompt engineers, product owners, business stakeholders, engineering leads, Intelligent Document Processing teams, and platform and AI partners across the organization.

04
MAKING THE WORKFLOW VISIBLE

Making the Workflow Visible

Before redesigning the interface, I mapped the end-to-end AI experimentation workflow: how users selected a model, created or reused prompts, uploaded documents, ran tests, reviewed outputs, compared results, and decided whether a prompt or model response was successful.

DEVELOPER WORKFLOW MAP

End-to-end AI experimentation workflow — from model selection to saved prompt version.

01Select model
02Choose template
03Edit prompt
04Upload document
05Run test
06Compare outputs
07Review confidence
08Validate pass/fail
09Inspect logs
10Save version
TECHNICAL DECISION POINT
Steps 01, 02, 03Model, prompt, or config choice
EVALUATION POINT
Steps 06, 07, 08Compare outputs, decide pass/fail
TROUBLESHOOTING SIGNAL
Steps 07, 09Logs, confidence, failed runs
SELF-SERVICE OPPORTUNITY
Steps 02, 10Template reuse, version sharing
05
EXPLORING THE EXPERIENCE

Exploring the Experience

AI Lab PlatformModelsPromptsDocumentsRuns

Platform IA

Organizing model selection, prompt templates, test runs, document upload, output review, and history into a clearer navigational structure.

v3CurrentACTIVE
v2Previousarchived
v1Initialarchived

Prompt Versioning

Exploring how users could create, edit, save, compare, and reuse prompt versions across test runs to support iteration without losing prior work.

UploadExtractReview

Document Analysis Flow

Designing how users upload documents, run analysis, and review AI-generated or extracted outputs against expectations.

Model A · v2
High
Model B · v3
Medium

Side-by-Side Output Comparison

Exploring how to compare responses across models, prompt versions, or test runs in a single view to support faster evaluation.

09:14:05OKExtraction complete · PASS
09:12:18WARNConfidence below threshold
09:11:44ERRTimeout on document parse
09:09:30OKRun complete · PASS

Logs and Run History

Designing visibility into past runs, system behavior, and troubleshooting details to support diagnosis and confidence in results.

01
Discover template
02
Configure prompt
03
Test run
04
Publish
Doc extraction v3Summary v2Classification v1

Self-Service Productization

Exploring how the platform could support product owners and business stakeholders without losing the technical depth that developers and prompt engineers needed.

06
THE SOLUTION

The Solution

The redesigned platform clarified the AI experimentation workflow and made it easier for users to move from setup to testing to evaluation. The experience supported technical users while making the platform more accessible as a self-service tool for broader product and business teams.

ABSTRACTED PLATFORM VISUALS · SCREENS RECREATED FOR PUBLIC PORTFOLIO
AI Innovation Lab — Redesigned Platform
SuccessRun #247 complete
AI LAB
Models
Prompts
Documents
Run Test
Compare
Logs & History
Templates
Model
LLM-Alpha-v2
Template
Document Extractionv3
v1v2v3
PROMPT EDITOR
system: You are a document analysis assistant.
task: Extract all key fields and validate against schema.
confidence_threshold: 0.82
output_format: structured_json
flags: anomaly_detection, field_validation
DOCUMENT UPLOAD
sample-document.pdf
124 KB
+ Add
OUTPUT COMPARISON
SuccessNeeds review
Model A · v2
Model B · v3
CONFIDENCE88%
LOGS · RUN #247
09:14:05OKExtraction complete · PASS
09:14:04INFODocument parsed · 12 fields
09:14:03INFOPrompt v3 loaded · 247 tokens
PLATFORM SIGNALSSuccessNeeds reviewFailed
High
Medium
Low
Template: Doc extraction v3Template: Summary v2Template: Classification v1

FEATURE AREA

Model and Prompt Setup

A clearer setup flow helped users select a model, choose a prompt template, edit prompt content, and prepare a test run — reducing setup friction and ambiguity.

FEATURE AREA

Prompt Versioning

Versioning patterns helped users track prompt changes, compare variations, and reuse successful prompt structures across test runs.

FEATURE AREA

Document Upload and Analysis

A structured document analysis flow helped users upload test documents and evaluate AI-generated or extracted outputs in a consistent, comparable format.

FEATURE AREA

Output Comparison

Side-by-side result comparison helped users evaluate differences across prompts, models, and runs — making quality judgments faster and more reliable.

FEATURE AREA

Confidence and Pass/Fail Review

Review patterns helped users understand confidence scores, validate results, and determine whether outputs met expectations with a clear pass/fail signal.

FEATURE AREA

Logs, Run History, and Templates

Logs and run history supported troubleshooting and diagnosis, while reusable prompt templates helped scale repeatable workflows across teams.

07
KEY DESIGN DECISIONS

Key Design Decisions

Design the workflow around experimentation

WHYUsers needed to iterate across prompts, models, documents, and outputs — not complete a single linear task. The tool had to support non-linear exploration.
TRADEOFFA simpler form-based interface would be easier to build but would not support the real experimentation workflows users were performing.
RESULTThe design made testing, comparison, and iteration central to the platform experience — matching how developers and prompt engineers actually work.

Make outputs easier to compare

WHYUsers needed to judge whether one model, prompt version, or test run performed better than another — often across multiple variations at once.
TRADEOFFShowing one result at a time would reduce screen complexity, but it would significantly increase cognitive load during evaluation.
RESULTSide-by-side comparison helped users evaluate output quality more efficiently and make clearer decisions about which direction to pursue.

Expose confidence, logs, and run history

WHYTechnical users needed visibility into what happened, what changed, and why a result might have failed — not just the final output.
TRADEOFFHiding system details would make the UI cleaner for non-technical users, but would reduce trust and troubleshooting value for the primary users.
RESULTThe design supported visibility, diagnosis, and more confident evaluation — building trust in the platform over time.

Productize without oversimplifying

WHYThe platform needed to serve a wider audience — including product owners and business stakeholders — without removing the technical depth that developers and prompt engineers relied on.
TRADEOFFA heavily simplified experience could help non-technical users but would frustrate the expert users who depended on the platform most.
RESULTThe design created a clearer self-service workflow while preserving the technical depth and system visibility needed for AI experimentation.
08
IMPACT

Impact

USER IMPACT

Reduced friction for developers, prompt engineers, product owners, and stakeholders testing AI document workflows — making the platform usable by a broader team without losing technical depth.

PRODUCT IMPACT

Helped evolve the AI Innovation Lab from a convoluted development tool toward a more productized self-service platform with a clearer experimentation workflow.

WORKFLOW IMPACT

Clarified model selection, prompt versioning, test runs, document analysis, output comparison, and evaluation workflows — reducing the steps between setup and insight.

PLATFORM IMPACT

Created reusable UX patterns for AI experimentation, prompt templates, run history, logs, and result validation that can scale across future AI product work.

09
REFLECTION

Reflection

What I learned

AI experimentation tools need to support iteration, comparison, and troubleshooting as first-class parts of the experience — not afterthoughts layered on top of a basic form interface.

What I would improve

I would continue refining how logs, confidence scores, and evaluation history help users diagnose issues without overwhelming them with raw technical output — the signal-to-noise balance is hard to get right.

How this shaped my design approach

This work reinforced that developer and AI platforms need to expose complexity in structured, legible ways instead of hiding it entirely. Users don't need less information — they need better-organized information.