How To Guide Archives - Thu, 04 Jun 2026 12:09:44 +0000 en-GB hourly 1 https://wordpress.org/?v=7.0.2 /wp-content/uploads/2026/02/cropped-TruStrata-Favicon-Web-App-sparkle-32x32.png How To Guide Archives - 32 32 Codeframe Classification https://trustrata.com/docs/codeframe-classification/ Thu, 02 Apr 2026 12:32:25 +0000 https://trustrata.com/?post_type=docs&p=1363 Getting Started How to classify verbatim responses against your own predefined codeframe using TruVerbatim. Codeframe Classification takes your existing coding […]

The post Codeframe Classification appeared first on .

]]>
Getting Started

How to classify verbatim responses against your own predefined codeframe using TruVerbatim.

Codeframe Classification takes your existing coding scheme – a structured list of themes with descriptions – and uses AI to classify every verbatim response against it. Instead of human coders reading each response and assigning a code, TruVerbatim does it for you, consistently and at scale.

Each response receives:

  • A primary code matching the most relevant theme in your codeframe
  • Secondary codes for any additional themes mentioned in the same response
  • Automatic quality filtering to remove low-confidence assignments

When to Use Codeframe Classification

Use it when:

  • You have an established codeframe from previous waves of research
  • You need results to be comparable across time periods or markets
  • Your client has specified exact themes they want coded against
  • You are running a tracking study that requires consistent categories

Use Topic Discovery instead when:

  • You are exploring the data for the first time
  • You do not have a predefined list of themes
  • You want to see what themes emerge naturally from the data

Before You Start

Preparing Your Codeframe

Your codeframe file should be a CSV or Excel file with at least two columns:

ColumnRequiredDescription
Theme nameYesThe name of each code/theme (e.g. “Customer Service”)
DescriptionRecommendedA description of what belongs in this theme

Example codeframe:

theme_namedescription
Customer ServiceMentions of interactions with support staff, call centres, help desks, or any customer-facing service experience
Product QualityComments about build quality, durability, reliability, or defects. NOT mentions of product features or design
Pricing & ValueReferences to cost, value for money, affordability, price comparisons, or discounts
Delivery & ShippingFeedback about delivery times, shipping costs, packaging condition, or courier experience
Website & AppMentions of the online shopping experience, app usability, checkout process, or website navigation

Adding descriptions significantly improves accuracy. The more detail you provide about what each theme covers (and what it does not cover), the better the AI can match responses.

Preparing Your Verbatim Data

RequirementDetail
File formatCSV or Excel (.xlsx)
Text columnOne column containing the verbatim responses
MetadataOptional extra columns (age, region, gender) enable cross-tabulation later

Step-by-Step Guide

Step 1: Select Codeframe Classification

When starting a new analysis, select Codeframe Classification as your analysis type. This tells TruVerbatim that you will be providing your own coding scheme rather than discovering themes from the data.

Step 2: Upload Your Verbatim Data

  1. Drag and drop your CSV or Excel file onto the upload area
  2. Select the column containing your verbatim text

Optional: Enable auto-cleaning to remove personal information, profanity, duplicates, and blank rows.

The system will prepare your data, applying any cleaning steps you selected. Wait for the “data ready” confirmation before proceeding.

Step 3: Upload Your Codeframe

Once your verbatim data is cleaned and ready, a codeframe upload prompt appears automatically in the chat. You will see an upload area specifically for your codeframe file.

  1. Drag and drop your codeframe file (CSV, Excel, or Parquet format)
  2. The system detects the columns in your file automatically
  3. Select the theme column – which column contains your theme/code names
  4. Select the description column – which column contains the theme descriptions
Codeframe classification with TruVerbatim

If your codeframe has only one column (theme names without descriptions), select the same column for both. Classification will still work, but descriptions improve accuracy.

Step 4: Start Classification

Click the Classify with Codeframe button. The classification begins automatically.

Real-time progress updates appear in the chat:

  1. Loading codeframe – your coding scheme is parsed and validated
  2. Preparing data – responses are organised for efficient processing
  3. Classifying – the AI processes your responses, with a progress percentage updating as it goes
  4. Validating – low-confidence assignments are automatically filtered out
  5. Correcting – any blank or invalid assignments are retried automatically

Step 5: View Your Results

When classification completes, the chat displays:

  1. Interactive bar chart – your codeframe themes ranked by the percentage of responses assigned to each
  2. Statistics summary – total responses classified, multi-label count, unique themes used
  3. Download button – click to download the full classified CSV

Chart interactions:

  • Hover over any bar to see exact counts and percentages
  • Export the chart as PNG, SVG, or PDF from the chart menu

Mention rank filtering (if applicable): If your verbatim data contained grouped columns that were unpivoted, toggle chips appear above the chart allowing you to filter by mention order (Total, 1st mention, 2nd mention, etc.).

Step 6: Download Your Results

Click the Download CSV button. The exported file includes:

ColumnDescription
Original verbatim textThe response as uploaded
themesAssigned theme(s) – comma-separated if multi-label
*Original columns*All metadata from your uploaded file

For single-label responses, the themes column contains one theme name. For multi-label responses, themes are comma-separated (e.g. “Customer Service, Product Quality”).

How It Works

Automatic “Uncodeable” Handling

The system automatically handles responses that do not fit any of your themes. Blank responses, gibberish, or text completely outside the scope of your coding scheme are marked as “Uncodeable/Ambiguous” rather than being forced into an inappropriate theme.

Multi-Label Classification

The AI assigns multiple themes where appropriate. Most responses that discuss more than one topic will receive 2-3 theme assignments. The AI only assigns a single theme when the response truly focuses on one specific topic.

Quality Filtering

After the initial classification, the system validates each assignment to check that it is a good fit. Low-confidence classifications are automatically removed, improving overall accuracy. If all assignments for a response are removed during filtering, it is marked as “Uncodeable/Ambiguous”.

Automatic Correction

Any responses left blank or assigned to a theme not in your codeframe are automatically retried. This typically recovers the vast majority of initially failed classifications, ensuring high coverage across your dataset.

After the Classification

Ask Questions

Type questions in the chat to explore your results:

  • “Show me the top 5 themes by count”
  • “What percentage were coded as Customer Service?”
  • “Show me verbatims coded to Product Quality”
  • “Show me a crosstab of themes by region”
  • “Which themes have the most multi-label assignments?”

Generate PowerPoint

Click “Generate PowerPoint” to create a presentation deck with your classification chart and AI-written insight subtitles.

Export Charts

Right-click any chart to export as PNG, SVG, PDF, or download the underlying data as CSV or Excel.

Writing Better Codeframes

The quality of your classification depends heavily on your codeframe. Here are guidelines for getting the best results:

Be Specific in Descriptions

The description tells the AI exactly what belongs in each theme.

Quality Example
Vague“Service”
Better“Customer Service – Mentions of interactions with support staff, call centres, help desks, or any customer-facing service experience”
Best“Customer Service – Mentions of interactions with support staff, call centres, help desks, or customer-facing service. Includes complaints about wait times, praise for helpful staff, and references to support channels. Does NOT include product complaints or delivery issues”

Include Exclusions

Telling the AI what does NOT belong is just as valuable as telling it what does:

  • “Product Quality – Mentions of build quality, durability, reliability, or defects. NOT mentions of product features or design”
  • “Pricing – References to cost and value. NOT mentions of promotions or marketing”

Keep Themes Distinct

Overlapping themes confuse the AI just as they would confuse a human coder. If two themes are hard to distinguish, consider:

  • Merging them into one broader theme
  • Adding explicit boundary descriptions to clarify where one ends and the other begins

Limit Codeframe Size

While TruVerbatim handles large codeframes, classification accuracy tends to be highest with 10-30 themes. If your codeframe has 50+ themes, consider whether some can be grouped into parent categories.

The post Codeframe Classification appeared first on .

]]>
Extraction Pipeline User Guide https://trustrata.com/docs/extraction-pipeline-user-guide/ Tue, 24 Mar 2026 14:40:22 +0000 https://trustrata.com/?post_type=docs&p=1278 How to extract and normalise key terms from short verbatim responses using TruVerbatim’s Entity Extraction pipeline. Overview The Extraction Pipeline […]

The post Extraction Pipeline User Guide appeared first on .

]]>
How to extract and normalise key terms from short verbatim responses using TruVerbatim’s Entity Extraction pipeline.

Overview

The Extraction Pipeline is designed for responses that mention specific things – brands, products, features, places, or other named entities. Instead of clustering responses by meaning, the pipeline reads each response and pulls out every entity mentioned, corrects spelling variations, and builds a clean frequency-ranked codeframe.

Each response is assigned:

  • One or more extracted entities (e.g. “Nike”, “Adidas”)
  • A raw mention preserving the original text before normalisation

Step-by-Step Guide

Step 1: Upload Your Data

  1. Open TruVerbatim and sign in
  2. Drag and drop your CSV or Excel file onto the upload area in the chat
  3. Select the column containing your verbatim responses

Optional: Enable auto-cleaning to remove personal information, profanity, duplicates, and blank rows.

Step 2: Review the Recommendation

TruVerbatim analyses your data and presents a triage recommendation. If your responses are short or sparse, the system will recommend Key Term Extraction as the best approach.

You will see:

  • A recommendation card with a confidence score
  • A brief explanation of why extraction was recommended
  • Key data metrics (median response length, short response rate, vocabulary diversity)

If the recommendation shows “Thematic Analysis” but you know your data contains entity mentions, you can override and select “Key Term Extraction” instead.

Step 3: Select Key Term Extraction

Click the Key Term Extraction pipeline card. The analysis begins immediately.

Step 4: Watch the Analysis Run

Real-time progress updates appear in the chat:

  1. Starting entity extraction pipeline – data is loaded and validated
  2. Understanding your data – the system samples your responses to understand the domain (e.g. brands, products, places)
  3. Extracting entities – the AI processes your responses in batches, with a progress percentage updating as it goes
  4. Building codeframe – extracted entities are counted and ranked by frequency
  5. Generating insight – the AI writes a brief summary of the findings

A progress bar shows the percentage complete throughout.

Step 5: View Your Results

When the extraction completes, the chat displays:

  1. Entity frequency chart – a bar chart showing your extracted entities ranked by how often they were mentioned
  2. AI-generated insight – a narrative summary of the top entities and their distribution
  3. Download button – click to download the full classified CSV

Step 6: Download Your Results

Click the Download CSV button. The exported file includes:

ColumnDescription
verbatim_idRow identifier
verbatim_textThe original response
EXTRACTED_ENTITYAll normalised entities found (semicolon-separated)
RAW_MENTIONThe original text before normalisation
THEMEPrimary entity
Original columnsAll metadata from your uploaded file

How It Works

Domain Understanding

Before processing your full dataset, the pipeline samples a selection of your responses to understand the domain. This helps the AI recognise whether responses are about brands, products, cities, features, or something else entirely – so it extracts the right type of entity.

Entity Extraction

For each group of responses, the AI:

  1. Uses the domain context to understand what types of entities to expect
  2. Reads each response carefully
  3. Extracts every named entity mentioned
  4. Splits multi-entity responses (e.g. “Nike and Adidas” becomes two separate entities)
  5. Corrects typos and spelling variations while preserving meaning
  6. Assigns a confidence level (high, medium, or low)

Normalisation

The AI automatically normalises variations:

  • “Nike”, “nike”, “NIKE”, “Nikee” all become “Nike”
  • “Customer Service”, “customer service”, “cust. service” are unified
  • The original mention is preserved in the RAW_MENTION column so you can always see what the respondent actually typed

Codeframe Building

Once extraction is complete, the system:

  1. Counts how many times each unique entity appears
  2. Ranks entities by frequency (most mentioned first)
  3. Calculates percentages of total responses
  4. Builds a codeframe compatible with the rest of TruVerbatim’s tools (Q&A, sentiment, PowerPoint)

After the Extraction

Ask Questions

Type questions in the chat to explore your results:

  • “What are the top 10 entities?”
  • “How many people mentioned Nike?”
  • “Show me a crosstab of entities by age group”
  • “Show me the verbatims that mentioned Adidas”

Handling Multi-Entity Responses

The extraction pipeline handles responses that mention multiple entities. For example:

ResponseExtracted entities
“Nike and Adidas”Nike; Adidas
“I bought shoes from Nike/Reebok”Nike; Reebok

Each entity is counted separately in the frequency chart. The CSV shows all entities in the EXTRACTED_ENTITY column (semicolon-separated) and the primary entity in the THEME column.

Troubleshooting

IssueLikely causeSolution
Entities not being splitUnusual separator in responsesThe AI handles “/”, “&”, “and”, and commas – other separators may need pre-processing
Too many variations of the same entityUnusual spellings or abbreviationsThe AI normalises most variations, but you can merge in post-processing
“Uncodeable” responsesBlank, gibberish, or truly non-extractable textThese are expected – review them to check they are genuinely non-responses
Unexpected entities extractedThe AI misinterpreted the domainThis can happen with ambiguous data. Try re-running the analysis
Analysis seems slowLarge dataset (5,000+ responses)Expected – the pipeline processes your data in stages. Progress updates show the current stage

Tips for Best Results

  • Short, specific responses work best – the pipeline is optimised for brand names, product mentions, and single concepts
  • Include context in the question – if your survey question mentions “brands”, the AI will focus on brand extraction
  • Review the “Uncodeable” items – they usually indicate blank or gibberish responses, but occasionally contain a valid entity the AI missed
  • Use mention rank filtering – if your data has grouped columns (brand_1, brand_2, brand_3), the chart will offer rank-based filtering to see first-choice vs second-choice mentions
  • Combine with Q&A – after extraction, use natural language questions to cross-tabulate entities against demographic variables

The post Extraction Pipeline User Guide appeared first on .

]]>
Theme Extraction https://trustrata.com/docs/theme-extraction/ Tue, 24 Mar 2026 14:33:34 +0000 https://trustrata.com/?post_type=docs&p=1277 How to discover themes in your verbatim data using TruVerbatim’s AI-powered topic modelling pipeline. Topic Modelling reads through all your […]

The post Theme Extraction appeared first on .

]]>
How to discover themes in your verbatim data using TruVerbatim’s AI-powered topic modelling pipeline.

Topic Modelling reads through all your open-ended survey responses and automatically groups them into meaningful themes. The system analyses the richness of your text and routes it through the most appropriate pipeline, so you always get the best results for your data.

Each response is assigned:

  • A primary theme (e.g. “Customer Service”)
  • A sub-theme (e.g. “Staff Friendliness”)
  • Secondary codes for additional themes mentioned in the same response

Before you start

RequirementDetail
File formatCSV or Excel (.xlsx)
Minimum rows50 responses (100+ recommended for best results)
Text columnOne column containing the verbatim responses
MetadataOptional extra columns (age, region, gender) enable cross-tabulation later

Data Preparation Tips

  • Each row should contain a single response
  • Use clean, simple column headers without special characters
  • Remove fully blank rows before uploading, or let TruVerbatim’s auto-cleaning handle them
  • If your data has grouped columns (e.g. brand_1, brand_2, brand_3), TruVerbatim will detect and unpivot them automatically

Step-by-Step Guide

Step 1: Upload Your Data
  1. Open TruVerbatim and sign in with your account
  2. In the chat interface, drag and drop your CSV or Excel file onto the upload area (or click to browse)
  3. Select the column containing your verbatim text (e.g. “Q5_Response”, “Open_Ended_Feedback”)

Auto data cleaning to:

  • Detect and anonymise personal information (names, emails, phone numbers)
  • Filter profanity
  • Remove duplicate responses
  • Remove blank rows

Step 2: Review the Triage Recommendation

Before running any analysis, TruVerbatim performs a richness check on your data. This evaluates several characteristics of your text:

What it checksWhat it means
Response lengthAre responses long enough for thematic clustering?
Short response rateWhat proportion of responses are very brief?
Vocabulary diversityIs the language varied enough to find distinct themes?
Text densityDo responses share enough common vocabulary?

Based on these checks, you will see a recommendation card with a clear explanation and a suggestion. Two pipeline options are presented:

  • Thematic Analysis – recommended for rich, detailed responses (paragraphs, full sentences)
  • Key Term Extraction – recommended for short responses mentioning specific entities (brands, products, places)

You are free to accept the recommendation or choose a different approach to get the best insights from your data.

Step 3: Start the Analysis

Click your chosen pipeline card. The analysis begins immediately and you will see real-time progress in the chat:

  1. Analysing language patterns (thematic analysis only) – each response is analysed for meaning and similarity
  2. Grouping responses – responses with similar meanings are grouped together into clusters
  3. Naming themes – the AI reads each cluster and gives it a descriptive, human-readable name
  4. Building hierarchy – clusters are organised into parent themes and sub-themes
  5. Multi-coding – every response is classified against the full codeframe, including secondary codes
  6. Quality check – the system evaluates how well the themes separate your data

A progress bar and status messages update in real time so you always know what stage the system has reached.

Step 4: View Your Results

Step 4: View Your Results

When the analysis completes, several elements appear in the chat:

  1. Interactive bar chart – your themes ranked by frequency with drill-down to sub-themes
  2. Theme summary table – a table listing each theme with its count, percentage, and top keywords
  3. AI-generated insight – a brief narrative summary of the key findings

Step 5: Interact with the Chart
  • Click any bar to drill down into its sub-themes
  • Click the breadcrumb at the top to navigate back to the parent level
  • Hover over a bar to see exact counts and percentages in the tooltip

Mention rank filtering (grouped data only): If your data contained grouped columns that were unpivoted, toggle chips appear above the chart:

  • Total – all mentions combined
  • 1st mention – first-choice responses only
  • 2nd mention, 3rd mention, etc.

This lets you see whether a theme is top-of-mind or typically a secondary consideration.

Tips for Best Results

  • More data is better – aim for at least 100 responses for thematic analysis
  • Trust the recommendation – the richness check is designed to route your data to the best pipeline
  • Review the outliers – always check the “Other” category for responses that did not fit a theme
  • Iterate – use merge and Q&A to refine the theme structure until it tells the story you need
  • Include metadata – additional columns like age, region, or segment enable much richer cross-tabulation in the Q&A step

The post Theme Extraction appeared first on .

]]>
When to use each pipeline https://trustrata.com/docs/when-to-use-each-pipeline/ Mon, 16 Mar 2026 13:31:25 +0000 https://trustrata.com/?post_type=docs&p=1054 The tables below provide an overview of each analysis pipeline and offer guidance on when and how to use them […]

The post When to use each pipeline appeared first on .

]]>
The tables below provide an overview of each analysis pipeline and offer guidance on when and how to use them effectively.

Overview 

Thematic AnalysisTranscript AnalysisKey Term ExtractionCodeframe Classification
PurposeDiscover the themes and topics hidden in open-ended responsesAnalyse long-form content such as interview transcripts, focus-group discussions, and multi-paragraph feedbackExtract and count the specific items (brands, products, features) that respondents mentionClassify responses against a predefined set of codes/categories provided by the user
Best forLong, descriptive responses – sentences and paragraphsVery long responses – multi-sentence and multi-paragraph (interviews, transcripts, focus groups)Short responses – single words, brand names, product mentionsAny response length – when you already know the categories you want to code against
What it producesA two-level hierarchy of parent themes and sub-themesA two-level hierarchy of parent themes and sub-themesA frequency-ranked list of normalised entitiesA flat frequency distribution of responses across your predefined codes
How it classifiesGroups responses by meaning – responses about similar topics end up in the same themeSplits each response into overlapping sentence windows, clusters the chunks by meaning, then aggregates back to document levelReads each response and pulls out every named item mentionedMatches each response to the most relevant code(s) from your uploaded codeframe
Multi-label supportYes – primary theme plus secondary codesYes – primary theme plus secondary codesYes – multiple entities per responseYes – multi-label is the default; most responses get multiple codes
Who defines the categories?The system discovers them from the dataThe system discovers them from the data (same as Thematic but optimised for long text)The data itself – entities are extracted as-isYou do – you upload a codebook of themes and descriptions
HierarchyYes – parent themes and sub-themesYes – parent themes and sub-themesFlat listFlat list (no sub-themes)

Scenario Comparisons

ScenarioThematic AnalysisTranscript AnalysisKey Term ExtractionCodeframe Classification
Responses are sentences or paragraphsBestBest Good
Responses are single words or short phrases BestGood
You want to understand the topics people are talking aboutBestNot designed for thisPartial – only finds topics in your codeframe
You want to count how many times each brand/product was mentionedNot designed for thisNot designed for thisBestPossible if your codeframe lists the brands
You want a parent/sub-theme hierarchyYesYes No – flat structure only
You already have a codebook and need responses coded against itNot designed for thisNot designed for thisNot designed for thisBest
You need consistent, repeatable coding categories across wavesThemes may vary slightly between runsThemes may vary slightly between runsGood – entities come from the dataBest – categories are locked to your codeframe
Grouped columns (e.g. brand_1, brand_2, brand_3)Yes – with mention rank filteringYes – with mention rank filteringYes – with mention rank filteringYes – with mention rank filtering
Fewer than 50 responsesNot enough data to find reliable patternsNot enough data to find reliable patternsWorks well – even small datasets produce useful frequency countsWorks well – predefined codes don’t need large samples to apply
You need to compare results across waves or marketsGood – but themes may vary slightly between runsGood – but themes may vary slightly between runsGood – entities are consistent because they come from the data itselfBest – same codeframe guarantees identical categories every time

The post When to use each pipeline appeared first on .

]]>
Getting Started https://trustrata.com/docs/getting-started/ Mon, 16 Mar 2026 12:41:03 +0000 https://trustrata.com/?post_type=docs&p=1051 TruVerbatim is an AI-powered tool designed to analyse and code open-ended survey responses. This guide walks you through everything you […]

The post Getting Started appeared first on .

]]>
TruVerbatim is an AI-powered tool designed to analyse and code open-ended survey responses. This guide walks you through everything you need to know to get the most out of the application.

Why Use TruVerbatim

  • Flexible Topic Discovery: Automatically detect themes using a hybrid blend of deep learning, machine learning, and large language models, or upload and apply your own existing codeframe.
  • Granular Insight: Explore themes and sub-themes at multiple levels of detail using dynamic, interactive visualisations.

Rich Exports: Generate PowerPoint reports, Excel exports, and structured codeframes ready for downstream analysis.

  • Powerful Analysis Layer: Crosstab topic percentages against profiling variables, or search for verbatim examples matching specific criteria.
  • Conversational Workflow: A fluid, agentic interface that allows you to quickly build a clear narrative from verbatim data.

System Requirements

  • Browser: Google Chrome, Microsoft Edge, Firefox, or Safari (latest versions recommended)
  • Internet Connection: Required for all features, as data processing takes place in the cloud.
  • File Format: Verbatims: CSV files (comma-separated values), Codeframes: CSV or Excel files

Accessing the Application

  1. Open your web browser. Navigate to: https://truverbatim.com
  2. The login screen displaying the TruVerbatim logo will appear.Access credentials will be provided by your system administrator.

Logging In

  1. You can login in using a Microsoft or Google account following the authentication steps on the screen
  2. Click the Login button.

Understanding the Interface

After logging in, you will see a chat-style interface, similar to modern messaging applications.

The main conversation area where the system displays:

  • System responses and results
  • Interactive charts and visualisations
  • File upload confirmations
  • Analysis progress updates
  • Q&A responses
  • Export options

Text Input Box Located at the bottom of the screen. You can use this to:

  • Type messages and questions
  • Send commands to the system

File Upload:

  • Drag and drop files into the upload area
  • Click the upload button to add a CSV or Excel file

Starter Buttons: Quick-start options displayed at the top of the interface:

  • Discover Themes: Start automatic topic discovery from your dataset.
  • Classify with Codeframe: Apply a pre-defined codeframe to classify responses.
Discover themes or classify with codeframe with TruVerbatim

The post Getting Started appeared first on .

]]>