Open Weight Models in Market Research: The License Decides
Open weight models in market research keep confidential transcripts on your own servers. The measured gap against closed models, and what the license allows.
A pharmaceutical client sends forty depth interviews about a product still under regulatory review. The contract says the recordings and the transcripts stay on the institute's own servers, and the client's legal team attaches a list of countries where the material may not be processed. The analysis still has to happen, on the same deadline. That clause is why open weight models in market research moved from a hobby topic to a question that comes up in procurement.
Key Takeaways
- An open weight model is one whose trained numbers are published for download, so it can run on hardware you control.
- Artificial Analysis and Epoch AI put the distance between the best open model and the best closed one at roughly 2 to 3.4 points on general indices, with a lag of about four months.
- That distance grows on harder tests. On Terminal Bench 2.1 two open models scored 88.3 and 88.2 against 88.8 for the closed leader. On version 3.0 the same comparison read 28.3 and 17.4 against 34.6.
- The Qwen license restricts operating the model as a service for other companies and says expressly that the requirement does not reach internal use.
- The group that published Meta's open models has posted nothing since 12 November 2025. Picking a family carries risk.
What an open weight model actually is
An open weight model is a language model whose trained numbers, called weights, are published as files anyone can download and run. The company that trained it keeps no switch over your copy. You supply the hardware. You supply the power. And the material you feed it never leaves the machine you chose.
A closed model is served only through an interface its maker controls, so every transcript you analyze is a transcript you transmitted.
Open weight is also not open source. Training data and training code are usually withheld. What you get is the finished artifact and a license saying what you may do with it, which is why the license file shipped with the model matters more here than any benchmark table.
How a research team handles confidential material today
PERSONAL EXPERIENCE Answer first: by moving people instead of moving software.
- The client flags the study as restricted, usually in the contract rather than in the brief.
- The institute assigns the transcripts to a named analyst and blocks the folder to a small group.
- That analyst reads and codes by hand, because the assistant everyone else uses sits on a server the contract does not allow.
- Work that could go faster, such as summarizing, first pass coding and cross study lookups, happens without help or happens twice.
- The report is assembled by hand, and the reviewer checks it against the raw material with both files open.
There is a second version of this, and it is worse. The analyst pastes an excerpt into a general assistant because the deadline is real and the excerpt looks harmless. Nobody logs it. The restriction survives on paper.
What the restricted study actually costs
The cost is not the hours. It is the two speeds.
An operation that automated the first pass on ordinary studies and still hand codes the restricted ones runs at two speeds. And the restricted work carries the most senior people, plus the tightest legal exposure. The senior analyst who was meant to be judging is back to clicking.
Whoever signs the confidentiality clause carries this, usually the account director, whose choice today is between a slow study and an undocumented shortcut. PERSONAL EXPERIENCE In our conversations with agencies, the restricted project is where the phrase comes out intact: we still do this one by hand.
How far behind the open models actually are
Close enough on ordinary work, not close enough on the hardest tasks. Both halves of that sentence are load bearing.
Artificial Analysis and Epoch AI put the distance between the best open model and the best closed one at roughly 2 to 3.4 points on the general composite indices, with the open side trailing by about four months. On the hardest exams in those same collections the distance is larger.
Then there is a result that looks like a tie and is not.
The tie appears on the old test and disappears on the new one
Two bar panels side by side, both scored out of one hundred, reported by Artificial Analysis for Terminal Bench versions 2.1 and 3.0. On the left, version 2.1: the first open model scores 88.3, the second open model scores 88.2, and the closed leader scores 88.8. The three bars are almost the same height, which reads as parity. On the right, version 3.0 of the same test: the first open model scores 28.3, the second open model scores 17.4, and the closed leader scores 34.6. All three bars collapse, and the closed leader now stands clearly above both open models. The reading below the panels is that the older version of the test was saturated, meaning every strong model was already at the top of it, so the tie measured the ceiling of the test rather than parity between the models.
Same benchmark, two versions. The parity on 2.1 is a property of the test, and it vanishes on 3.0.
Artificial Analysis publishes both versions. On the earlier one, two open models scored 88.3 and 88.2 against 88.8 for the closed leader. On the newer version of the same benchmark the comparison reads 28.3 and 17.4 against 34.6. The older test was saturated, meaning every serious model already sat near its ceiling, so the apparent tie between models was a tie against the test. Both numbers are real. Read the version and the testing method first.
The useful conclusion survives that correction. On what a research operation actually asks for, summarizing a transcript, applying a codeframe, drafting against a tabulation plan, the open models are close. On long chains of tool use with many dependent steps, they are not.
What the license lets a research institute do
The license decides, and most teams read it last.
The proprietary licenses attached to the open Chinese families restrict who may operate the model as a service for other companies, and they state expressly that the requirement does not reach internal use. The Qwen license file is published next to the model and can be read in five minutes.
For an institute that downloads a model, runs it inside its own network and analyzes transcripts for its own client work, that reading is favorable. For one that wants to open a portal where the end client talks to the model directly, the analysis has to be redone against the actual wording. That is a different arrangement.
None of this is legal advice and should not be treated as any. The reading passes through the client's legal team, the same way the confidentiality clause did.
The family that went quiet
One risk in choosing an open family has nothing to do with quality.
The group that published Meta's open models has posted nothing since 12 November 2025, per the last modified field on Hugging Face's public model API, consulted on 2 September 2026. The site for the old line now redirects to a page about a different family.
Note what that is. Nobody announced an ending. The observable fact is silence, and silence is what an institute plans around. A tracker that runs three years on a family nobody publishes to any more ages without anyone deciding that it should.
The defense is in the design. If each stage can swap its model without being rewritten, a family going quiet costs a weekend of testing. Welded to one model, it costs a move.
Where the material runs, stage by stage
UNIQUE INSIGHT The useful question is not which model to use. It is which stages are allowed to leave the building.
Splitting research stages by the confidentiality boundary
A flow drawn in two zones separated by a vertical boundary line. On the left, inside a solid box labelled inside your own infrastructure, four stages run in sequence on an open weight model: transcription of the recordings, removal of names and identifying details, first pass coding against the codeframe, and search across the client's own archive. Below that box a note says the raw recordings and transcripts never cross the boundary. On the right, inside a dashed box labelled outside, over an external service, three stages run on a closed model: scanning published literature, drafting language for the report from findings that carry no identifying detail, and a second reading by a model from a different maker. Between the two zones sits a gate labelled the de-identified extract, and an arrow crosses from left to right through that gate only. A return arrow crosses back carrying drafted text. The rule stated at the bottom is that the boundary is a contract clause, so it is drawn before any model is chosen.
Stages one to four hold material the contract protects, so they run on hardware the institute controls. Everything to the right of the gate carries no identifying detail.
Split the study into stages and ask the contract about each one. Transcription touches the raw recording, so it stays inside. Stripping names stays inside by definition, since that stage produces the extract allowed to travel. First pass coding touches full verbatims, so it stays inside. Searching the client's archive touches the client's archive.
Everything past the gate is a different conversation, because what crosses is no longer protected material. Consider scanning published literature, which never touched it, or drafting report language from de-identified findings, which does not need it.
Once the split is drawn, the model choice per stage gets small. Stages inside get an open model, the only kind that can run there. Stages outside get whichever model is best at that thing, a separate question covered in [one model, one blind spot](/blog/single-model-bias-market-research).
Where Cassi.ai comes in
Cassi.ai is a software engineering company specialized in the pains of market research, innovation and insights, working with research agencies and corporate insights teams. The relevant product here is the Super Agents line, deployed inside the client's own environment: on premises, in a private cloud, or air gapped.
Air gapped means the hardware running the model is physically disconnected from any outside network, so material cannot leave even by accident. Clients under strict confidentiality ask for it by name.
What we build there is the boundary itself: which stage runs where, which model serves each stage, and a record of what ran, who ran it and against which study. The rest is in the Cassi.ai portfolio. The platform argument sits in [six gaps a chat assistant leaves open](/blog/why-chatgpt-is-not-a-research-platform), and the policy this deployment has to satisfy is in [what the standard now demands in writing](/blog/written-ai-policy-for-research-agencies).
When the closed model is the right call
Sometimes running your own model is the wrong answer.
| If your situation is | Then | Because |
|---|---|---|
| No confidentiality clause restricting where data is processed | Use the closed model over an interface | You are paying for hardware to solve a problem you do not have |
| Long chains of dependent tool use, many steps | Closed model, for now | This is where the measured distance is widest |
| Restricted transcripts, ordinary analysis tasks | Open model inside your own network | The distance on these tasks is small and the clause is absolute |
| You want to offer the model to end clients directly | Stop and read the license against that arrangement | The internal use carve out does not describe this |
| No engineer available to keep hardware running | Closed model, and revisit later | A model nobody maintains is worse than a model nobody owns |
Index distances and version comparisons come from Artificial Analysis and Epoch AI. The mapping from those numbers to research stages is our own reading, built from client deployments.
FAQ
What is an open weight model?
A language model whose trained numbers are published as downloadable files, so it can run on hardware you control instead of on the maker's servers. The training data and the training code are usually not published, which is why open weight and open source are not the same thing.
Are open models good enough for market research work?
For summarizing transcripts, applying a codeframe and drafting against a tabulation plan, the distance is small. Artificial Analysis and Epoch AI measure roughly 2 to 3.4 points on the general composite indices, with the open side about four months behind. On long chains of dependent tool use the distance is much larger, and that is where a closed model still earns its place.
Does the license of an open Chinese model allow commercial use?
The Qwen license restricts operating the model as a service for other companies and says expressly that the requirement does not reach internal use. Running it inside your own network on your own client work reads favorably against that text. Opening a portal for end clients is a different arrangement and needs its own reading, by your legal team rather than by a vendor.
What happens if the model family I picked stops being published?
Plan for it. The group behind Meta's open models has published nothing since 12 November 2025 on Hugging Face's public API, with no announcement either way. If each stage can swap its model without being rewritten, that silence costs a weekend of testing rather than a migration.
Should we run the model on premises or in a private cloud?
UNIQUE INSIGHT Ask the contract, not the engineering team. On premises and private cloud are both answers to the same clause, and which one applies is written in the contract terms and in the list of countries the client attached. Air gapped is the answer when the clause says the material may not touch a network at all.
The decision that matters here was never open against closed. It is knowing which stages of your study are allowed to leave the building, and that answer is already written in a contract somebody signed. It does not change when a new model is released.
If you have a restricted study running now, the hour worth spending this week is with the confidentiality clause and the stage list side by side, marking which stage touches protected material. Everything else follows from that page.
Published by Cassi.ai. Read the full article at https://www.cassiai.com/blog/open-weight-models-in-market-research.