Research

Research we can stand behind

Rewriting a vendor's documentation is not research. What we publish here comes from building and testing, with the method and the data attached. This page sets the standard, links the studies we have completed, and lists the ones we have registered and not yet run.

Completed studies

What every study must include

Research question
One sentence, written before any data is collected.
Methodology
Enough detail that someone else could repeat it.
Sample
What was tested, how it was chosen and how many.
Exact dates
When each measurement was taken, because platforms change week to week.
Environment
Product, plan, surface, model and versions in use.
Repetitions
How many times each case was run. Model output varies, so one run is an anecdote.
Raw and derived data
Published as CSV or JSON where it is safe to do so, with client data removed.
Limitations
What the study cannot tell you, stated next to the findings.
Methodology changes
A dated log of anything changed after the first publication.
How to cite
A stable URL, a title, a date and a suggested citation.

Two rules sit above the list. We do not manufacture data. And we do not present a planned study as a finished one.

Registered studies

Registered

ChatGPT plugin review log

Question
How long does OpenAI's plugin review take, and what feedback does it return, for plugins built by one studio?
Method
For every plugin we submit: timestamps for upload, automated checks, review start, each piece of feedback and the decision. Feedback is recorded in full with client names removed. No sampling, every submission is included.
When it publishes
As each review completes, then updated with every submission. A single case will be labelled as a single case.
Registered, not started

Tool description wording and tool selection

Question
For the same MCP tool, how much does the wording of its description change how often ChatGPT selects it for direct, indirect and negative prompts?
Method
One test server, a fixed set of prompts in three groups, several description variants, repeated runs per variant on a stated plan and surface. Precision and recall recorded per variant, following the evaluation approach in OpenAI's metadata guide.
When it publishes
When the protocol has been run in full. Prompt set and results as JSON.
Registered, not started

Sign in friction in plugin first use

Question
What share of people who invoke a plugin for the first time complete account linking, and does offering a tool that needs no sign in change it?
Method
Server side counts of first tool calls, authorization starts and authorization completions on plugins we operate or are given access to, with the owner's permission. Reported as counts, not only percentages.
When it publishes
Only once there is enough volume to report without identifying a client.

Sourced reference guides

These are not studies. They are guides built from primary sources, where each article ends with a ledger of every claim, its source, the source date and the date we verified it.

Corrections

If something we published is wrong or out of date, write to sales@houseofmvps.com with the page and the source. We correct the text, add a dated note saying what changed, and keep the original publication date. The service these studies support is ChatGPT plugin development.