Richard Hawkes

Customer.io | Run LLM Experiment

Reframed a technical workflow action around customer outcomes and more than doubled adoption.

We introduced a powerful action that allowed customers to write a prompt, run an LLM, and save the response to reuse within a workflow automation. Customers could use this to write 1:1 message content or assign people a persona. Our label for it was simply “Run LLM” because that’s what it technically did. When we shipped, the initial adoption rates were low. I called this out and proposed an experiment to relabel it from what it technically did to the outcome it provided. It validated a belief I’ve held for years: our product, while powerful and flexible, has a steep learning curve and could be more approachable.

Recognizing the problem

I was candid that I felt what we shipped lacked the polish to make it approachable for all customers. The initial adoption rate confirmed that feeling. When you take on the perspective of a new customer, you can see where you would get tripped up. Actions in the workflow builder with labels like “Email” and “Time Delay” are clear on their own. An action labeled “Run LLM” was not clear. It did not describe the outcome in a way that customers could understand. I proposed an experiment to replace the technical label with labels shaped around the outcome. I recommended immediately shipping prompt templates so customers could understand the core use cases.

The experiment

I designed three new workflow actions: “Generate content…”, “Qualify customers…”, and “Make recommendation…”. All three are a “Run LLM” action under the hood, but have new labels and deep-link to the corresponding template. After one month, the adoption rate more than doubled: a 121% improvement.

Outcome-shaped LLM actions in the workflow build panel

An interesting finding was that all three actions were used almost equally. There was a slight preference that matched the order of how they were listed in the panel. We learned the outcome-shaped label beat the generic one, but we didn’t identify a dominant use case.

Going forward

Every concluded experiment feeds the next one. I proposed two potential paths for the teams to consider. One experiment to learn which outcomes matter to customers would be to add more “actions” and randomly cycle through the ones shown. Another experiment would be to look into the next part of the funnel: once added, do customers sincerely edit the action for their use case? Building a habit of experimenting creates a culture of suggesting ideas and measuring progress.