The bit I don’t want to rush
Using several LLMs is teaching me which work I want to spend time on, and which work to leave alone. My brain feels like it’s getting a workout. I’m still learning when to put the weights down.
I went home from an alumni event with an idea for a research project and started building it with language models. I’m keeping the research itself for another day. For now, I’m learning quite a lot about which bits of the work I actually want to do.
I love the methodology. Psychometrics and study design are the bits I keep coming back to: what a measure captures, where a category begins and ends, and whether a study could answer the question I’ve asked of it.
I can apparently spend an unreasonable amount of time arguing about what a category means. I am having a lovely time.
This has now reached the group chat. I’d just tried a new brisket place that had opened and my mate Andy asked for a rating.
Andy: Rating out of 10?
Me: What's the instrument being used to measure?
Andy: Quality of beef brisket
Me: 7
My immediate thought was whether I was rating the quality of the brisket, the speed of delivery or the meal as a whole. What was the number supposed to mean?
Once he’d specified brisket quality, I started thinking about how we could quantify the subjective bits of it. Moisture and seasoning, presumably, but how much was down to the experience of the person doing the smoking? Then there was the person rating it, with their own preferences and expectations. If delivery took longer, or different packaging changed the condition it arrived in, what exactly would their score be a rating of?
Andy wanted a number. I had concerns about construct validity. The brisket got a seven.
I knew I liked research. I hadn’t quite appreciated how much of the enjoyment was in working out how to do it properly. The part where someone else might reasonably lose patience with the definitions is often where I’m getting interested.
My main workspace is Visual Studio Code, using LLMs from Anthropic and OpenAI alongside Python scripts and local files. There is quite a lot to learn before that setup is useful. I need to give the agents jobs they can do, keep track of the files and get back something I can check. I’ve got better at it by doing it, including getting it wrong.
Keeping up
I’ve written about the usage bar before. At first I struggled to keep up with one model. Later I was running several threads and getting irritated that they weren’t keeping up with me. I filled the wait for one by going to another, until there was always something waiting for me.
Some of that is getting more skilled. Some of it is being very good at generating things I now have to read. I would prefer to count all of it as progress, for obvious reasons.
Several outputs arriving together do not give me several brains to read them with. I still have to remember what I asked for and understand what came back. That takes time even when generating the file doesn’t.
I’ve even started using “context” when I talk to people about my own understanding. What I know about a problem, what I’m missing, what someone needs to bring me up to speed on. I’ve caught myself doing it in ordinary conversation, well away from the editor.
There is a feeling that goes with it, though, which I’m interested in. The more context I hold in my head, the more my memory feels like a muscle stretching to accommodate it. I seem able to hold more of the work together. That may just be familiarity with a complicated project. I haven’t measured it, and I’m trying not to announce a cognitive upgrade on the strength of a feeling and a subscription.
The practical skill I need is controlling the pace. If an agent has a well-defined mechanical job, I want it to get on with it while I do something else. If we’re changing what a study measures, I want to stop and think before sending it off again. I’m learning to specify where the work stops as carefully as what it does.
Otherwise I can spend the whole session keeping up with the agents and barely touch the methodological thinking I started the project for.
Leaving it alone
I’m also learning how much skill there is in choosing what not to do.
An idea can become a task almost as soon as I’ve had it. I ask for another comparison or a different version of the instructions, and soon there is a file to read. Giving something a filename is a surprisingly effective way of making it feel like an obligation.
Before I ask another model, I need to know what its answer could settle. Sometimes I want a different reading or a check of a particular claim. Sometimes I suspect I’m asking because I’d like to feel more certain. There is no obvious limit to how many files I could produce in pursuit of that feeling.
The same problem comes up when I revise the method. A model can help me think through a definition or show me a case that doesn’t fit. I still have to decide what to do about it. I need to work through why a definition would be better and what changing it would do to the study. That can take a while. I’m usually quite happy doing it.
I’m trying to finish the sentence “I need this output because…” before I start the next job. Exploring something because I’m interested is a perfectly good answer. I want to keep doing that. But an interesting question can belong in a different project, or sit in a note until I have time for it. It doesn’t need a running agent by the end of the sentence.
The checking needs time too. I have already written about what happens when I don’t do it properly, so I have rather spoiled my chances of pleading ignorance. If I’m going to use an output, I need to understand and verify it. The optional extra comparison can wait until I’ve checked the one I already have.
I can’t claim that I’m faster overall. I haven’t measured that, and I’ve criticised the assumption before. It would be convenient if the rules changed when I was enjoying myself, but I don’t think they do.
What I do know is that I want more time for the methodology. I like thinking about measurement. I want to learn more about study design. These tools have let me spend a lot of time on both, and I’m pleased about that. I also need to let myself sit with a question for a while without commissioning something else to fill the gap.
At some point I’ll have to get used to leaving an agent idle. I realise this is a peculiar problem to have bought for myself.