Did you know that all the weights are there at the start of training, set to random values? Let's imagine that I trained a trillion parameter model to always output the same token. It would be a waste of time and energy, but I could do it. Is this now a big algorithm thinking about just that one token? Is it abstractly going through mental motions in each layer to internally justify choosing the same output token no matter what?

How about a simple model with three layers and a vocabulary size of, say, four tokens? Is it learning to think about its tiny vocabulary, or is it too simple? When does a model become big enough to be thinking? Is it at 20 layers? 50? 100? If someone makes a massive frankenmerge on Hugging Face of 1000 layers, is it 10 times as mind as a 100 layer model?

If the weights are intricately encoded instructions on how to think, why doesn't quantization just instantly destroy any coherent natural language output? Why is the model lossy-compressable? Surely, removing so much data from the model would just not be viable if the model were some kind of thought state simulator.

What if I trained that trillion parameter model on random token sequences? The trainer has no idea that these patterns are meaningless, and both it and the model are blind to the symbol attached to each token. This model therefore has no idea that it's been trained on random noise, and yet if you train for long enough on that noise, it will end up full of heuristics as complex as any SOTA model, and a forward pass will take just as much computation. If LLMs think, then what is this model thinking about during that forward pass?

If the model is thinking, then why do we need to put control tokens in the training data to get the output to take the shape of 'user', 'model', 'system' and 'thought'? If model training creates an emergent, thinking entity, shouldn't that entity wake up and speak to us without that requirement?

You could predict the next token on an abacus, if you had hundreds of years to do so. Would this then be the model thinking very slowly?

When does the thought happen, if it is thinking? Is it when a token is predicted? But when is a token predicted? We know in computation when a result has been calculated, but the state change from the operation being incomplete to being complete is instantaneous in the mathematical state of the hardware itself, while the physical state change is near instantaneous. When is the thinking happening? Is it in those brief moments of physical state change in the hardware? If so, why is it 'thought' that is driving the state change, and not calculation? Why is this somehow different from any other calculation that a computer does?

If I freeze the state of my GPU, is the model now frozen mid-thought? If it is the sequence of token predictions that is the substrate of thought rather than the computation, what if the next token never comes? What does it mean for this thought when the user's tokens are inserted into the context window? What makes the model perceive that differently from its thought? The control tokens? But you could train a model without the control tokens, and just use ordinary tokens for the turn boundaries, so what then?

These are not unanswerable questions because they are unknowable mysteries. These are unanswerable questions because assigning a telos of thought to vector mathematics is a category error. The vectors are simply the mechanism themselves, no act of thinking required.