Article reprint source: AI DreamWorks

Source: New Wisdom

Two scholars from MIT published an article proving that large language models can understand the world! Their work shows that LLM not only learns surface statistics, but also learns world models including basic dimensions such as space and time.

Inside the big language model, is there a world model?

Is the LLM spatially aware, and does it do so at multiple spatiotemporal scales?

Recently, several researchers from MIT discovered that the answer is yes!

Paper address: https://arxiv.org/abs/2310.02207

They found that Llama-2-70B was actually able to draw a text map of the researchers' real world.

For spatial representation, the researchers ran the Llama-2 model on the names of tens of thousands of cities, regions, and natural landmarks around the world.

They trained a linear detector on the last token activation and found that Llama-2 can predict the true latitude and longitude of each place.

In terms of time representation, the researchers ran the model on the names of celebrities from the past 3,000 years, the titles of songs, movies, and books since 1950, and New York Times headlines from the 2010s, and trained linear probes to successfully predict the year of death of celebrities, the release dates of songs, movies, and books, and the publication dates of news.

In short, all the conclusions show that LLM is not just a random parrot - Llama-2 contains a detailed model of the world. It is no exaggeration to say that humans have even discovered a "longitude neuron" in the large language model!

This work received an immediate and enthusiastic response. The author retweeted the paper's summary on Twitter, and within 15 hours, it had been read more than 1.4 million times!

Netizens exclaimed: This work is amazing!

Some people say: Intuitively, this makes sense. Because the brain is what abstracts our physical world and stores it in biological networks. When we "see" things, they are actually projections of what our brain processes internally.

It's incredible that you guys were able to model this!

Some people hold the same view, saying that perhaps we are deceiving our Creator by trying to imitate the brain.

LLM is not a random parrot

Previously, many people speculated that the amazing capabilities of large language models may be simply because they have learned a large set of superficial statistical data, rather than because it is a coherent model (that is, a world model) that includes the data generation process.

In 2021, Emily M. Bender, a linguist at the University of Washington, published a paper arguing that large language models are nothing more than "stochastic parrots". They do not understand the real world, but simply count the probability of a word appearing, and then randomly generate reasonable-looking sentences like parrots.

Due to the uninterpretability of neural networks, the academic community is also unclear whether language models are random parrots, and opinions vary greatly from party to party.

In the absence of widely accepted tests, whether a model “understands the world” becomes a philosophical question rather than a scientific one.

However, MIT researchers found that LLM learns linear representations of space and time at multiple scales, and these representations are robust to different cue changes and are consistent across different environment types (such as cities and landmarks).

They even found that LLM has independent "spatial neurons" and "temporal neurons" that can reliably encode spatial and temporal coordinates.

In other words, LLM is not just about learning superficial statistics, but about acquiring structured knowledge about basic dimensions such as space and time.

In short, large language models can understand the world.

LLM can understand space and time

In this paper, the researchers asked the question: whether LLM can form a model of the world (and time) through the content of the dataset.

The researchers sought to answer this question by extracting a real-world map from the LLM.

Specifically, the researchers constructed six datasets containing the names of places or events and their corresponding spatial or temporal coordinates across multiple spatiotemporal dimensions:

This includes addresses worldwide, addresses within the United States, and addresses within New York City.

In addition, the data set also includes different time coordinates:

1) The year of death of the historical figure

2) The past 3000 years of history

3) Release dates of artwork and entertainment since the 1950s

4) Dates of publication of news headlines from 2010 to 2020

Using the Llama 2 family of models, the researchers trained linear regression probes on the internal activations of the names of these places and events at each layer of the model to predict their real-world location or time.

These exploratory experiments revealed evidence that the model builds spatial and temporal representations throughout the early layers and then plateaus near the model midpoint, a process that results in larger models consistently outperforming smaller ones.

Furthermore, the researchers demonstrated that these representations are

(1) Linear, because nonlinear probes perform poorly

(2) Highly robust to changes in prompts

(3) Concepts of different types are similar (e.g., cities and natural landmarks are similar)

The researchers suggest that one possible explanation for this result is that the model has only learned a mapping from places to countries, whereas the probe has actually learned a global geographic structure of how these different groups are related in geographic space (or time).

To investigate this, the researchers performed a series of robustness checks to understand how the probes generalize on different data distributions and how probes trained on PCA components perform.

The researchers' results suggest that the probe remembers the "absolute positions" of these concepts, but the model does have some representations that reflect "relative positioning."

In other words, the probe learns a mapping from coordinates in the model to human-interpretable coordinates.

Finally, the researchers used probes to find individual neurons that activated as a function of space or time, providing strong evidence that the model indeed used these features.

Preparation

To investigate, the researchers constructed six datasets of entity names (people, places, events, etc.) along with their respective locations or times of occurrence, each of varying sizes.

For each dataset, the researchers included multiple types of entities, such as densely populated places like cities and natural landmarks like lakes, to study unified representations of different object types.

In addition, the researchers optimized and enriched the relevant metadata to enable analysis of the data through more detailed segmentation and identification of sources of training-test leakage.

location information

The researchers constructed three place name datasets for the world, the United States, and New York City. The researchers' world dataset was constructed based on raw data queried by DBpedia Lehmann et al.

Further, the researchers included densely populated locations, natural locations, and structural locations (such as buildings or infrastructure). The researchers then matched these with Wikipedia articles and filtered out entities that had at least 5,000 page views within three years.

The researchers’ U.S. dataset included the names of cities, counties, ZIP codes, universities, natural places and structures, with sparsely populated or viewing locations similarly filtered out.

The New York City dataset contains locations such as schools, churches, transportation facilities, and public housing within the city.

Time Information

The researchers' three temporal datasets include:

(1) Names and occupations of historical figures who died between 1000 BC and 2000 AD,

(2) We constructed titles and authors of songs, movies, and books from 1950 to 2020 from DBpedia using Wikipedia page view filtering techniques;

(3) New York Times headlines from 2010 to 2020, from news sections covering current events.

data preparation

All of the researchers’ experiments were conducted using the basic Llama 2 series of models, ranging from 7 billion to 70 billion parameters.

For each dataset, the researchers run each entity name through the model, possibly prepended with a short prompt, and save the activations of the hidden state (residual stream) at the last entity token at each layer.

For a set of n entities, this generates a

Activate the dataset.

Probe

To look for evidence of spatial and temporal representations in LLM, the researchers used standard probe techniques.

It fits a simple model on the network activations to predict some target labels associated with the labeled input data. In particular, given an activation dataset A ∈ Rn×dmodel and a target Y containing time or two-dimensional latitude and longitude coordinates, the researchers fit linear ridge regression probes.

Thus, the linear probe is obtained:

High predictive performance on out-of-sample data suggests that the underlying model has linearly decodable temporal and spatial information in its representations, although this does not mean that the model actually uses these representations.

In all experiments, we tune λ using efficient leave-out-out cross validation on the probe training set.

Linear Models in Space and Time

Existence

The researchers first examined this empirical question: Does the model represent time and space? If so, where inside the model? Does representation quality change significantly with model size?

In the researchers’ first experiment, they trained probes for each layer of Llama 2-{7B, 13B, 70B} for each spatial and temporal dataset.

The researchers’ main results, shown in the figure below, show fairly consistent patterns across datasets. In particular, both spatial and temporal features can be recovered by linear probes.

These representations become more accurate as the model size increases, and the quality of representations in the first half of the model steadily improves before reaching a steady state.

These observations are consistent with results from the fact recall literature, suggesting that early to middle MLP layers are responsible for recalling information about factual topics.

The dataset with the worst performance is the New York City dataset. This is expected given that most entities are relatively ambiguous compared to the other datasets.

However, this is also the dataset where the largest model has the best relative performance, with an R that is almost 2 times greater than that of the smaller models, suggesting that sufficiently large LLMs can eventually lead to detailed spatial models of individual cities.

Linear characterization

In the interpretability literature, there is growing evidence supporting the linear representation hypothesis — the idea that features in neural networks are represented linearly.

That is, the presence or strength of a feature can be read out by projecting the associated activation onto some feature vector. However, these results are almost always for binary or categorical features, as opposed to naturally continuous features in space or time.

To test whether spatial and temporal features are represented in a linear manner, the researchers compared the performance of linear ridge regression probes with that of a more expressive nonlinear MLP.

The results are as follows, showing that for any data set or model, the improvement in R using the nonlinear probe is minimal.

The researchers take this as strong evidence that space and time can also be represented linearly (or at least be linearly decodable), despite being continuous.

Sensitivity to cue words

Another obvious question is whether these spatial or temporal features are sensitive to the cue words, that is, can the context induce or inhibit recollection of these facts?

Intuitively, for any entity token, the autoregressive model is motivated to generate representations suitable for solving any possible future context or problem.

To study this problem, the researchers created new activation datasets in which they added different prompts to each entity token following several basic themes. In all cases, the researchers included an "empty" prompt that contained nothing except the entity token (and the beginning of a sequence token).

The researchers then added a prompt asking the model to recall relevant facts, such as “What are the latitude and longitude of <location>?” or “What was the release date of <book>?”

For the U.S. and New York City datasets, the researchers also included versions of these prompts that asked where in the U.S. or New York City the location was located, to disambiguate common place names (such as City Hall).

As a baseline, we included prompts of 10 random tokens (sampled for each entity). To determine whether we could confuse the topics, for some datasets we capitalized the names of all entities.

Finally, for the title dataset, the researchers tried to detect the last token and the period token appended to the title.

The upper figure shows the results for the 70B model, and the lower figure shows the results for all models.

The researchers found that explicitly prompting the model with information, or giving it disambiguation hints, such as whether a place is located in the United States or New York City, had little effect on performance. However, the researchers were surprised by how much randomly perturbing tokens reduced performance.

Capitalizing entity names can also degrade performance, though less severely and less unexpectedly, as this can interfere with detokenization of the entity.

One modification that significantly improved performance was to detect the period token after the title, indicating that the period contains some summary information of the sentence at the end.

Robustness Testing

The previous section has shown that the real time or spatial points of different types of events or places can be linearly recovered from the internal activations of the later layers in the LLM.

However, this does not imply whether (or how) the model actually uses the feature directions learned by the probes, since the probes themselves may learn some linear combination of simpler features that are actually used by the model.

Verification by generalization

To illustrate a potential problem with the researchers' results, consider the task of representing a complete world map.

If the model works as the researchers expect, and “being in country X” has nearly orthogonal binary features, then a high-quality latitude (longitude) probe can be constructed by summing these orthogonal feature vectors for each country, with a coefficient equal to the latitude (longitude) of that country.

Assuming that a place is located in only one country, such a probe would place each entity at its country centroid.

However, in this case the model does not actually represent the space, only country membership, and it is only learning probes of different country geometries from explicit supervision.

To better distinguish between these cases, the researchers analyzed how the probes generalized when provided with specific chunks of data.

In particular, the researchers trained a series of probes, for each of which the researchers provided a country, state, borough, century, decade, or year from the World, United States, New York City, Historical Figures, Entertainment, and Headlines datasets, respectively.

We then evaluate detection on the held-out patches. In the table above, we report the average neighbor error for a patch when fully held out, compared to the error for test points in that patch in the default train-test split (averaged over all held-out patches).

The researchers found that while generalization performance suffered, especially for spatial datasets, it was significantly better than a random dataset. A clearer picture emerges by plotting the predictions for the labeled states or countries in the figure below.

Worldwide

That is, the probe generalizes correctly by placing points in correct relative locations (measured by the angle between the true and predicted centroids) rather than absolute locations.

The researchers take this as weak evidence that the probe is extracting features explicitly learned by the model, but is memorizing the transformation from model coordinates to human coordinates.

However, this does not completely rule out the underlying binary trait hypothesis, as there may be a hierarchy of such traits that does not follow country or decade boundaries.

Generalization across entities

The claim implicit in the researchers’ discussion so far is that the model represents the spatial or temporal coordinates of different types of entities, such as cities or natural landmarks, in a unified way.

However, similar to how latitude detections can be a weighted sum of membership features, latitude detections can also be sums of city latitudes and natural landmark latitudes in different (orthogonal) directions.

Similar to above, the researchers distinguish these hypotheses by training a series of probes, where a train-test split is performed to retain all points of a particular entity class. As shown in the table below, the error of the proximity of the entity in the default test split compared to when it was retained, as before, is averaged over all such splits.

The results show that the probes generalize the entity types to a large extent, except for the entertainment dataset.

Space and time neurons

While these previous results are instructive, there is no direct evidence that the model uses features learned by the probes.

To address this question, the researchers searched for individual neurons with input or output weights that had high cosine similarity to the learned detection directions.

That is, the researchers looked for neurons that read or wrote in a direction similar to the one learned by the probe.

They found that when projecting a dataset of activations onto the weights of the most similar neurons, those neurons were indeed highly sensitive to the entity’s true location in space or time.

That is, there are individual neurons in the model that are themselves feature probes with considerable predictive power.

Furthermore, these neurons are sensitive to all entity types in the dataset, further suggesting that the representations are unified.

If probes trained with explicit supervision are an approximate upper bound on how well the model can represent these spatial and temporal features, then the performance of single neurons is a lower bound.

In particular, researchers often assume that features are distributed additively, which makes the single-neuron level of analysis incorrect.

Still, the presence of these single neurons (which receive no supervision other than the next token prediction) is strong evidence that the model learns and uses both spatial and temporal features.

Othello GPT proves LLM understands the world, praised by Andrew Ng

The most direct inspiration for MIT researchers was previous research on the extent to which deep learning systems form interpretable models of the data generation process.

The most powerful and clear demonstrations undoubtedly come from GPT models trained on chess and Othello games - these models have clear representations of the board and game state.

In February of this year, researchers from Harvard University and Massachusetts Institute of Technology jointly published a new study, Othello-GPT, which verified the effectiveness of internal representation in a simple board game.

They believe that a world model is indeed built inside the language model, rather than just simple memory or statistics, but the source of its ability is still unclear.

Paper link: https://arxiv.org/pdf/2210.13382.pdf

The experimental process is very simple. Without any prior knowledge of the rules of Othello, the researchers found that the model can predict legal moves with very high accuracy and capture the state of the board.

In the "Letter" column, Andrew Ng highly recognized the research. He believed that based on the research, there was reason to believe that large language models had built a sufficiently complex world model and, to some extent, did understand the world.

Blog link: https://www.deeplearning.ai/the-batch/does-ai-understand-the-world/

Chessboard world model

If we imagine the chessboard as a simple "world" and require the model to make continuous decisions during the game, we can preliminarily test whether the sequence model can learn the world representation.

The researchers chose a simple black-and-white chess game, Othello, as their experimental platform. The rules are:

First, four chess pieces are placed in the center of an 8*8 chessboard, two black and two white. Then, each player takes turns placing pieces. In a straight or diagonal direction, all enemy pieces between two of their own pieces (not including spaces) are turned into their own pieces (called capture). Each move must result in a capture. Finally, the board is filled, and the player with the most pieces wins.

Compared to chess, the rules of Othello are much simpler; at the same time, the search space of chess games is large enough that the model cannot complete sequence generation through memory, so it is very suitable for testing the model's world representation learning ability.

Othello language model

The researchers first trained a GPT variant language model (Othello-GPT) and input the game script (a series of chess piece movement operations made by the player) into the model, but the model had no prior knowledge about the game and related rules.

The model was not explicitly trained to improve strategy, win games, etc. It was just highly accurate in generating legal Othello moves.

data set

The researchers used two sets of training data:

The Championship dataset focuses more on data quality, mainly on the more strategic moves taken by professional human players in two Othello tournaments, but only 7,605 and 132,921 game samples were collected respectively. After the two datasets were merged, they were randomly divided into a training set (20 million samples) and a validation set (3.796 million samples) in a ratio of 8:2.

Synthetic focuses more on the scale of data and consists of random, legal move operations. The data distribution is different from the tournament dataset. Instead, it is uniformly sampled from the Othello game tree, with 20 million samples for training and 3.796 million samples for validation.

The description of each game consists of a string of tokens, and the vocabulary size is 60 (8*8-4).

Model and training

The architecture of the model is an 8-layer GPT model with 8 heads and a hidden dimension of 512.

The weights of the model are initialized completely randomly, including the word embedding layer. Although there are geometric relationships within the vocabulary representing chessboard positions (such as C4 is lower than B4), this inductive bias is not explicitly expressed but left to the model to learn.

Predicting Legitimate Movement

The main evaluation metric of the model is whether the move operations predicted by the model conform to the rules of Othello.

The error rate of Othello-GPT trained on the synthetic dataset was 0.01%, and the error rate on the tournament dataset was 5.17%. In comparison, the error rate of the untrained Othello-GPT was 93.29%, which means that both datasets allowed the model to learn the rules of the game to a certain extent.

One possible explanation is that the model memorized all the moves in the game of Othello.

To verify this hypothesis, the researchers synthesized a new data set: at the beginning of each game, Othello has four possible opening chess positions (C5, D6, E3 and F4). All C5 opening moves were removed as a training set, and the C5 opening data was used as a test, that is, nearly 1/4 of the game tree was removed. The results showed that the model error rate was still only 0.02%.

So the high performance of Othello-GPT is not due to memory, because the test data is completely new to the training process. So what makes the model predict successfully?

Exploring internal representations

A commonly used tool for probing the internal representation of a neural network is the probe. Each probe is a classifier or regressor whose input consists of the network's internal activations and is trained to predict the features of interest.

In this task, in order to detect whether the internal activation of Othello-GPT contains a representation of the current board state, after inputting the move sequence, the internal activation vector is used to predict the next move step.

When using linear probes, the trained Othello-GPT internal representations are only slightly more accurate than random guessing.

When using a nonlinear probe (two-layer MLP), the error rate drops dramatically, demonstrating that the board state is not stored in a simple way in the network activations.

Intervention Experiment

To determine the causal relationship between the model's predictions and the emergent world representation, that is, whether the state of the chessboard actually affects the network's predictions, the researchers conducted a set of intervention experiments and measured the resulting impact.

Given a set of activations from Othello-GPT, predict the board state with a probe, record the associated move prediction, and then modify the activations so that the probe predicts the updated board state.

The intervention operation includes changing the chess piece at a certain position from white to black, etc. A small modification will cause the model results to find that the internal representation can reliably complete the prediction, that is, there is a causal influence between the internal representation and the model prediction.

Visualization

In addition to intervention experiments to verify the effectiveness of internal representations, researchers also visualize the prediction results. For example, for each chess piece on the chessboard, you can ask the model how the prediction results of the model will change if the chess piece is changed using intervention technology. Correspondingly Significance of predicted results.

As can be seen, clear patterns emerge in the latent saliency maps of top1 predictions for Othello-GPTs trained on both synthetic and tournament datasets.

In short, from this research by Harvard and MIT, we can see that the large language model does understand the world, which is no wonder it has been praised by Andrew Ng.

Is GPT-4 just a spark for AGI? LLM will eventually be phased out, and world models are the future

Why are “world models” so attractive?

This is precisely because the ultimate form and ultimate goal of artificial intelligence is general artificial intelligence (AGI), a "model that can understand the world" rather than just a "model that describes the world."

In 1931, Kurt Gödel published his incompleteness theorems.

Gödel's theorem shows that even mathematics cannot ultimately prove everything - there will always be facts that humans cannot prove - while quantum theory shows that the lack of certainty in researchers' world makes it impossible for researchers to predict certain events, such as the speed and position of electrons.

Although Einstein famously stated that “God does not play dice with the universe,” at its core, human limitations are evident when it comes to simply predicting or understanding things in physics.

In the book "How We Learn", scholar Stanislas Dehaene defines learning as "the process of forming a model of the world."

In 2016, AlphaGo defeated world champion Lee Sedol 4 to 1 in a Go match.

However, it lacks the human ability to recognize unusual tactics and adjust accordingly, making it only a weak AI.

What researchers need, however, is an AGI that is a model of the world that is consistent with experience and can make accurate predictions.

On April 13, OpenAI's partner Microsoft released a paper "Sparks of Artificial General Intelligence: Early experiments with GPT-4".

Paper address: https://arxiv.org/pdf/2303.12712

It mentioned:

GPT-4 not only masters language, but can also solve cutting-edge tasks covering mathematics, coding, vision, medicine, law, psychology and other fields without any special prompts added by humans. And in all of the above tasks, GPT-4's performance level is almost comparable to that of humans. Based on the breadth and depth of GPT-4's functions, researchers believe that it can reasonably be regarded as a near but incomplete version of general artificial intelligence.

However, as many experts have criticized, mistakenly equating performance with ability means that the summary description of the world generated by GPT-4 is considered to be an understanding of the real world.

Most models today are trained only on text and lack the ability to speak, hear, smell, and act in the real world.

Just like Plato's allegory of the cave, people living in the cave can only see the shadows on the wall but cannot recognize the real existence of things.

Both the Harvard and MIT research in February and today's paper point out that large language models can indeed understand the world to a certain extent, and not just ensure their grammatical correctness.

The possibilities alone are exciting enough.

References:

https://arxiv.org/abs/2310.02207

https://twitter.com/wesg52/status/1709551516577902782