Tag: embedding
Teaching Computers Context: Making Meaning Mathematical with Embeddings
Posted by bsstahl on 2026-09-12 and Filed Under: development
Introduction
Imagine trying to teach a computer the subtle difference between the phrases "I need to address this issue" and "What's your address?" We recognize the intended meaning immediately because we understand the context in which the word is used. A standard text encoding, however, captures the characters in the text, not the concepts those characters represent.
This is the gap that embeddings help us bridge. An embedding represents a concept as a numerical vector, and once we have that mathematical representation, we can compare concepts, group them, and even perform arithmetic on some of their relationships. That is the real power of embeddings: they make aspects of meaning computable.
For example, consider the word "Ram." Depending on its context, it can represent several entirely different concepts:
- Computing: "I'm getting more RAM for my PC" (computer memory)
- Automotive: "I'm getting a Ram so I can pull my boat" (truck model)
- Agriculture: "I'm getting a ram and a ewe" (male sheep)
The characters alone cannot tell us which of these concepts is intended, so how can a computer distinguish among them? An embedding can represent each usage based on the ideas that surround it, giving us a mathematical way to make that distinction.
From Encoded Text to Represented Meaning
Before a machine learning model can process text, that text must be converted into numerical values. ASCII and UTF encodings represent characters as numbers, while model-specific tokenization processes assign numbers to tokens. A model can operate directly on any of these values, but there is an important limitation: the numerical order of the values does not inherently tell us anything about the relationships among the text they represent. Token 101, for example, is not necessarily more similar in meaning to token 102 than it is to token 900.
An embedding gives us an additional, learned representation. It maps a token, word, phrase, sentence, or other item to a vector, which is an ordered list of numbers that we can treat as a point in a high-dimensional space. Unlike the arbitrary relationships implied by the original numeric encoding, the position of this point can capture relationships learned from the data.
Models learn these representations through exposure to large amounts of data. The underlying idea, known as the distributional hypothesis, is that text used in similar contexts tends to have related meanings and should therefore develop related representations. The exact training process varies by model; some models create a single representation for a word, while others create representations that change based on the context. For our purposes, however, the important result is the same: concepts can be placed into a space where we can examine their relationships mathematically.
Visualizing the 'Ram' Example
Let's return to our three uses of "Ram." To visualize the idea, imagine a deliberately simplified embedding space with only three dimensions, represented as a cube. For this example, the axes correspond to the computing, automotive, and agricultural domains. Real embedding spaces usually have many more dimensions, and those dimensions generally do not have meanings that are this clear or human-readable, but this abstraction gives us a useful place to start.
- "I need to buy more RAM for my PC" would be positioned near concepts such as "Memory," "Storage," and "ROM."
- "I need to buy a Ram to haul my boat" would be positioned near "Truck," "Vehicle," and "Dodge."
- "I put the ram in the pen with the chickens" would be positioned near "Sheep," "Farm," and "Livestock."
Each sentence uses the same three letters, but its embedding occupies a different region of the space because it represents a different concept. We now have something that was not available when the text was merely encoded: a geometry that describes mathematical relationships among meanings.
Operating on Meaning
Once concepts are represented as vectors, standard mathematical operations become tools for working with meaning. What kinds of operations can we perform, and what can they tell us?
Measuring Similarity
Distance and similarity calculations tell us how close two vectors are. The embedding for the computing use of "RAM," for example, should be closer to "Memory" than to "Truck," while the automotive use should reverse that relationship. Depending on the model and the task, a system might use Euclidean distance, cosine similarity, a dot product, or another metric. These calculations are not interchangeable in every situation, but each can turn the geometric relationship between vectors into a useful numerical score.
Finding Groups
Clustering algorithms can identify regions that contain related vectors. Even without explicit labels, a collection might form distinct groups around computing, transportation, and agriculture. A classification system can then compare a new vector with known examples to determine which group it most closely resembles.
Examining Directions and Arithmetic
The differences between vectors can sometimes capture relationships between concepts, producing familiar examples such as:
king - man + woman ≈ queenparis - france + italy ≈ rometeacher - school + university ≈ professor
These analogies illustrate how a direction through the vector space can represent a conceptual relationship. They should not be treated as universal laws; whether they work depends on the model, its training data, and the specific concepts involved. Even with that limitation, they demonstrate a broader and more interesting point: embeddings support more than lookup or comparison. They give us a mathematical structure in which some conceptual relationships can be manipulated.
What This Makes Possible
These operations are the foundation for many practical systems:
- Semantic search compares the embedding of a query with the embeddings of documents or passages, allowing it to retrieve related ideas even when they use different words.
- Classification compares new text with known categories or examples, supporting scenarios such as sentiment analysis and spam detection.
- Clustering and recommendations group related items and help surface content similar to something a user already values.
- Question answering identifies passages related to a question, recognizing, for example, that "Who created C#?" and "Who is C#'s inventor?" express nearly the same intent.
Of course, putting these ideas into production introduces additional choices. Which embedding model and similarity metric should we use? Should each vector represent a word, a sentence, a paragraph, or some other chunk of content? How will the vectors be stored and searched? Multilingual and multimodal models extend the same geometric approach across languages, images, audio, and other kinds of data. The implementation details vary, but they all build on the same foundation: represent concepts as vectors, then use mathematics to work with their relationships.
Limitations and Challenges
It is important to remember that embeddings do not contain objective or complete definitions of concepts. They are learned representations shaped by a model's architecture, training data, and objectives. As a result, they can reproduce biases, miss domain-specific meanings, become less useful as language changes, or place concepts together for reasons that are difficult to interpret.
Similarity is also highly dependent on both the model and the task. Two vectors being close together means that a particular model considers them related; it does not explain the relationship or guarantee that the relationship is useful for our scenario. High-dimensional vectors can also require substantial storage and computation at scale. We should therefore evaluate embeddings against the actual task we need to perform rather than assume that geometric elegance guarantees correct results.
Conclusion
ASCII, UTF, and tokenization make text available to computation by assigning it numbers. Embeddings take us an important step further by giving learned concepts positions in a mathematical space. Within that space, the computing, automotive, and agricultural meanings of "Ram" can occupy different neighborhoods even though the original text is identical.
That shift, from encoded text to represented meaning, is what makes embeddings so powerful. Concepts become vectors, similarity becomes distance, categories become clusters, and some relationships become directions that we can explore using arithmetic. Embeddings do not give computers human understanding, but they do make useful aspects of meaning available to mathematics, and that gives us an extraordinary set of tools for building systems that operate on concepts rather than just characters.
Tags: ai algorithms data-structures embedding ml math
The Depth of GPT Embeddings
Posted by bsstahl on 2023-10-03 and Filed Under: tools
I've been trying to get a handle on the number of representations possible in a GPT vector and thought others might find this interesting as well. For the purposes of this discussion, a GPT vector is a 1536 dimensional structure that is unit-length, encoded using the text-embedding-ada-002 embedding model.
We know that the number of theoretical representations is infinite, being that there are an infinite number of possible values between 0 and 1, and thus an infinite number of values between -1 and +1. However, we are not working with truly infinite values since we need to be able to represent them in a computer. This means that we are limited to a finite number of decimal places. Thus, we may be able to get an approximation for the number of possible values by looking at the number of decimal places we can represent.
Calculating the number of possible states
I started by looking for a lower-bound for the value, and incresing fidelity from there. We know that these embeddings, because they are unit-length, can take values from -1 to +1 in each dimension. If we assume temporarily that only integer values are used, we can say there are only 3 possible states for each of the 1536 dimensions of the vector (-1, 0 +1). A base (B) of 3, with a digit count (D) of 1536, which can by supplied to the general equation for the number of possible values that can be represented:
V = BD or V = 31536
The result of this calculation is equivalent to 22435 or 10733 or, if you prefer, a number of this form:
10000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000
Already an insanely large number. For comparison, the number of atoms in the universe is roughly 1080.
We now know that we have at least 10733 possible states for each vector. But that is just using integer values. What happens if we start increasing the fidelity of our value. The next step is to assume that we can use values with a single decimal place. That is, the numbers in each dimension can take values such as 0.1 and -0.5. This increases the base in the above equation by a factor of 10, from 3 to 30. Our new values to plug in to the equation are:
V = 301536
Which is equivalent to 27537 or 102269.
Another way of thinking about these values is that they require a data structure not of 32 or 64 bits to represent, but of 7537 bits. That is, we would need a data structure that is 7537 bits long to represent all of the possible values of a vector that uses just one decimal place.
We can continue this process for a few more decimal places, each time increasing the base by a factor of 10. The results can be found in the table below.
| B | Example | Base-2 | Base-10 |
|---|---|---|---|
| 3 | 1 | 2435 | 733 |
| 30 | 0.1 | 7537 | 2269 |
| 300 | 0.01 | 12639 | 3805 |
| 3000 | 0.001 | 17742 | 5341 |
| 30000 | 0.0001 | 22844 | 6877 |
| 300000 | 0.00001 | 27947 | 8413 |
| 3000000 | 0.000001 | 33049 | 9949 |
| 30000000 | 0.0000001 | 38152 | 11485 |
| 300000000 | 0.00000001 | 43254 | 13021 |
| 3000000000 | 0.000000001 | 48357 | 14557 |
| 30000000000 | 1E-10 | 53459 | 16093 |
| 3E+11 | 1E-11 | 58562 | 17629 |
This means that if we assume 7 decimal digits of precision in our data structures, we can represent 1011485 distinct values in our vector.
This number is so large that all the computers in the world, churning out millions of values per second for the entire history (start to finish) of the universe, would not even come close to being able to generate all of the possible values of a single vector.
What does all this mean?
Since we currently have no way of knowing how dense the representation of data inside the GPT models is, we can only guess at how many of these possible values actually represent ideas. However, this analysis gives us a reasonable proxy for how many the model can hold. If there is even a small fraction of this information encoded in these models, then it is nearly guaranteed that these models hold in them insights that have never before been identified by humans. We just need to figure out how to access these revelations.
That is a discussion for another day.