来源:AI Alignment Forum 2026-07-28 17:30

凝结研究方向:客观性的多样性

of the and to condensation

This is the first part of a survey of various ways that I’d like to see work on the theory of condensation develop. Condensation is a mathematical theory dealing with the organization of descriptions of the world into conceptual parts; some of the existing work on it is presented in the paper (Eisenstat 2025). This sequence will draw from that paper the definitions of random variable models and latent variable models, the notation for indexing subfamilies of variables in such models, and the objectivity theorem, Theorem 6.8. In the condensation paper, Section 3, Ideas, gives an overview of all of this. For more context on condensation, readers can refer to LessWrong posts including (Demski 2025, 2026; Gillen and Chiang 2026; Kirchner 2026). The first two posts of this sequence will introduce some central directions of current work—mainly, the concepts of almost perfect condensation and Kolmogorov (or algorithmic-information) condensation—which will be used in other sections, but the different parts are mostly independent beyond that.

I’d encourage those making a serious effort on any of these problem to contact me for further thoughts and coordination.

Thanks to Kaarel Hänni and James Cook for some of the ideas behind this research program, and to James Cook and Jeremy Gillen for comments on this document.

1. Varieties of objectivity

For the theory of condensation to be useful, we’d like hypotheses that we can assume about latent variable models that are general enough so that such latent variable models exist in cases of interesting random variable models, but narrow enough to imply objectivity-like conclusions—statements of approximate uniqueness along the lines of Theorem 6.8 of (Eisenstat 2025). We’ll refer to hypotheses of this kind as condensation properties. To know that such properties apply to interesting data, we’d like to come up with interesting random variable models satisfying them. (We can also look for interesting string models, which will be the analogue of random variable models when we introduce Kolmogorov condensation in the next post.) Work in these directions will also affect the problems discussed in later sections, which almost all either need to assume the existence of a latent variable model (or a latent string model in the Kolmogorov case) with good condensation properties, or else ask for such latent variable models in some particular context.

1.1. Almost perfect condensation

One strong condensation property is almost perfect condensation. I’ve defined this somewhat differently in different places, and there isn’t yet a good reference for this. We’ll look at a definition here, and we’ll discuss informally the consequences that follow from and motivate the definition, without giving a theorem statement or proof sketch.

Definition 1. Suppose that are the random variables of a random variable model, and that we have an associated latent variable model , with latent variables , where . Then, is an -almost perfect condensation if (1—size) , (2—reconstruction) for all and all , if , then , and (3—Markov) for all -upward closed sets , we have . Further, we say that is -well-separated if (4) for all distinct , if is not a singleton, then at least one of the sets and has cardinality greater than .

Here, we should have an objectivity theorem as follows. If and are two well-separated almost perfect condensations of the same string model, then this should induce a bijection between those latents of and , which should have good objectivity-style mutual determination. More generally, if and are both almost perfect condensations, and is well-separated, then we get a partition of the latents of into some sets, and these sets admit a bijection with the latents of with similar properties to the previous case. Finally, if we only assume that and are almost perfect condensations, without any well-separatedness assumption, we can still conclude that each latent of is determined by some small set of latents of , and vice versa.

While we can prove theorems like these, I’d like to better understand their range of applicability. I’d like to have better developed examples to which these theorems would apply, and to draw from them in order to better formulate similar theorem statements that more fully bring out what these methods can tell us.

We can look at a simple such example of almost perfect condensation. In the following, the latents will be the biases of dice, and the given variables will be observations that give us some information about these dice. That is, we will have a set of latents , each of which is valued in the space of probability measures on , and which are jointly independent. The idea here is that the bias of each die is itself random (we can think of drawing it from a bag of biased dice), but each particular die, with that bias, is rolled multiple times, giving us information about the bias. Now, for the construction, fix some sets , and a relation . Construct a joint distribution on as described. Define by making a roll of die , and setting

We need further latents to account for the idiosyncratic information in , so set and , extending to by setting also . We can check that for appropriate parameters, is an almost perfect condensation of . We want sufficiently large relative to ; not too many such that for each fixed (e.g., we could choose such that there are exactly two such for each , that is, each observation is the sum of two die rolls); and sufficiently evenly distributed, avoiding properties like two different contributing to the same, or nearly the same, set of .

It would be helpful to have examples of almost perfect condensation that are a bit less artificial, and better demonstrate its relevance to constructing interesting latents. This could be in either the Shannon entropy or the Kolmogorov complexity setting. It would also be good if there were better ways to think about the rather long definition given, in which many somewhat arbitrary choices were made. For example, in Definition 1, many things that we want to be small are individually bounded by the same . Any bounds would give some kind of objectivity theorem. Is there a better way of thinking about adapting these upper bounds to the different quantities bounded? This might be a situation where the bounds should be adapted to particular latent variable models being studied. But even then, there could be some general insight to be learned from looking at interesting such examples.

1.2. Beyond almost perfect condensation

Almost perfect condensation has been constructed so as to derive reasonably strong and simple bounds from the objectivity theorem, and more generally aid in the interpretation of its conclusions. However, there may be many other cases where the objectivity theorem can tell us that some latent model is approximately unique enough to be interesting. We will look at some examples and discuss why almost perfect condensation will be too strict to allow the sort of analysis that we’d want, and argue that there’s reasonable hope for other ideas here. But first, we’ll review the objectivity theorem so that we can see how it applies here. We’ll work in the Shannon setting for simplicity, but these problems also exist in the Kolmogorov setting.

The general idea of the objectivity theorem is as follows. Let’s say is a random variable model with latent variables , and and are two associated latent variable models with latent variables and , respectively, where . The objectivity theorem says that for any family , we have

Now, suppose that and are -almost perfect condensations; they satisfy the three conditions in Definition 1. We can use this to pick well in order to simultaneously control and the right-hand side of the inequality. Using condition (1) for , we know that . If we want to control , we really only need to control , and if we remove elements from one at a time, this can take at most -many steps. Recalling the definition

we can see that we can achieve any desired set with a set of size at most . That is, suppose that we want to bound for some desired upward-closed set . We can let

which will have size at most , and will satisfy .

By the definition of in the objectivity theorem, we have . Further, since satisfies the approximate Markov condition, condition (3) from the definition of almost perfect condensation, each of the terms in the sum over is at most . Thus, this inequality implies the simpler statement

From this, we can motivate condition (2) in the definition of almost perfect condensation. Each set on the right-hand-side sum is the complement of a set in that we want to ensure is absent from . Further, we want to be reconstructable from—to have low conditional entropy given—. So, one simple condition to impose is that the intersection of and is sufficiently large. For a fixed , we can take to be the collections of sets such that , which we can regard as an approximation of the upward cone of . This definition of is related to well-separatedness, though it doesn’t quite get us there. Then, the collection of sets that we have to remove consists of sets with large , i.e. sets whose complement has large intersection with . But this is exactly those sets for which (2) promises that the corresponding entropy terms are small. So, we get

We can make two key observations here. First, we didn’t actually need to assume that for every sufficiently large , we have small. It would suffice for this to be true of the particular sets that we need based on the relationship between and . So, we can hope for much more general condensation properties, that let us adapt to different kinds of latent variable models, giving us something much more generally applicable while still implying meaningful objectivity properties. Second, the development here was guided by the goal of something like a bijection between and , giving pairs of corresponding latent variables. While this is the most obviously compelling objectivity story, sets of variable with somewhat weaker relations may still be unique enough to tell us something interesting about how to understand the joint distribution.

I expect that there is at least some, and possibly a lot, of progress to be made in these directions, which would extend the space of distributions that we can say something interesting about. There are plausibly many unanticipated phenomena here, but I’ll give some indication of a few directions to try.

1.2.1. Almost perfect condensation with a relation

In the definition of almost perfect condensation, we imposed the condition that when . This makes some sort of sense, since we can interpret it as saying that has enough elements that are informative about . But often the cardinality of this set is less closely tied to its informativeness.

We can produce examples by thinking about situations where has some geometric structure. Suppose is a latent variable that specifies the approximate shape of some region, and that each observed variable for is the colour of a point. If we measure many points that are close together, we are less informed about the same than we would be if we measured the colour at a set of points that had the same cardinality, but were more geometrically scattered.

This suggests an idea like the following. Instead of asking about those such that , we can ask about those such that every point of is within a distance of some point of . This leads to a more general definition of almost perfect condensation, where we replace the condition by a new relation . We do not assume that is an order; instead, the idea here is that if “almost contains” , in a sense that we are free to choose. We will recover the more restrictive sense of almost perfect condensation above if we define to mean that . Alternatively, we can follow our geometric idea, and say perhaps that if contains a ball of a certain radius in , or equivalently that if every point of is within a certain radius of a point in .

Definition 2. Let be the random variables of a random variable model and the latent variables of an associated latent variable model , where . Further, let be a relation on .

We say that is an -almost perfect condensation if (1—size) , (2—reconstruction) for all and all , either or , and (3—Markov) for all -upward closed sets , we have .

This definition has many of the same good properties as almost perfect condensation, for the same reasons. It is a little more opaque though; it would be easier to understand if we had appropriate examples. The geometric ideas mentioned above seem somewhat promising as a source, though there are some limitative reasons to think that many natural such random variable models, such as joint distributions coming from the colours at different points of a visual scene, will not admit this kind of almost perfect condensation. We will discuss this in the next subsection. However, there may still be some useful geometric idea. We can also put other kinds of structure on the set of observed variables. For example, maybe we have many experimental subjects, and we measure many properties of each subject, giving us a set of observations forming a 2-dimensional table. Note here that the different experimental subjects are being treated as giving us different random variables in our joint distribution, rather than being treated as different i.i.d. draws from the same distribution. This could therefore be a better fit for Kolmogorov condensation, which we will discuss in Section 2.

This also doesn’t say anything about where comes from. For a particular choice of , the objectivity theorem tells us something interesting, but it would be more satisfying if the choice of was in some sense itself approximately unique.

As before, we may do better by adapting within a latent variable model, rather than using the same everywhere, as is done here for simplicity.

1.2.2. Medium-scale effects

The hypothesis of well-separatedness is overly restrictive around what we might think of as medium scale effects, in a sense that we’ll look at next. The parameter—either or —that governs whether latent variables can be reconstructed from observed variables also controls how far apart latent variables must be from each other.

In the case of a visual scene, we might imagine trying to construct a latent variable model using different latent variables for different properties of the objects visible in the scene. For example, we can have a latent with a large contribution set specifying the approximate position and orientation of a tree, with smaller latents filling in details about its shape, e.g. its branches, which are more local and thus have smaller contribution sets. If and are the contribution sets of the tree latent and the branch latent, then on the one hand we’d expect , to be large, since if we don’t make any measurement of the area with the branch, we won’t have a good idea of the approximate shape of the tree. On the other hand, we need this quantity to be small, since we want to rule out the information in the branch latent from being necessary for recovering the tree latent. If we have good separation of scales—e.g., the tree scale is much larger than the branch scale—we can go with the second option, small; we don’t actually need a sample from the branch to have sufficiently good coverage to infer the large-scale shape. But we’d like to say as much as we can even when we don’t have this kind of separation of scales.

While this visual scene idea is described informally, we can come up with particular distributions in which these sort of problems exist, which can be good test cases for refining these ideas. Here I have in mind distributions like those coming from statistical physics, like a random walk or a Ising model near its critical temperature. Whether or not the more specific suggestions above are pointing in a useful direction, it would be good to have any kind of analysis of these distributions. Near the critical temperature, a typical Ising model configuration has a scale-free structure, where it is made up of large domains separated by walls, but then these domains have smaller islands with the same shapes. In order to understand these, physicists introduce coarse-grained variables in renormalization schemes, which are rather like the latent variables that we care about here. One version of this therefore is that we want to know what the objectivity theorem can tell us about the relationship between two different renormalization schemes.

The random walk example is simpler to describe. Given some number , we will construct a joint distribution on . Let be a constant, and define , where each is an independent Rademacher variable, equally likely to be . It’s tempting to take latent variables that are certain averages of the s, but even more simply we can get interesting latent variable models just by making latent variable models into copies of certain s. We would like both these kinds of latent variable models to be related by some appropriate objectivity theorem, but we’ll discuss the latter kind here for simplicity. Make a binary tree of subintervals of , where the root is labelled and for each node whose interval has more than one element, its two children are labelled and for some . We’ll assume that the tree is reasonably shallow—not too much deeper than —but let the tree be otherwise arbitrary. Now, we can construct a latent variable model. Let be the set of labels of the tree, and define by letting , where is the chosen cut point of the interval , if , and otherwise . Because of the Markov property of the random walk, this latent variable model satisfies its own Markov property exactly, whichever tree we use.

Now, for an objectivity theorem, we don’t expect any two such trees to correspond exactly; they have different cut points. However, given any variable at layer of one tree (counting up from at the leaves), we can be reasonably confident of its value given only variables at layers at least , where the exact conditional entropy depends on the choice of . Instead of a one-to-one correspondence of variable, we have a few-to-one function, where each variable in one tree is determined by a small number of variables in the other tree.

This corresponds to what the objectivity theorem should tell us in the case of almost perfect condensation with a relation. While together with well-separatedness properties we can get a bijection on latent variables, the same ideas give us these few-to-one relations without the well-separatedness assumption. Is this the best we can do? Is this the right way to understand this example? What other class of latent variable models for a random walk does this analysis generalize to? Does it generalize further to the critical Ising model, or to other such examples that we can come up with? It would be good to understand these examples better.

1.3. Almost perfect condensation and scoring functions

In Eisenstat (2025) and in Demski (2025), condensation is expressed in terms of some functions, the simple score and the conditioned score, which relate it to compression. Almost perfect condensation and related ideas are helpful for understanding why the objectivity theorem gives good bounds in certain situations, but the hypotheses of almost perfect condensation don’t have a clear relation to compression. It would be interesting if these two perspectives could be unified.

References

Demski, Abram. 2025. “Condensation. LessWrong.”

Demski, Abram. 2026. “Condensation & Relevance. LessWrong.”

Eisenstat, Sam. 2025. “Condensation: A Theory of Concepts.” ODYSSEY 2025 Conference.

Gillen, Jeremy, and Daniel Chiang. 2026. “A Summary of Condensation and Its Relation to Natural Latents. LessWrong.” March 4.

Kirchner, Jan. 2026. “Elementary Condensation. LessWrong.”



Discuss

相关文章推荐

返回首页