<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Theory | Lei Zhang</title><link>https://zhanglei.page/tags/theory/</link><atom:link href="https://zhanglei.page/tags/theory/index.xml" rel="self" type="application/rss+xml"/><description>Theory</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 12:30:00 +0000</lastBuildDate><image><url>https://zhanglei.page/media/icon_hu_102d14ed545eed19.png</url><title>Theory</title><link>https://zhanglei.page/tags/theory/</link></image><item><title>When a Language Model Refers to Itself, What Is "Itself"?</title><link>https://zhanglei.page/blog/self-reference-in-llms/</link><pubDate>Thu, 01 Oct 2026 12:30:00 +0000</pubDate><guid>https://zhanglei.page/blog/self-reference-in-llms/</guid><description>&lt;p&gt;This is the second half of a reading-group talk on self-reference and recursive self-improvement (RSI).
separated three things that all get called a &amp;ldquo;fixed point&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Kleene&amp;rsquo;s second recursion theorem&lt;/strong&gt;: a program obtaining its own description (a quine is the simplest case). Constructed once, no iteration.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Y combinator&lt;/strong&gt;: a function able to call itself. Also constructed once.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The limit of an iteration&lt;/strong&gt; $x_{t+1} = f(x_t)$: a question about dynamics, with convergence only under extra conditions.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here I use that vocabulary on large language models.&lt;/p&gt;
&lt;h2 id="the-paper-that-got-me-started"&gt;The paper that got me started&lt;/h2&gt;
&lt;p&gt;The talk grew out of reading Zhang, Yuan, and Zhang, &lt;em&gt;Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement&lt;/em&gt; (Entropy, 2026). Its thesis, as I read it: sustainable self-improvement needs a functional analogue of von Neumann&amp;rsquo;s complexity threshold for self-reproducing automata. The authors call this &lt;em&gt;introspection&lt;/em&gt;, the system&amp;rsquo;s capacity to simulate its own operations and its intended modifications. Kleene&amp;rsquo;s second recursion theorem shows that introspective programs exist in principle. Current LLMs show partial, &amp;ldquo;quasi-introspective&amp;rdquo; behavior and face structural limits on the way to the real thing.&lt;/p&gt;
&lt;p&gt;I like this paper. It takes a word that is usually used loosely and ties it to a theorem, which is what makes a claim checkable. It is also the reason I went back and relearned the recursion theorems properly. What follows are three questions I kept asking as I read, and where I ended up on each. They are a reader&amp;rsquo;s notes, and I may have misread things. I first read the arXiv preprint; the journal version is more careful in several of the places I discuss, and that is the version I am working from.&lt;/p&gt;
&lt;h2 id="question-1-which-fixed-point-is-doing-the-work"&gt;Question 1: which fixed point is doing the work?&lt;/h2&gt;
&lt;p&gt;On the existence side, the paper is explicit. The recursion theorem is used once, to obtain the first introspective program $S_0$. Everything after that, $S_0 \to S_1 \to S_2 \to \cdots$, is a trajectory produced by running it, where each $S_{i+1}$ is the &lt;em&gt;output&lt;/em&gt; of $S_i$. That is the definitional/dynamical split from part 1, stated cleanly.&lt;/p&gt;
&lt;p&gt;Where I would keep the same care is on the limitation side. One limit discussed for Transformers is that a single forward pass cannot execute fixed-point iterations. Iterating to a fixed point is the third kind in my list, the Banach kind. Kleene&amp;rsquo;s fixed point is the first kind: the s-m-n theorem builds it in one shot, and a quine does not iterate toward its source. So I would treat these as two independent properties of a system:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Can it iterate a map until it converges?&lt;/li&gt;
&lt;li&gt;Does it support Kleene-style self-reference?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A limit on the first does not, by itself, tell us about the second.&lt;/p&gt;
&lt;p&gt;A related point. The paper takes true introspection to require unbounded recursion, with the system modeling its own self-modeling to arbitrary depth. I read that as a chain:&lt;/p&gt;
&lt;p&gt;introspection → self-simulation → recursive self-simulation → unbounded recursive self-simulation&lt;/p&gt;
&lt;p&gt;Each arrow is a real strengthening. Kleene&amp;rsquo;s construction gives a program access to its own description without any tower of &amp;ldquo;simulating how I simulate how I simulate my next action.&amp;rdquo; Whether &lt;em&gt;improvement&lt;/em&gt; needs that tower seems to me an open question, and one worth arguing for directly.&lt;/p&gt;
&lt;h2 id="question-2-how-binding-is-the-single-pass-bound"&gt;Question 2: how binding is the single-pass bound?&lt;/h2&gt;
&lt;p&gt;The bound in question is from Merrill and Sabharwal (2023): a log-precision Transformer, in one forward pass, computes only functions in uniform $\mathsf{TC}^0$. This is correct and important. It is also a statement about &lt;em&gt;one forward pass&lt;/em&gt;, as the paper itself says, noting that chain-of-thought partially evades it.&lt;/p&gt;
&lt;p&gt;How far does the evasion go? The models we call LLMs are autoregressive decoders, so the relevant theory is the one for decoder-only Transformers with intermediate generation. Merrill and Sabharwal (2024) give a ladder indexed by the number of chain-of-thought steps: logarithmically many steps stays within log-space; linearly many adds real power (all regular languages) while staying within context-sensitive languages; polynomially many gives exactly $\mathsf{P}$.&lt;/p&gt;
&lt;p&gt;Beyond that there is a line of Turing-completeness results: Pérez et al. (2021), then Schuurmans (2023) and Schuurmans et al. (2024), then Li and Wang (2025), who show that a Transformer with a fixed number of parameters &lt;em&gt;and&lt;/em&gt; fixed numerical precision is Turing complete, provided the context window can grow as needed and the autoregressive computation can run long enough. The full picture has more moving parts than this summary (hard versus soft attention, the choice of positional encoding), but for the system as it is actually used, the natural reference point is Turing completeness.&lt;/p&gt;
&lt;p&gt;Two caveats, so that I do not overclaim in the other direction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Execution is not discovery.&lt;/strong&gt; Turing completeness says the model can &lt;em&gt;carry out&lt;/em&gt; a program, including a program that refers to its own description. RSI needs something else: given only a goal (prove this theorem, design this algorithm, find this bug), &lt;em&gt;find&lt;/em&gt; the useful program. Turing completeness is silent on that.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What counts as &amp;ldquo;the program&amp;rdquo;?&lt;/strong&gt; Early results build a separate Transformer for each computable function. Qiu et al. (2025) show something more relevant: fix one Transformer, never touch its weights again, and by changing only the prompt you can compute any computable function. In that picture, the prompt is the program code of a Turing-complete model of computation. This matters for the next question.&lt;/p&gt;
&lt;h2 id="question-3-so-can-an-llm-refer-to-itself"&gt;Question 3: so can an LLM refer to itself?&lt;/h2&gt;
&lt;p&gt;If Transformers are Turing complete, and Kleene&amp;rsquo;s second recursion theorem holds in any acceptable programming system, then the answer should be yes.&lt;/p&gt;
&lt;p&gt;It is yes. But it may not be the self-reference you had in mind.&lt;/p&gt;
&lt;p&gt;Think about what a quine running on my laptop can and cannot print. It can print its own source code. It cannot print the transistor layout of the processor, the CPU microcode, the contents of RAM, the operating system kernel, or the binary of the Python interpreter that is running it. The recursion theorem is about program descriptions relative to a fixed interpreter. The interpreter stays in the background and is never asked to read itself.&lt;/p&gt;
&lt;p&gt;Now line this up with an LLM:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Kleene&amp;rsquo;s setting&lt;/th&gt;
&lt;th&gt;LLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;a fixed universal interpreter $\varphi$&lt;/td&gt;
&lt;td&gt;weights plus architecture&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a program code $e$, given as input&lt;/td&gt;
&lt;td&gt;the prompt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;$\varphi_e$, the behavior of code $e$&lt;/td&gt;
&lt;td&gt;the model&amp;rsquo;s behavior on that prompt&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This matches the setting of Qiu et al. Under it, I would expect a statement like the following to hold. I state it informally, and I have not verified every detail of the acceptable-numbering condition (the s-m-n property) for the specific constructions:&lt;/p&gt;
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;Suppose a fixed Transformer $T_\theta$, with sufficiently scalable context and computation, implements an acceptable universal programming system over program descriptions encoded as prompts. Then for every computable transformation $g$ on such descriptions there is a prompt $c^\ast$ with $T_\theta(c^\ast) = g(c^\ast)$. Taking $g$ to be the identity gives a prompt quine, $T_\theta(c^\ast) = c^\ast$.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So &amp;ldquo;Kleene&amp;rsquo;s theorem implies a Transformer can execute self-referential programs&amp;rdquo; is fine. &amp;ldquo;Kleene&amp;rsquo;s theorem implies a Transformer knows its own parameters&amp;rdquo; does not follow. The theorem has no necessary connection to $\theta$ at all.&lt;/p&gt;
&lt;p&gt;Prompt-level self-access is also close to free. In the classical setting a program is not handed its own code; it obtains it indirectly, through the diagonal construction. In a Transformer, attention makes the entire prompt visible to the computation at every step.&lt;/p&gt;
&lt;p&gt;It helps to step back and see the stack. The GPU executes inference code. The inference code reads the weights. The weights act on the prompt. The prompt produces behavior. The weights are a program from the point of view of the inference code, and an interpreter from the point of view of the prompt. &amp;ldquo;Self&amp;rdquo; depends on which level you are standing on. (Li and Wang&amp;rsquo;s main theorem, for instance, is in a one-model-per-task setting, where the weights play the role of the program.)&lt;/p&gt;
&lt;p&gt;The paper is after weight-level self-reference: for a bare LLM, the functional identity lives in the weights, so the reflexive structure would have to be realized over the weights. That is a reasonable thing to want. My point is only about which level the theorem lives on. The formal development identifies a program with its source code, and in the Turing-completeness results for a fixed model, the object playing that role is the prompt. Getting from there to the weights needs its own argument.&lt;/p&gt;
&lt;h2 id="without-kleene-can-parameters-refer-to-themselves"&gt;Without Kleene: can parameters refer to themselves?&lt;/h2&gt;
&lt;p&gt;Set Kleene aside, then, and ask directly. It helps to distinguish program self-reference from parameter self-&lt;em&gt;description&lt;/em&gt; and parameter self-&lt;em&gt;modification&lt;/em&gt;.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Parameter self-description: output your own weights&lt;/td&gt;
&lt;td&gt;Yes, with a cost: neural quines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parameter self-modification: change your own weights&lt;/td&gt;
&lt;td&gt;Yes, within limits: self-referential weight matrices&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;An ordinary Transformer reading its own $\theta$ during a forward pass&lt;/td&gt;
&lt;td&gt;No; $\theta$ is not available to the computation as data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strong closure, $S_{t+1} = \Phi_{S_t}(S_t)$: the rule doing the modifying is itself among the things modified&lt;/td&gt;
&lt;td&gt;Open; Kleene does not supply it, and it may need a stronger form of computational reflection&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="neural-quines"&gt;Neural quines&lt;/h3&gt;
&lt;p&gt;Chang and Lipson (2018) trained a network to output its own weights. The naive version, $f_\Theta(x) = (\theta_1, \ldots, \theta_N)$, fails by counting: the output layer alone would need more parameters than the network has. Their solution is to query by index. Given a coordinate $c$, the network returns the single weight $\theta_c$, and training minimizes&lt;/p&gt;
$$L_{SR}(\Theta) = \sum_c \lVert f_\Theta(c) - \theta_c \rVert^2$$&lt;p&gt;It is supervised learning where the labels are the weights themselves.&lt;/p&gt;
&lt;p&gt;Two things they report are worth knowing. Trained on self-replication alone, the network is strongly attracted to the &lt;em&gt;zero quine&lt;/em&gt;: all weights zero, which satisfies the equation trivially. And when an auxiliary task (MNIST classification) is added, the self-replication loss rises over training instead of converging, because the network prioritizes the task. The quine-plus-task network reaches 90.41% test accuracy, against 96.33% for an identical network trained on the task alone.&lt;/p&gt;
&lt;p&gt;Which kind of fixed point is this? The condition $f_\Theta(c) = \theta_c$ is definitional: it is an equation to be satisfied. But there is no Kleene-style construction that hands you a solution. You have to approach one by iteration, which here means training.&lt;/p&gt;
&lt;p&gt;That contrast is the main thing I took away. In Kleene&amp;rsquo;s setting, self-reference is free. The theorem constructs the program, once, and the program still does whatever $F$ asks of it. In parameter space, self-reference is bought with optimization, it competes with the task for capacity, and the price can be measured.&lt;/p&gt;
&lt;h3 id="self-referential-weight-matrices"&gt;Self-referential weight matrices&lt;/h3&gt;
&lt;p&gt;Irie, Schlag, Csordás, and Schmidhuber (2022) go from description to modification. A single weight matrix $W$ is partitioned into four blocks that produce, respectively, the output $y$, a key $k$ (where to write), a query $q$ (what to write), and a learning rate $\beta$ (how strongly):&lt;/p&gt;
$$y_t,\, k_t,\, q_t,\, \beta_t = W_{t-1}\,\phi(x_t)$$$$\bar v_t = W_{t-1}\,\phi(k_t), \qquad v_t = W_{t-1}\,\phi(q_t)$$$$W_t = W_{t-1} + \sigma(\beta_t)\,(v_t - \bar v_t) \otimes \phi(k_t)$$&lt;p&gt;Eliminating the intermediate variables gives $W_t = H(W_{t-1};\, x_t)$. The current weights really do take part in producing the next weights, and the new weights are used immediately in the next step. That is one level beyond a deep equilibrium model, where only the activations change.&lt;/p&gt;
&lt;p&gt;The rank-1 form is what makes this affordable. Producing a full update $\Delta W \in \mathbb{R}^{m \times n}$ from a hypernetwork would need on the order of $d \cdot mn$ parameters. An outer product avoids that blow-up, and it does so while staying entirely in continuous weights, without serializing anything into discrete symbols.&lt;/p&gt;
&lt;p&gt;The limits are just as instructive. The modification is input-conditioned. The &lt;em&gt;form&lt;/em&gt; of the update, a rank-1 additive write, is fixed by the designers. The initial matrix $W_0$ is found by gradient descent. And the changes live within an episode,&lt;/p&gt;
$$W_0 \xrightarrow{x_1} W_1 \xrightarrow{x_2} W_2 \xrightarrow{x_3} \cdots \xrightarrow{x_t} W_t$$&lt;p&gt;without being carried over permanently to the next one.&lt;/p&gt;
&lt;h2 id="what-stays-outside-the-loop"&gt;What stays outside the loop&lt;/h2&gt;
&lt;p&gt;The same question can be put to a range of systems that have some flavor of $x = f(x)$ or of self-modification: what changes, and what is left untouched?&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;What changes&lt;/th&gt;
&lt;th&gt;What stays outside&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deep equilibrium models&lt;/td&gt;
&lt;td&gt;the activation state $z$, solved from $z^\ast = f_\theta(z^\ast, x)$&lt;/td&gt;
&lt;td&gt;the weights and the solver&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DARTS&lt;/td&gt;
&lt;td&gt;weights $w$ and architecture parameters $\alpha$, by bilevel optimization&lt;/td&gt;
&lt;td&gt;the search space and the evaluation criterion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Graph hypernetworks&lt;/td&gt;
&lt;td&gt;the weights of a &lt;em&gt;target&lt;/em&gt; network&lt;/td&gt;
&lt;td&gt;the generator itself and its training objective&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-referential weight matrix&lt;/td&gt;
&lt;td&gt;fast weights within an episode&lt;/td&gt;
&lt;td&gt;the form of the update rule; $W_0$; persistence across episodes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Neural quine&lt;/td&gt;
&lt;td&gt;weights, to match their own description&lt;/td&gt;
&lt;td&gt;the training objective and the optimizer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Darwin Gödel Machine&lt;/td&gt;
&lt;td&gt;the agent&amp;rsquo;s own code&lt;/td&gt;
&lt;td&gt;the outer archive and the selection procedure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The agent case deserves a remark, because it is where current practice is. A coding agent&amp;rsquo;s loop looks like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# a bounded function call that returns text&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# executed by a genuinely Turing-complete interpreter&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This system is Turing complete for a plain reason: Python is. The LLM has become a bounded subroutine inside a universal machine. Recent work that formalizes agent harnesses as λ-calculi (Liu&amp;rsquo;s $\lambda_A$; the LLMbda calculus of Garby, Gordon, and Sands) makes the structure explicit, and the self-reference in these calculi is mainly the Y-combinator kind, recursive invocation, rather than the second-recursion-theorem kind.&lt;/p&gt;
&lt;p&gt;In every row, whatever is left outside the loop is what the system cannot improve. So the honest version of the RSI equation carries one more argument:&lt;/p&gt;
$$S_{t+1} = \Phi(S_t;\, A)(S_t)$$&lt;p&gt;where $A$ is everything fixed from outside. $A$ is never empty. Hardware, physics, and the Python interpreter are always in it. The useful question is what else is.&lt;/p&gt;
&lt;h2 id="what-this-adds-up-to"&gt;What this adds up to&lt;/h2&gt;
&lt;p&gt;Three clarifications, in the vocabulary of part 1:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Criterion.&lt;/strong&gt; Whether a system &lt;em&gt;can&lt;/em&gt; self-refer is settled by universality, through Kleene&amp;rsquo;s theorem. Whether an iteration converges is a separate matter.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scope.&lt;/strong&gt; The $\mathsf{TC}^0$ bound covers a single forward pass. Autoregressive use is a different regime.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Object.&lt;/strong&gt; What Kleene delivers for a fixed model is self-reference to its instructions, the prompt. Self-reference to weights is a different thing. Where it has been built, it was paid for.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And one thing this does not do. The paper&amp;rsquo;s conclusion, that current systems fall short of sustainable recursive self-improvement, is untouched by any of the above. I have not refuted it, and I have not proved it. It may well be right. My guess is that the road to it runs through the dynamical question more than through self-reference.&lt;/p&gt;
&lt;p&gt;So I end where part 1 began, with $S_{t+1} = \Phi(S_t)(S_t)$ and two questions to put to any proposed system:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Where does $\Phi$ come from, and how much of it is really inside $S$?&lt;/li&gt;
&lt;li&gt;What does the iteration do when it runs: stabilize, diverge, oscillate, or stop somewhere?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Self-reference itself is not the hard part, with two qualifications. It matters which kind you have, because the three kinds come with different prices: Kleene&amp;rsquo;s and Y&amp;rsquo;s are free, constructed directly by a theorem, and the parameter kind is bought by iteration at a measurable cost. And having it is not the end of the story, because something always remains outside.&lt;/p&gt;
&lt;p&gt;My thanks to the authors for a paper worth arguing with, and to the reading group for the questions.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Zhang, J., Yuan, B., &amp;amp; Zhang, Q. (2026). Self-reference in large language models: The introspection threshold for recursive self-improvement. &lt;em&gt;Entropy&lt;/em&gt;, 28(9), 951.&lt;/li&gt;
&lt;li&gt;Merrill, W., &amp;amp; Sabharwal, A. (2023). The parallelism tradeoff: Limitations of log-precision transformers. &lt;em&gt;TACL&lt;/em&gt;, 11, 531–545.&lt;/li&gt;
&lt;li&gt;Merrill, W., &amp;amp; Sabharwal, A. (2024). The expressive power of transformers with chain of thought. &lt;em&gt;ICLR&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Pérez, J., Barceló, P., &amp;amp; Marinkovic, J. (2021). Attention is Turing-complete. &lt;em&gt;JMLR&lt;/em&gt;, 22(75), 1–35.&lt;/li&gt;
&lt;li&gt;Schuurmans, D. (2023). Memory augmented large language models are computationally universal. arXiv:2301.04589.&lt;/li&gt;
&lt;li&gt;Schuurmans, D., Dai, H., &amp;amp; Zanini, F. (2024). Autoregressive large language models are computationally universal. arXiv:2410.03170.&lt;/li&gt;
&lt;li&gt;Li, Q., &amp;amp; Wang, Y. (2025). Constant bit-size transformers are Turing complete. &lt;em&gt;NeurIPS&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Qiu, R., Xu, Z., Bao, W., &amp;amp; Tong, H. (2025). Ask, and it shall be given: On the Turing completeness of prompting. &lt;em&gt;ICLR&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Chang, O., &amp;amp; Lipson, H. (2018). Neural network quine. &lt;em&gt;Artificial Life Conference Proceedings&lt;/em&gt;, 234–241.&lt;/li&gt;
&lt;li&gt;Irie, K., Schlag, I., Csordás, R., &amp;amp; Schmidhuber, J. (2022). A modern self-referential weight matrix that learns to modify itself. &lt;em&gt;ICML&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Zhang, J., Hu, S., Lu, C., Lange, R., &amp;amp; Clune, J. (2026). Darwin Gödel Machine: Open-ended evolution of self-improving agents. &lt;em&gt;ICLR&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Liu, Q. (2026). $\lambda_A$: A typed lambda calculus for LLM agent composition. arXiv:2604.11767.&lt;/li&gt;
&lt;li&gt;Garby, Z., Gordon, A. D., &amp;amp; Sands, D. (2026). The LLMbda calculus: AI agents, conversations, and information flow. arXiv:2602.20064.&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Three Fixed Points Hiding in "Recursive Self-Improvement"</title><link>https://zhanglei.page/blog/three-fixed-points/</link><pubDate>Thu, 01 Oct 2026 12:00:00 +0000</pubDate><guid>https://zhanglei.page/blog/three-fixed-points/</guid><description>&lt;p&gt;I recently gave a talk at a reading group on self-reference and AI. The announced topic was recursive self-improvement (RSI), but most of my time went to a narrower question: when we say a system &amp;ldquo;refers to itself,&amp;rdquo; what exactly is being referred to, and what machinery makes that possible?&lt;/p&gt;
&lt;p&gt;This post is the first half of that talk, the background.
applies it to language models.&lt;/p&gt;
&lt;h2 id="where-this-started-promptprompt"&gt;Where this started: prompt(prompt)&lt;/h2&gt;
&lt;p&gt;A higher-order function takes a function and returns a function:&lt;/p&gt;
$$F(f, \text{data}) \to f'$$&lt;p&gt;People building LLM systems now write the same shape with prompts. A meta-prompt takes a prompt, together with the result of running it, and returns a better prompt:&lt;/p&gt;
$$P(p, \mathrm{run}(p)) \to p'$$&lt;p&gt;In both cases, something that is normally &lt;em&gt;executed&lt;/em&gt; becomes something that can also be passed around, inspected, and transformed. Lisp programmers call this code-as-data:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-lisp" data-lang="lisp"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;; code: evaluate it, get 3&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="o"&gt;&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;; data: keep the expression itself&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;For functions, we know what happens when you close the loop and feed a function to itself: you get recursion, and the Y combinator is the classical way to build it. For prompts I did not have an equally crisp answer, and that is what sent me back to the theory.&lt;/p&gt;
&lt;h2 id="the-shape-of-recursive-self-improvement"&gt;The shape of recursive self-improvement&lt;/h2&gt;
&lt;p&gt;Ordinary improvement looks like this:&lt;/p&gt;
$$S_{t+1} = I(S_t)$$&lt;p&gt;$S$ is the state of the system (weights, code, memory, tool configuration) and $I$ is one improvement step, chosen by someone outside the system: an engineer, an optimizer, a training script.&lt;/p&gt;
&lt;p&gt;Recursive self-improvement asks for more. The improvement step is itself read off the current state:&lt;/p&gt;
$$I_t = \Phi(S_t), \qquad S_{t+1} = \Phi(S_t)(S_t)$$&lt;p&gt;$S_t$ appears twice: once as the thing that produces the rule, and once as the thing the rule is applied to. This is the same shape as the λ-term $\lambda f.\, f\, f$, which takes a function and applies it to itself.&lt;/p&gt;
&lt;p&gt;Equations of this shape invite the phrase &amp;ldquo;fixed point.&amp;rdquo; The trouble is that at least three different mathematical objects go by that name, and they answer different questions.&lt;/p&gt;
&lt;h2 id="fixed-point-one-a-program-that-has-its-own-description"&gt;Fixed point one: a program that has its own description&lt;/h2&gt;
&lt;p&gt;Here is a Python quine, a program that prints its own source:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;s=&lt;/span&gt;&lt;span class="si"&gt;%r&lt;/span&gt;&lt;span class="s1"&gt;;print(s&lt;/span&gt;&lt;span class="si"&gt;%%&lt;/span&gt;&lt;span class="s1"&gt;s)&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="nb"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Quines are not a trick of Python. They exist in every reasonable programming language because of &lt;strong&gt;Kleene&amp;rsquo;s second recursion theorem&lt;/strong&gt;: for every total computable transformation $F$ on program codes, there is a program $e$ such that&lt;/p&gt;
$$\varphi_e = \varphi_{F(e)}$$&lt;p&gt;In words: whatever you plan to do to a program&amp;rsquo;s text, some program already behaves as if that had been done to &lt;em&gt;its own&lt;/em&gt; text. The practical corollary is that any program can be written as though it had access to its own description and could compute with it. A quine is the special case where the computation is &amp;ldquo;print it.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Two features of this theorem matter later.&lt;/p&gt;
&lt;p&gt;First, it is about &lt;em&gt;descriptions&lt;/em&gt;. The thing being referred to is the code $e$, the text of the program. I will call this &lt;strong&gt;representational self-access&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Second, the proof is a construction. The s-m-n theorem builds $e$ explicitly, once. Nothing is run until it converges. A quine does not approximate its source over many rounds; it prints it in a single execution.&lt;/p&gt;
&lt;h2 id="fixed-point-two-a-function-defined-in-terms-of-itself"&gt;Fixed point two: a function defined in terms of itself&lt;/h2&gt;
&lt;p&gt;The λ-calculus has three kinds of expressions (variables, abstractions $\lambda x.\, e$, and applications $e_1\, e_2$) and one rule of computation, β-reduction:&lt;/p&gt;
$$(\lambda x.\, e)\; a \;\to\; e[x := a]$$&lt;p&gt;There are no names. That creates a problem for recursion. In Python, &lt;code&gt;fact&lt;/code&gt; can call &lt;code&gt;fact&lt;/code&gt; because the body can mention the name. An anonymous function has nothing to call:&lt;/p&gt;
$$\lambda n.\ \text{if } n = 0 \text{ then } 1 \text{ else } n \cdot \text{???}(n-1)$$&lt;p&gt;The way out takes three steps.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 1: turn the recursive call into a parameter.&lt;/strong&gt; Define a one-step operator that is not itself recursive:&lt;/p&gt;
$$G \triangleq \lambda f.\, \lambda n.\ \text{if } n = 0 \text{ then } 1 \text{ else } n \cdot f(n-1)$$&lt;p&gt;$G$ says: give me something that handles smaller inputs, and I will handle one more step. Nowhere does $G$ call $G$.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 2: ask for a fixed point of $G$.&lt;/strong&gt; If some $f$ satisfies $f = G(f)$, then substituting the definition of $G$ gives&lt;/p&gt;
$$f = \lambda n.\ \text{if } n = 0 \text{ then } 1 \text{ else } n \cdot f(n-1)$$&lt;p&gt;which is the recursive definition we wanted.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 3: build the fixed point.&lt;/strong&gt; The Y combinator is&lt;/p&gt;
$$Y \equiv \lambda f.\, (\lambda x.\, f\,(x\, x))\,(\lambda x.\, f\,(x\, x))$$&lt;p&gt;and for any $g$ it satisfies $Y g = g\,(Y g)$. So $\mathit{fact} \triangleq Y\, G$. Running it peels off one layer of $G$ per reduction:&lt;/p&gt;
$$Y\,G \;\to\; G\,(Y\,G) \;\to\; G\,(G\,(Y\,G)) \;\to\; \cdots$$&lt;p&gt;and each layer performs one multiplication: $\mathit{fact}(3) \to 3 \cdot \mathit{fact}(2) \to 3 \cdot 2 \cdot \mathit{fact}(1) \to \cdots \to 6$.&lt;/p&gt;
&lt;p&gt;The engine is the self-application $x\, x$ inside $Y$: two identical copies of $\lambda x.\, f\,(x\,x)$, one fed to the other.&lt;/p&gt;
&lt;p&gt;What this shows is that inside the pure λ-calculus, using only abstraction and application, recursion can be &lt;em&gt;constructed&lt;/em&gt;. The language does not need to supply a &lt;code&gt;rec&lt;/code&gt; primitive, function names, mutable state, or a call stack.&lt;/p&gt;
&lt;p&gt;It is also a different kind of self-reference from the quine. $Y$ never gives a function its own source code. It gives the function &lt;em&gt;itself&lt;/em&gt;, as a value to call. I will call this &lt;strong&gt;behavioral self-reference&lt;/strong&gt;. Its theoretical counterpart is Kleene&amp;rsquo;s &lt;em&gt;first&lt;/em&gt; recursion theorem: every computable functional $\Phi$ on partial functions has a least fixed point $g$ with $\Phi(g) = g$. This is the semantic basis of recursive definitions, and in Scott&amp;rsquo;s denotational semantics $Y$ is interpreted as exactly the least-fixed-point operator.&lt;/p&gt;
&lt;p&gt;So Kleene&amp;rsquo;s two theorems split along a line that is easy to miss. The first works at the level of functions (extensional): what the program computes. The second works at the level of descriptions (intensional): what the program&amp;rsquo;s text is.&lt;/p&gt;
&lt;h2 id="fixed-point-three-where-an-iteration-settles"&gt;Fixed point three: where an iteration settles&lt;/h2&gt;
&lt;p&gt;Start from $x_0 = 1$ and repeatedly press the cosine key on a calculator:&lt;/p&gt;
$$1 \;\to\; 0.5403 \;\to\; 0.8576 \;\to\; 0.6543 \;\to\; \cdots \;\to\; 0.7390851332\ldots$$&lt;p&gt;The limit satisfies $x^* = \cos(x^*)$. Notice the two roles here. The fixed point is &lt;em&gt;defined&lt;/em&gt; by the equation. The iteration is one &lt;em&gt;method&lt;/em&gt; for finding it.&lt;/p&gt;
&lt;p&gt;These are two different statements:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Algebraic:&lt;/strong&gt; does $x^* = f(x^*)$ have a solution?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dynamical:&lt;/strong&gt; does $x_{t+1} = f(x_t)$, started from a given $x_0$, converge to it?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Banach fixed-point theorem connects them under an extra condition: a contraction mapping has a unique fixed point, and iteration from any starting point converges to it. Cosine is a contraction on $[0, 1]$ because $|\cos'(x)| = |\sin x| \lt 1$ there. Without the condition, the two statements come apart. Take $f(x) = 2x$. Zero is a fixed point, but starting from $x_0 = 1$ the iteration goes $1 \to 2 \to 4 \to 8 \to \cdots$ and never gets there.&lt;/p&gt;
&lt;h2 id="keeping-them-apart"&gt;Keeping them apart&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Kleene&amp;rsquo;s second theorem&lt;/th&gt;
&lt;th&gt;Y combinator (Kleene&amp;rsquo;s first)&lt;/th&gt;
&lt;th&gt;Banach-style iteration&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What is fixed&lt;/td&gt;
&lt;td&gt;a program code $e$&lt;/td&gt;
&lt;td&gt;a function&lt;/td&gt;
&lt;td&gt;a point in a metric space&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kind of self-reference&lt;/td&gt;
&lt;td&gt;representational: access to one&amp;rsquo;s own description&lt;/td&gt;
&lt;td&gt;behavioral: the ability to call oneself&lt;/td&gt;
&lt;td&gt;none; it is a statement about dynamics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;How you get it&lt;/td&gt;
&lt;td&gt;constructed once&lt;/td&gt;
&lt;td&gt;constructed once, unfolds when run&lt;/td&gt;
&lt;td&gt;iterate, with convergence only under extra conditions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Question it answers&lt;/td&gt;
&lt;td&gt;can a program obtain and compute on its own text?&lt;/td&gt;
&lt;td&gt;can a function be defined in terms of itself?&lt;/td&gt;
&lt;td&gt;where does repeated application end up?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Now go back to $S_{t+1} = \Phi(S_t)(S_t)$. There are two separate things to ask about it.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Can $\Phi$ be written down at all?&lt;/strong&gt; Does the system have the means to refer to, describe, and rewrite itself? This is a definitional question, and Kleene&amp;rsquo;s theorems and the Y combinator speak to it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What happens when it runs?&lt;/strong&gt; After $\Phi$ has acted many times, does $S_t$ stabilize, diverge, oscillate, or park at some equilibrium? This is a question about the whole trajectory, and it is the Banach-style question.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;An answer to one does not transfer to the other. A system can have perfect access to its own description and a trajectory that goes nowhere useful. And a system can have a beautifully convergent iteration with no self-reference in it at all: cosine does not know it is cosine.&lt;/p&gt;
&lt;p&gt;My own view, after working through this, is that the first question is the easier of the two. In
I look at what each one says about language models.&lt;/p&gt;</description></item></channel></rss>