-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathindex.html
More file actions
644 lines (619 loc) · 59.4 KB
/
Copy pathindex.html
File metadata and controls
644 lines (619 loc) · 59.4 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Textual Planning with Explicit Latent Transitions</title>
<meta name="generator" content="PEO 1.10.7">
<meta name="description" content="EmbedPlan: a transition model for LLM planning built on frozen text embeddings, and a map of how far it generalizes across 9 classical planning domains.">
<meta name="author" content="Eliezer Shlomi, Ido Levy, Eilam Shapira, Michael Katz, Guy Uziel, Segev Shlomov, Nir Mashkif, Roi Reichart, Sarah Keren">
<meta name="robots" content="index, follow, max-image-preview:large, max-snippet:-1">
<link rel="canonical" href="https://embedplan.github.io/">
<link rel="alternate" type="text/plain" href="llms.txt" title="llms.txt">
<meta name="keywords" content="EmbedPlan, LLM planning, transition model, world model, frozen text embeddings, classical planning, PDDL, ACPBench, generalization">
<meta name="google-site-verification" content="wtOqeB5giQGwmDSPrehhJVguSECuzVMXL7mrX5oypWg">
<!-- Google Scholar (Highwire Press) -->
<meta name="citation_title" content="Textual Planning with Explicit Latent Transitions">
<meta name="citation_author" content="Shlomi, Eliezer">
<meta name="citation_author_institution" content="Technion – Israel Institute of Technology">
<meta name="citation_author" content="Levy, Ido">
<meta name="citation_author_institution" content="IBM">
<meta name="citation_author" content="Shapira, Eilam">
<meta name="citation_author_institution" content="Technion – Israel Institute of Technology">
<meta name="citation_author" content="Katz, Michael">
<meta name="citation_author_institution" content="IBM">
<meta name="citation_author" content="Uziel, Guy">
<meta name="citation_author_institution" content="IBM">
<meta name="citation_author" content="Shlomov, Segev">
<meta name="citation_author_institution" content="IBM">
<meta name="citation_author" content="Mashkif, Nir">
<meta name="citation_author_institution" content="IBM">
<meta name="citation_author" content="Reichart, Roi">
<meta name="citation_author_institution" content="Technion – Israel Institute of Technology">
<meta name="citation_author" content="Keren, Sarah">
<meta name="citation_author_institution" content="Technion – Israel Institute of Technology">
<meta name="citation_publication_date" content="2026/02/04">
<meta name="citation_arxiv_id" content="2602.04557">
<meta name="citation_pdf_url" content="https://arxiv.org/pdf/2602.04557">
<meta name="citation_language" content="en">
<meta name="citation_abstract_html_url" content="https://embedplan.github.io/">
<meta name="citation_keywords" content="EmbedPlan; LLM planning; transition model; world model; frozen text embeddings; classical planning; PDDL; ACPBench; generalization">
<!-- Open Graph and Twitter cards -->
<meta property="og:type" content="article">
<meta property="og:site_name" content="EmbedPlan">
<meta property="og:title" content="Textual Planning with Explicit Latent Transitions">
<meta property="og:description" content="When a large language model serves as a planner's transition model, every next state is generated token by token, which makes searching over many possible futures slow and expensive. EmbedPlan embeds the state and action with a frozen LLM, predicts the next-state embedding with a lightweight learned network, and returns the closest real state, evaluated on 9 classical planning domains under six settings that hold out progressively more of the data.">
<meta property="og:url" content="https://embedplan.github.io/">
<meta property="og:image" content="https://embedplan.github.io/assets/social.png">
<meta property="og:image:width" content="1280">
<meta property="og:image:height" content="640">
<meta property="og:image:alt" content="Textual Planning with Explicit Latent Transitions: the EmbedPlan architecture, where a frozen LLM encoder embeds a state and an action and a learned network predicts the next state in a latent space.">
<meta property="og:locale" content="en_US">
<meta name="twitter:card" content="summary_large_image">
<meta name="twitter:title" content="Textual Planning with Explicit Latent Transitions">
<meta name="twitter:description" content="When a large language model serves as a planner's transition model, every next state is generated token by token, which makes searching over many possible futures slow and expensive. EmbedPlan embeds the state and action with a frozen LLM, predicts the next-state embedding with a lightweight learned network, and returns the closest real state, evaluated on 9 classical planning domains under six settings that hold out progressively more of the data.">
<meta name="twitter:image" content="https://embedplan.github.io/assets/social.png">
<meta name="twitter:image:alt" content="Textual Planning with Explicit Latent Transitions: the EmbedPlan architecture, where a frozen LLM encoder embeds a state and an action and a learned network predicts the next state in a latent space.">
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "ScholarlyArticle",
"@id": "https://embedplan.github.io/#paper",
"headline": "Textual Planning with Explicit Latent Transitions",
"name": "Textual Planning with Explicit Latent Transitions",
"abstract": "Planning requires a transition model that predicts how each action changes the current state. When a large language model (LLM) plays this role, every next state is generated token by token, which makes searching over many possible futures slow and expensive. Existing alternatives either still query an LLM at every step or require a symbolic model of the domain. We propose EmbedPlan, a transition model built on frozen text embeddings: it embeds natural language descriptions of the state and the action with a frozen LLM, predicts the embedding of the next state with a lightweight learned network, and returns the closest real state. Because this network can be trained on top of any encoder, EmbedPlan also provides a controlled way to compare text representations for learning transitions. We evaluate it on 9 classical planning domains, under six settings that hold out progressively more of the data, from transitions to entire domains, and against baselines ranging from predicting no change to learning symbolic action rules. On planning problems seen during training, EmbedPlan almost always ranks the true next state among its top five guesses, still does so for most queries even when every observed state is a candidate, and retains 92–99% of its single-step accuracy when predicting several steps ahead from its own outputs. Given the same candidate states as GPT-5.4, it picks the true next state more often while taking about 0.17 ms per transition with cached embeddings. Accuracy is lower on unseen problems and near chance on unseen domains, and the controlled comparison traces this limit to the state representation rather than to the learned transition.",
"description": "When a large language model serves as a planner's transition model, every next state is generated token by token, which makes searching over many possible futures slow and expensive. EmbedPlan embeds the state and action with a frozen LLM, predicts the next-state embedding with a lightweight learned network, and returns the closest real state, evaluated on 9 classical planning domains under six settings that hold out progressively more of the data.",
"keywords": [
"EmbedPlan",
"LLM planning",
"transition model",
"world model",
"frozen text embeddings",
"classical planning",
"PDDL",
"ACPBench",
"generalization"
],
"author": [
{
"@type": "Person",
"name": "Eliezer Shlomi",
"givenName": "Eliezer",
"familyName": "Shlomi",
"affiliation": [
{
"@type": "Organization",
"name": "Technion – Israel Institute of Technology"
}
]
},
{
"@type": "Person",
"name": "Ido Levy",
"givenName": "Ido",
"familyName": "Levy",
"affiliation": [
{
"@type": "Organization",
"name": "IBM"
}
],
"url": "https://orcid.org/0009-0005-7400-3452",
"sameAs": [
"https://orcid.org/0009-0005-7400-3452",
"https://github.com/dolev31",
"https://scholar.google.com/citations?user=Ok_7M80AAAAJ",
"https://www.linkedin.com/in/idolevi31/"
]
},
{
"@type": "Person",
"name": "Eilam Shapira",
"givenName": "Eilam",
"familyName": "Shapira",
"affiliation": [
{
"@type": "Organization",
"name": "Technion – Israel Institute of Technology"
}
]
},
{
"@type": "Person",
"name": "Michael Katz",
"givenName": "Michael",
"familyName": "Katz",
"affiliation": [
{
"@type": "Organization",
"name": "IBM"
}
]
},
{
"@type": "Person",
"name": "Guy Uziel",
"givenName": "Guy",
"familyName": "Uziel",
"affiliation": [
{
"@type": "Organization",
"name": "IBM"
}
]
},
{
"@type": "Person",
"name": "Segev Shlomov",
"givenName": "Segev",
"familyName": "Shlomov",
"affiliation": [
{
"@type": "Organization",
"name": "IBM"
}
]
},
{
"@type": "Person",
"name": "Nir Mashkif",
"givenName": "Nir",
"familyName": "Mashkif",
"affiliation": [
{
"@type": "Organization",
"name": "IBM"
}
]
},
{
"@type": "Person",
"name": "Roi Reichart",
"givenName": "Roi",
"familyName": "Reichart",
"affiliation": [
{
"@type": "Organization",
"name": "Technion – Israel Institute of Technology"
}
]
},
{
"@type": "Person",
"name": "Sarah Keren",
"givenName": "Sarah",
"familyName": "Keren",
"affiliation": [
{
"@type": "Organization",
"name": "Technion – Israel Institute of Technology"
}
]
}
],
"inLanguage": "en",
"isAccessibleForFree": true,
"url": "https://embedplan.github.io/",
"mainEntityOfPage": "https://embedplan.github.io/",
"image": "https://embedplan.github.io/assets/social.png",
"about": [
{
"@type": "DefinedTerm",
"name": "EmbedPlan",
"description": "A transition model built on frozen text embeddings: it embeds natural language descriptions of the state and the action with a frozen LLM, predicts the embedding of the next state with a lightweight learned network, and returns the closest real state."
},
{
"@type": "DefinedTerm",
"name": "Transition model",
"description": "A model that predicts how each action changes the current state, which planning requires."
},
{
"@type": "DefinedTerm",
"name": "Hit@k",
"description": "The fraction of queries whose true next state ranks in the top k of a candidate pool. Unless stated otherwise the pool holds 128 states, the true successor and 127 distractors, so chance is 3.9% for Hit@5 and 0.8% for Hit@1."
},
{
"@type": "DefinedTerm",
"name": "Head",
"description": "The learned state and action projection heads and the transition network, together. Only the head is trained, and the LLM encoder stays frozen."
},
{
"@type": "DefinedTerm",
"name": "Snapping",
"description": "In multi-step rollout, each state prediction is replaced by the nearest real state the model retrieves, right or wrong, and that state becomes the next input."
},
{
"@type": "DefinedTerm",
"name": "Domain and problem",
"description": "A domain fixes the action schemas, and a problem fixes the objects, the initial state, and the goal."
},
{
"@type": "DefinedTerm",
"name": "Grounded action",
"description": "An instance of a lifted action schema with its objects filled in, such as pick-up(C) for the schema pick-up(?x)."
},
{
"@type": "DefinedTerm",
"name": "Interpolation",
"description": "The protocol that trains on 80% of a domain's transitions and tests on the other 20%, held out at random, so test transitions come from observed problems, with distractors from the whole domain."
},
{
"@type": "DefinedTerm",
"name": "Plan-Variant",
"description": "The protocol that trains on some optimal plans and tests on other optimal plans of the same problems, with distractors from successors along alternative optimal plans, the hardest distractors the paper uses."
},
{
"@type": "DefinedTerm",
"name": "Extrapolation",
"description": "The protocol that trains on about 80% of a domain's problems and tests on the other 20%, so test transitions come from unseen problems, with distractors from the query's own problem."
},
{
"@type": "DefinedTerm",
"name": "Multi-Domain",
"description": "Extrapolation with one model for all nine domains at once."
},
{
"@type": "DefinedTerm",
"name": "Cross-Domain",
"description": "The protocol that trains on one domain and tests on another, with distractors from the target domain."
},
{
"@type": "DefinedTerm",
"name": "Leave-One-Out",
"description": "The protocol that trains on eight domains and tests on the ninth, with distractors from the target domain."
},
{
"@type": "DefinedTerm",
"name": "Identity and Offset",
"description": "Two non-learned floors that need no training. Identity predicts the current state's embedding, and Offset adds the mean training displacement of the ground action or of its schema."
},
{
"@type": "DefinedTerm",
"name": "Lifted STRIPS induction",
"description": "A reference method that infers each action schema's add and delete effects from the symbolic facts of each state. It serves as an oracle upper bound."
},
{
"@type": "DefinedTerm",
"name": "Action disambiguation",
"description": "A training objective that separates the effects of different actions on the same state: the true action must yield a prediction closer to the next state than up to 50 alternative actions applied to the same state."
}
],
"identifier": "arXiv:2602.04557",
"sameAs": [
"https://arxiv.org/abs/2602.04557"
],
"datePublished": "2026-02-04"
},
{
"@type": "SoftwareSourceCode",
"codeRepository": "https://github.com/embedplan/EmbedPlan",
"@id": "https://github.com/embedplan/EmbedPlan",
"name": "Code",
"url": "https://github.com/embedplan/EmbedPlan",
"subjectOf": {
"@id": "https://embedplan.github.io/#paper"
}
},
{
"@type": "FAQPage",
"@id": "https://embedplan.github.io/#questions",
"about": {
"@id": "https://embedplan.github.io/#paper"
},
"mainEntity": [
{
"@type": "Question",
"name": "What is EmbedPlan?",
"acceptedAnswer": {
"@type": "Answer",
"text": "EmbedPlan is a transition model for planning built on frozen text embeddings. Given a state and an action described in natural language, a frozen LLM embeds both, a lightweight learned network predicts the embedding of the next state, and the closest real state is returned. It is a transition component, not a complete planner: search composes many transitions, and the paper studies the single transition that search queries repeatedly."
}
},
{
"@type": "Question",
"name": "Why not let an LLM generate the next state?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Because generating every next state token by token makes search slow and expensive. When an LLM serves as the world model of a planner, each next state is generated with a full forward pass per token, which makes multi-step lookahead and rollout-based search prohibitively expensive in latency and cost. EmbedPlan leaves the semantic encoding to the frozen LLM and learns only the cheap dynamics, so each prediction is one learned vector step followed by retrieval."
}
},
{
"@type": "Question",
"name": "Is EmbedPlan better than asking an LLM like GPT-5.4?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Within observed problems, at picking the true next state, yes. Given the identical ranking task over the same 128 candidates, GPT-5.4 selects the true successor for 44% of Ferry and 92% of Logistics queries, against 99.0% and 96.3% for EmbedPlan. On Logistics the margin is small next to the spread across the paper's runs (91.2–96.3%). The comparison is matched in task, not in information: EmbedPlan has observed other transitions of the same problems, whereas the LLM is zero-shot, and on unseen problems its Ferry Hit@1 falls to 12.0%."
}
},
{
"@type": "Question",
"name": "How fast is EmbedPlan?",
"acceptedAnswer": {
"@type": "Answer",
"text": "About 0.17 ms per transition with cached state embeddings, 11,215× faster than generating the next state through an LLM API (1.9 s per transition). End to end, including encoding the query texts with BGE-M3, it takes 18.6 ms, 101× faster. Training is a one-time cost per domain, an embedding pass plus about 15 minutes of training, which breaks even against autoregressive generation after about 4,300 transitions with Llama-3.3-70B, so it pays off for a domain that is planned in repeatedly."
}
},
{
"@type": "Question",
"name": "What does Hit@5 measure, and what is chance?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Hit@5 is the fraction of queries whose true next state ranks in the top 5 of a candidate pool. Unless stated otherwise the pool holds 128 states, the true successor and 127 distractors, so chance is 3.9% for Hit@5 and 0.8% for Hit@1. The paper emphasizes Hit@5 because it is robust to the pool size, and a top-5 list is also useful where a verifier can filter candidates."
}
},
{
"@type": "Question",
"name": "How well does EmbedPlan generalize to new problems and new domains?",
"acceptedAnswer": {
"@type": "Answer",
"text": "It depends on what the model has seen. With the Llama-3.3-70B encoder, Hit@5 is 99.7% on held-out transitions of observed problems, 54.6% on unseen problems of an observed domain, and 6.6% on unseen domains, near the 3.9% chance level. Training on the other eight domains (Leave-One-Out) reaches 9.2%. One model trained on all nine domains reaches 37.2% on their unseen problems."
}
},
{
"@type": "Question",
"name": "What limits transfer to unseen problems and domains?",
"acceptedAnswer": {
"@type": "Answer",
"text": "The paper traces the limit to the frozen state representation rather than to the learned transition. With the head, data and pool fixed, character 3–5 grams of the same text reach 79.4% Hit@5 on unseen problems of three domains, against 56.8% with Llama-3.3-70B embeddings. On unseen domains, the paper sees two plausible causes that its experiments do not separate: frozen embeddings organize states by surface form rather than by structural role, and the learned action semantics are domain-specific."
}
},
{
"@type": "Question",
"name": "Does EmbedPlan stay accurate over several steps?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Yes, over short horizons, when each prediction is snapped to a real state. If each predicted state is replaced by the nearest real state before the next step, rollout keeps 92–99% of the accuracy obtained with the true state at each step, on Ferry, Logistics and Blocksworld within observed problems. Without snapping, the predicted embedding drifts off the real states. The test trajectories are short (mean 2.2 steps on Ferry), so this establishes stability over short horizons only."
}
},
{
"@type": "Question",
"name": "Does the true next state have to be among the candidates?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Yes. Retrieval assumes a candidate pool that contains the true successor. The paper measures how accuracy degrades as the pool grows to every state observed in a domain. On Ferry (46,205 states), Hit@5 falls from 100.0% to 86.9%, and on Logistics (13,373 states) from 99.9% to 74.9%, while Hit@1 falls to 35.2% and 32.6%. The true successor thus stays in a short list against the full pool, though not reliably at rank one."
}
},
{
"@type": "Question",
"name": "Which domains and encoders does the paper use?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Nine classical PDDL domains from ACPBench, with states rendered as natural language: Blocksworld, Depot, Ferry, Floortile, Goldminer, Grid, Logistics, Rovers and Satellite, 2.97M transitions over 67 problems. The frozen encoders are MPNet (all-mpnet-base-v2), BGE-M3, Qwen2.5-7B and Llama-3.3-70B. Larger encoders extrapolate better, from 26.8% Hit@5 on unseen problems with MPNet to 54.6% with Llama-3.3-70B, but none closes the gap."
}
},
{
"@type": "Question",
"name": "When is EmbedPlan worth using?",
"acceptedAnswer": {
"@type": "Answer",
"text": "When states and actions are available as text but no symbolic model is, so that the alternative is querying an LLM at every expansion. Action selection and search remain external to EmbedPlan. It must be trained per domain, with a one-time embedding pass plus about 15 minutes of training, so it pays off for a domain that is planned in repeatedly. The paper's classical domains have exact simulators, which makes them a controlled testbed."
}
},
{
"@type": "Question",
"name": "When EmbedPlan misses, how far off is its top guess?",
"acceptedAnswer": {
"@type": "Answer",
"text": "By a few facts. When the true successor misses the top five on unseen problems, the top-ranked state typically differs from it in one to three facts (median 2), usually involving object locations or holdings within the same problem instance. These errors suggest the model captures coarse transition structure (correct problem context, approximate state region) but struggles to resolve fine-grained predicate changes, particularly when multiple objects undergo similar transformations."
}
},
{
"@type": "Question",
"name": "Would fine-tuning the encoder help?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Adapting the encoder already helps. Low-rank adaptation (LoRA) of BGE-M3 raises full-pool Hit@1 on unseen problems from 3.8% to 9.3% when the head is first trained with the encoder frozen, and to 6.2% with a cold joint start. Larger frozen encoders also extrapolate better. Because the encoder is a plug-in, the paper names a concrete target the framework can evaluate directly: an object-aware, order-invariant state encoder."
}
},
{
"@type": "Question",
"name": "Can EmbedPlan tell apart the effects of different actions?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Within observed problems it mostly can, and on unseen problems far less often. The paper applies every action applicable in a state and ranks the actions by how close their predictions come to the true next state. Acc@k is the fraction of queries whose true action ranks in the top k. Mean Acc@5 is 85.3% within observed problems (Acc@1 32.0%) and 16.2% on unseen problems, where training without the action loss reaches only 4.8%."
}
},
{
"@type": "Question",
"name": "How often does EmbedPlan get a whole plan right?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Under Plan-Variant, where the plans are unseen but their problems are observed, 22.6% of plans are correct at every step and mean per-step Hit@5 is 51.2%. Each step is ranked against successors along alternative optimal plans, the hardest distractors the paper uses: they are reachable and often differ from the target in a single fact. Plans of unseen problems fare far worse, at 10.1% per step and 3.3% of plans."
}
}
]
}
]
}
</script>
<link rel="icon" href="favicon.svg" type="image/svg+xml">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500&family=IBM+Plex+Sans:ital,wght@0,400;0,500;0,600;1,400&family=IBM+Plex+Serif:wght@500;600&display=optional" media="print" onload="this.media='all'">
<noscript><link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500&family=IBM+Plex+Sans:ital,wght@0,400;0,500;0,600;1,400&family=IBM+Plex+Serif:wght@500;600&display=optional"></noscript>
<style>
:root {
color-scheme: light;
--bg: #FFFFFF;
--panel: #F3F6FA;
--ink: #111C2A;
--ink-2: #37465A;
--muted: #5C6A7C;
--line: #E0E6EE;
--accent: #1B5EA8;
--accent-soft: #E4ECF5;
--before: #B0521A;
--serif: "IBM Plex Serif", Georgia, "Times New Roman", serif;
--sans: "IBM Plex Sans", system-ui, -apple-system, "Segoe UI", Helvetica, Arial, sans-serif;
--mono: "IBM Plex Mono", ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
}
@media (prefers-color-scheme: dark) {
:root {
color-scheme: dark;
--bg: #0D131A;
--panel: #141D27;
--ink: #E7EDF4;
--ink-2: #C3CDD9;
--muted: #97A4B4;
--line: #25313E;
--accent: #82A6CF;
--accent-soft: #10263E;
--before: #F09A62;
}
}
* { box-sizing: border-box; }
html { -webkit-text-size-adjust: 100%; }
body { margin: 0; background: var(--bg); color: var(--ink); font: 400 17px/1.65 var(--sans); -webkit-font-smoothing: antialiased; }
a { color: var(--accent); text-underline-offset: 3px; text-decoration-thickness: 1px; }
a:focus-visible, button:focus-visible { outline: 2px solid var(--accent); outline-offset: 3px; border-radius: 4px; }
img { max-width: 100%; height: auto; display: block; }
sup { font-size: .68em; line-height: 0; }
code { font-family: var(--mono); font-size: .86em; background: var(--panel); padding: .1em .35em; border-radius: 4px; }
.wrap { max-width: 1040px; margin: 0 auto; padding-inline: 20px; }
.text { max-width: 720px; margin-inline: auto; }
header.hero { padding-block: 64px 36px; text-align: center; }
.eyebrow { font: 500 12.5px/1 var(--mono); letter-spacing: .12em; text-transform: uppercase; color: var(--muted); margin: 0 0 22px; }
h1 { font: 600 clamp(28px, 4.6vw, 44px)/1.16 var(--serif); letter-spacing: -.012em; margin: 0 auto 22px; max-width: 900px; text-wrap: balance; }
h1 .line2 { display: block; color: var(--accent); }
.authors { font-size: 18px; margin: 0; }
.authors a { color: inherit; text-decoration-color: var(--accent); }
.authors .sep { color: var(--muted); padding-left: .3em; }
.authors .au { white-space: nowrap; }
.affil { margin: 6px 0 0; color: var(--muted); font-size: 15.5px; }
.affil span + span { margin-left: 1.1em; }
nav.links { display: flex; flex-wrap: wrap; justify-content: center; gap: 10px; margin-top: 28px; }
.btn { display: inline-flex; align-items: center; gap: 9px; padding: 10px 17px; border-radius: 999px; font: 500 15.5px/1 var(--sans); text-decoration: none; color: var(--bg); background: var(--ink); border: 1px solid var(--ink); }
.btn:hover { background: var(--accent); border-color: var(--accent); }
.btn svg { width: 17px; height: 17px; flex: none; }
.btn .emoji { font-size: 17px; line-height: 1; }
.btn small { font: 400 12.5px/1 var(--mono); opacity: .75; }
.btn.is-soon { background: transparent; color: var(--muted); border-color: var(--line); border-style: dashed; cursor: default; }
figure { margin: 0; }
.plate { background: #FFFFFF; border: 1px solid var(--line); border-radius: 14px; padding: clamp(10px, 2.4vw, 22px); }
.plate a { display: block; }
figcaption { color: var(--ink-2); font-size: 15.5px; line-height: 1.6; margin: 14px auto 0; max-width: 820px; }
figcaption strong { color: var(--ink); }
.short { margin-top: 56px; }
.short p.lead { font: 400 21px/1.55 var(--serif); margin: 0; color: var(--ink); }
.stats { display: grid; grid-template-columns: repeat(auto-fit, minmax(220px, 1fr)); gap: 14px; margin-top: 30px; }
.stat { background: var(--panel); border-radius: 12px; padding: 18px 18px 16px; }
.stat .num { font: 600 30px/1.1 var(--serif); letter-spacing: -.01em; font-variant-numeric: tabular-nums; color: var(--accent); }
.stat .num .was { color: var(--before); }
.stat .num .arrow { color: var(--muted); font-weight: 500; padding-inline: .15em; }
.stat p { margin: 8px 0 0; font-size: 14.5px; line-height: 1.5; color: var(--ink-2); }
section.block { margin-top: 72px; }
h2 { font: 600 28px/1.25 var(--serif); letter-spacing: -.005em; margin: 0 0 14px; text-wrap: balance; }
h3 { font: 600 17px/1.4 var(--sans); margin: 0 0 6px; }
section.block p { margin: 0 0 14px; color: var(--ink-2); }
section.block p strong, section.block li strong { color: var(--ink); font-weight: 600; }
section.block ul, section.block ol { color: var(--ink-2); padding-left: 1.3em; margin: 0 0 14px; }
section.block li { margin-bottom: 6px; }
.abstract p { font-size: 16.5px; }
.table-wrap { overflow-x: auto; margin: 22px 0 8px; border: 1px solid var(--line); border-radius: 12px; }
table { border-collapse: collapse; width: 100%; font-size: 15.5px; }
th, td { padding: 12px 16px; text-align: left; border-top: 1px solid var(--line); vertical-align: top; }
thead th { border-top: 0; font: 500 12px/1.3 var(--mono); letter-spacing: .06em; text-transform: uppercase; color: var(--muted); background: var(--panel); }
td { font-variant-numeric: tabular-nums; }
.figs { display: grid; gap: 22px; margin-top: 30px; }
dl.terms { display: grid; grid-template-columns: minmax(0, 1fr); gap: 14px; margin: 18px 0 0; }
dl.terms div { border-left: 3px solid var(--accent-soft); padding-left: 14px; }
dl.terms dt { font-weight: 600; color: var(--ink); }
dl.terms dd { margin: 2px 0 0; color: var(--ink-2); }
.faq { display: grid; gap: 22px; margin-top: 20px; }
.faq div p { margin: 0 !important; }
pre { margin: 0; }
.cite { position: relative; }
.cite pre { background: var(--panel); border-radius: 12px; padding: 18px 20px; overflow-x: auto; font: 400 13.5px/1.6 var(--mono); color: var(--ink); }
.copy { position: absolute; top: 10px; right: 10px; font: 500 13px/1 var(--sans); padding: 8px 12px; border-radius: 8px; border: 1px solid var(--line); background: var(--bg); color: var(--ink); cursor: pointer; }
.copy:hover { border-color: var(--accent); color: var(--accent); }
footer.foot { margin-top: 88px; padding-block: 28px 44px; border-top: 1px solid var(--line); color: var(--muted); font-size: 14.5px; }
footer.foot .row { display: flex; flex-wrap: wrap; justify-content: space-between; gap: 10px 24px; }
footer.foot a { color: var(--ink-2); }
footer.foot .counted { flex-basis: 100%; font-size: 13px; }
.consent { position: fixed; left: 16px; right: 16px; bottom: 16px; max-width: 560px; margin: 0 auto; padding: 14px 16px; border-radius: 12px; border: 1px solid var(--line); background: var(--bg); color: var(--ink); box-shadow: 0 6px 24px rgba(0,0,0,.12); font: 15px/1.45 var(--sans); }
.consent p { margin: 0 0 10px; }
.consent button { font: 500 14px/1 var(--sans); padding: 9px 14px; border-radius: 999px; border: 1px solid var(--ink); background: var(--ink); color: var(--bg); cursor: pointer; }
.consent button + button { background: transparent; color: var(--ink); }
footer.foot .credit { font-size: 13px; }
@media (max-width: 760px) {
body { font-size: 16px; }
header.hero { padding-top: 44px; }
section.block { margin-top: 56px; }
h2 { font-size: 24px; }
.short p.lead { font-size: 19px; }
}
@media (prefers-reduced-motion: no-preference) {
.btn, .copy { transition: background-color .15s, border-color .15s, color .15s; }
}
</style>
</head>
<body>
<div class="wrap">
<header class="hero">
<p class="eyebrow">Preprint · 2026</p>
<h1>Textual Planning with Explicit Latent Transitions</h1>
<p class="authors"><span class="au">Eliezer Shlomi<sup>1,*</sup></span><span class="sep">·</span> <span class="au"><a href="https://orcid.org/0009-0005-7400-3452">Ido Levy</a><sup>2,*</sup></span><span class="sep">·</span> <span class="au">Eilam Shapira<sup>1</sup></span><span class="sep">·</span> <span class="au">Michael Katz<sup>2</sup></span><span class="sep">·</span> <span class="au">Guy Uziel<sup>2</sup></span><span class="sep">·</span> <span class="au">Segev Shlomov<sup>2</sup></span><span class="sep">·</span> <span class="au">Nir Mashkif<sup>2</sup></span><span class="sep">·</span> <span class="au">Roi Reichart<sup>1</sup></span><span class="sep">·</span> <span class="au">Sarah Keren<sup>1</sup></span></p>
<p class="affil"><span><sup>1</sup>Technion – Israel Institute of Technology</span><span><sup>2</sup>IBM</span><span><sup>*</sup>Equal contribution</span></p>
<nav class="links" aria-label="Paper, code and resources">
<a class="btn" href="https://arxiv.org/abs/2602.04557" data-goatcounter-click="paper"><svg viewBox="0 0 16 16" aria-hidden="true"><path d="M3.5 1.5h6l3 3v10h-9z" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linejoin="round"/><path d="M9.5 1.5v3h3" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linejoin="round"/></svg>Paper <small>arXiv</small></a>
<a class="btn" href="https://github.com/embedplan/EmbedPlan" data-goatcounter-click="code"><svg viewBox="0 0 16 16" aria-hidden="true"><path fill="currentColor" d="M8 0c4.42 0 8 3.58 8 8a8.013 8.013 0 0 1-5.45 7.59c-.4.08-.55-.17-.55-.38 0-.27.01-1.13.01-2.2 0-.75-.25-1.23-.54-1.48 1.78-.2 3.65-.88 3.65-3.95 0-.88-.31-1.59-.82-2.15.08-.2.36-1.02-.08-2.12 0 0-.67-.22-2.2.82-.64-.18-1.32-.27-2-.27-.68 0-1.36.09-2 .27-1.53-1.03-2.2-.82-2.2-.82-.44 1.1-.16 1.92-.08 2.12-.51.56-.82 1.28-.82 2.15 0 3.06 1.86 3.75 3.64 3.95-.23.2-.44.55-.51 1.07-.46.21-1.61.55-2.33-.66-.15-.24-.6-.83-1.23-.82-.67.01-.27.38.01.53.34.19.73.9.82 1.13.16.45.68 1.31 2.69.94 0 .67.01 1.3.01 1.49 0 .21-.15.45-.55.38A7.995 7.995 0 0 1 0 8c0-4.42 3.58-8 8-8Z"/></svg>Code</a>
<a class="btn" href="#citation" data-goatcounter-click="bibtex"><svg viewBox="0 0 16 16" aria-hidden="true"><path fill="currentColor" d="M2.5 3.5h4.25v4.25H4.9c0 1.45.55 2.35 1.85 2.95l-.62 1.55C3.7 11.4 2.5 9.75 2.5 7.2zm6.75 0h4.25v4.25h-1.85c0 1.45.55 2.35 1.85 2.95l-.62 1.55c-2.43-.85-3.63-2.5-3.63-5.05z"/></svg>BibTeX</a>
</nav>
</header>
<main>
<figure class="teaser"><div class="plate"><a href="assets/figure1.png"><img src="assets/figure1.png" width="1740" height="330" fetchpriority="high" alt="EmbedPlan architecture. A Blocksworld state and the action pick-up(C) enter a frozen LLM encoder E. Learned heads project the two embeddings into a latent space, a learned transition network predicts the next-state embedding, and the nearest real state is returned as the next state, at about 0.17 ms per transition with cached embeddings."></a></div><figcaption><strong>EmbedPlan.</strong> A frozen LLM encoder embeds the state and the action, learned heads project them into a latent space where the transition is computed, and the nearest real state is retrieved as the successor. Search itself is external.</figcaption></figure>
<section class="short text" aria-label="In short"><p class="lead">When a large language model serves as a planner's transition model, every next state is generated token by token, which makes searching over many possible futures slow and expensive. EmbedPlan embeds the state and action with a frozen LLM, predicts the next-state embedding with a lightweight learned network, and returns the closest real state, evaluated on 9 classical planning domains under six settings that hold out progressively more of the data.</p></section>
<div class="stats" role="list"><div class="stat" role="listitem"><div class="num">99.7%</div><p>Hit@5 within observed problems: the true next state ranks in the top 5 of 128 candidates (Interpolation, Llama-3.3-70B, chance 3.9%).</p></div><div class="stat" role="listitem"><div class="num">92–99%</div><p>of step accuracy kept when each prediction is fed back as the next input, snapped to the nearest real state (Ferry, Logistics, Blocksworld, observed problems).</p></div><div class="stat" role="listitem"><div class="num">6.6%</div><p>Hit@5 on unseen domains (Cross-Domain), near the 3.9% chance level, against 54.6% on unseen problems of an observed domain.</p></div><div class="stat" role="listitem"><div class="num">0.17 ms</div><p>per transition with cached state embeddings, against 1.9 s for generating the next state through an LLM API.</p></div></div>
<section class="block text abstract" id="abstract"><h2>Abstract</h2><p>Planning requires a transition model that predicts how each action changes the current state. When a large language model (LLM) plays this role, every next state is generated token by token, which makes searching over many possible futures slow and expensive. Existing alternatives either still query an LLM at every step or require a symbolic model of the domain. We propose EmbedPlan, a transition model built on frozen text embeddings: it embeds natural language descriptions of the state and the action with a frozen LLM, predicts the embedding of the next state with a lightweight learned network, and returns the closest real state. Because this network can be trained on top of any encoder, EmbedPlan also provides a controlled way to compare text representations for learning transitions. We evaluate it on 9 classical planning domains, under six settings that hold out progressively more of the data, from transitions to entire domains, and against baselines ranging from predicting no change to learning symbolic action rules. On planning problems seen during training, EmbedPlan almost always ranks the true next state among its top five guesses, still does so for most queries even when every observed state is a candidate, and retains 92–99% of its single-step accuracy when predicting several steps ahead from its own outputs. Given the same candidate states as GPT-5.4, it picks the true next state more often while taking about 0.17 ms per transition with cached embeddings. Accuracy is lower on unseen problems and near chance on unseen domains, and the controlled comparison traces this limit to the state representation rather than to the learned transition.</p></section>
<section class="block text" id="the-idea"><h2>The idea</h2><p>Planning needs a transition model that predicts how each action changes the current state. When a large language model plays this role, each next state is generated token by token, with a full forward pass per token, which makes multi-step lookahead and search expensive in latency and cost.</p>
<p>EmbedPlan keeps the semantic encoding in a frozen LLM and learns only the cheap dynamics. It predicts the embedding of the next state with a small network and returns the closest real state, so every prediction is a real state. Because the same head can be trained over any encoder, the framework also serves as a controlled probe of what a frozen text representation supports for learning dynamics.</p></section>
<section class="block text" id="how-it-works"><h2>How it works</h2><ul><li>A frozen LLM encoder embeds a natural language description of the state and of the action, for example a Blocksworld state and the action <code>pick-up(C)</code>. Each state prompt holds the problem and goal description followed by the facts that hold.</li><li>Learned state and action heads project both embeddings into a 128-dimensional latent space.</li><li>A residual MLP with fewer than 500K parameters predicts the embedding of the next state.</li><li>The nearest state in a candidate pool, by cosine similarity, is returned as the successor. Fed back as the next input, it supports multi-step rollout.</li><li>Two contrastive objectives train the head: state prediction, which identifies the correct next state among candidates, and action disambiguation, which separates the effects of different actions on the same state.</li></ul>
<p>EmbedPlan is a transition component, not a complete planner. Search composes many transitions, and the paper studies the single transition that search queries repeatedly. Training is per domain: a one-time embedding pass plus about 15 minutes of training.</p></section>
<section class="block text" id="how-it-was-tested"><h2>How it was tested</h2><p>The data are nine classical PDDL domains from ACPBench, with states rendered as natural language: Blocksworld, Depot, Ferry, Floortile, Goldminer, Grid, Logistics, Rovers and Satellite. Together they give 2.97M transitions over 67 problems (259K unique states). A domain fixes the action schemas, and a problem fixes the objects, the initial state, and the goal.</p>
<p>Four frozen encoders are compared: MPNet (all-mpnet-base-v2), BGE-M3, Qwen2.5-7B and Llama-3.3-70B. Results use Llama-3.3-70B unless another encoder is named.</p>
<p>Six protocols form a ladder of exposure, ordered by what the test data share with training:</p>
<ul><li>Interpolation: held-out transitions of observed problems.</li><li>Plan-Variant: unseen optimal plans of observed problems.</li><li>Extrapolation: unseen problems of an observed domain.</li><li>Multi-Domain: unseen problems, with one model for all nine domains.</li><li>Cross-Domain: an unseen domain, after training on one other domain.</li><li>Leave-One-Out: an unseen domain, after training on the other eight.</li></ul>
<p>Every query is ranked against 128 states, the true successor and 127 distractors. The reference methods keep the head and the pool and vary only the state representation, from no-change floors to lifted STRIPS induction, which serves as an oracle upper bound.</p></section>
<section class="block text" id="what-it-finds"><h2>What it finds</h2><p>Within observed problems, EmbedPlan has the properties search needs. Hit@5 is 99.7% with Llama-3.3-70B. When the candidate pool grows to every observed state of the domain, Hit@5 falls from 100.0% to 86.9% on Ferry (46,205 states) and from 99.9% to 74.9% on Logistics (13,373 states). Snapping each prediction to the nearest real state keeps multi-step rollout within 92–99% of the accuracy obtained with the true state at each step.</p>
<p>Within observed problems, EmbedPlan matches or beats every LLM tested. Given the identical ranking task over the same 128 candidates, GPT-5.4 selects the true successor for 44% of Ferry and 92% of Logistics queries, against 99.0% and 96.3% for EmbedPlan. The comparison is matched in task, not in information: EmbedPlan has observed other transitions of the same problems, whereas the LLM is zero-shot. It ranks successors 101× faster end to end than generation through an LLM API (1.9 s per transition), and takes about 0.17 ms per transition with cached state embeddings.</p>
<p>Accuracy falls with each step down the exposure ladder. On unseen problems of an observed domain, Hit@5 is 54.6%, 14 times chance, and it varies widely by domain (26–76%). Larger encoders extrapolate better, from 26.8% for MPNet to 54.6% for Llama-3.3-70B, but none closes the gap. On unseen domains, Hit@5 is 6.6% (Cross-Domain) and 9.2% (Leave-One-Out), near the 3.9% chance level.</p>
<p>With the head, data and pool fixed, the state representation decides transfer to unseen problems. On Ferry, Logistics and Goldminer, character 3–5 grams of the same text reach 79.4% Hit@5, against 56.8% with Llama-3.3-70B embeddings, and lifted STRIPS induction, given each state's symbolic facts, reaches 99.8%. On these templated states an action edits only a few facts of an otherwise identical text, which sparse and symbolic representations capture directly and a pooled LLM embedding barely registers. The paper names a concrete target for transfer to new problems: an object-aware, order-invariant state encoder.</p></section>
<section class="block text" id="limitations"><h2>Limitations</h2><ul><li>EmbedPlan is a transition component for discrete, domain-specific settings, not a domain-general planner, and must be trained per domain because cross-domain transfer fails.</li><li>Retrieval assumes a candidate pool that contains the true successor.</li><li>The domains are templated, so the paper has not yet tested the regime the text interface is meant for: states without a clean literal decomposition, where symbolic induction and lexical features would break.</li><li>Pool scaling, multi-step rollout, and LLM ranking are measured within observed problems, and rollouts cover short test trajectories (mean 2.2 steps).</li><li>Extrapolation rests on one or two held-out problems per domain, and the reference comparison covers three domains, without ablating the problem and goal text that every prompt shares.</li><li>Integrating EmbedPlan into beam or tree search over long horizons on held-out problems is the direct next step.</li></ul></section>
<section class="block text" id="use-it-on-your-data"><h2>Use it on your data</h2><p>The code includes EmbedPlan as a scikit-learn style estimator for any domain whose states and actions can be written as text, such as planning problems, game logs, web or UI agent traces, or lab protocols.</p>
<ul><li><code>fit</code> learns from (state, action) pairs and their next states, with optional groups such as problem ids.</li><li><code>predict</code> returns the most likely next state as text, ranked among candidate states, and <code>rollout</code> predicts several steps, each snapped to the nearest real state.</li><li><code>evaluate</code> follows the paper's protocol, 128 candidates per query, and reports the chance level of the same pools.</li><li>The encoder can be a hashing encoder of word n-grams that needs no download, any sentence-transformers or Hugging Face model name, or your own function from texts to vectors.</li></ul>
<p>Try it with no install in the <a href="https://colab.research.google.com/github/embedplan/EmbedPlan/blob/main/examples/quickstart.ipynb">Colab notebook</a>, or start from the <a href="https://github.com/embedplan/EmbedPlan">repository</a>. On unseen problems, expect lower accuracy than on seen ones: that is the paper's main finding.</p></section>
<section class="block text" id="results"><h2>Results by protocol</h2><p>Hit@5 (%) with the Llama-3.3-70B encoder, mean ± SE across the nine domains. Every query is ranked against 128 candidates, so chance is 3.9%. Rows are ordered by what the test data share with training, and the gain is over chance in percentage points (pp). Plan-Variant faces the hardest distractors. From Table 5 of the paper.</p><div class="table-wrap"><table><thead><tr><th scope="col">Protocol</th><th scope="col">Observed in training</th><th scope="col">Hit@5 (%)</th><th scope="col">Gain over chance (pp)</th></tr></thead><tbody><tr><td>Interpolation</td><td>The problems</td><td><strong>99.7 ± 0.1</strong></td><td>+96</td></tr><tr><td>Plan-Variant</td><td>The problems</td><td>51.2 ± 5.5</td><td>+47</td></tr><tr><td>Extrapolation</td><td>The domain</td><td>54.6 ± 5.5</td><td>+51</td></tr><tr><td>Multi-Domain</td><td>The domain</td><td>37.2 ± 3.8</td><td>+33</td></tr><tr><td>Leave-One-Out</td><td>8 other domains</td><td>9.2 ± 1.2</td><td>+5.3</td></tr><tr><td>Cross-Domain</td><td>1 other domain</td><td>6.6 ± 0.5</td><td>+2.7</td></tr><tr><td>Chance</td><td></td><td>3.9</td><td></td></tr></tbody></table></div></section>
<section class="block text" id="terms"><h2>Terms</h2><dl class="terms"><div><dt>EmbedPlan</dt><dd>A transition model built on frozen text embeddings: it embeds natural language descriptions of the state and the action with a frozen LLM, predicts the embedding of the next state with a lightweight learned network, and returns the closest real state.</dd></div><div><dt>Transition model</dt><dd>A model that predicts how each action changes the current state, which planning requires.</dd></div><div><dt>Hit@k</dt><dd>The fraction of queries whose true next state ranks in the top k of a candidate pool. Unless stated otherwise the pool holds 128 states, the true successor and 127 distractors, so chance is 3.9% for Hit@5 and 0.8% for Hit@1.</dd></div><div><dt>Head</dt><dd>The learned state and action projection heads and the transition network, together. Only the head is trained, and the LLM encoder stays frozen.</dd></div><div><dt>Snapping</dt><dd>In multi-step rollout, each state prediction is replaced by the nearest real state the model retrieves, right or wrong, and that state becomes the next input.</dd></div><div><dt>Domain and problem</dt><dd>A domain fixes the action schemas, and a problem fixes the objects, the initial state, and the goal.</dd></div><div><dt>Grounded action</dt><dd>An instance of a lifted action schema with its objects filled in, such as <code>pick-up(C)</code> for the schema <code>pick-up(?x)</code>.</dd></div><div><dt>Interpolation</dt><dd>The protocol that trains on 80% of a domain's transitions and tests on the other 20%, held out at random, so test transitions come from observed problems, with distractors from the whole domain.</dd></div><div><dt>Plan-Variant</dt><dd>The protocol that trains on some optimal plans and tests on other optimal plans of the same problems, with distractors from successors along alternative optimal plans, the hardest distractors the paper uses.</dd></div><div><dt>Extrapolation</dt><dd>The protocol that trains on about 80% of a domain's problems and tests on the other 20%, so test transitions come from unseen problems, with distractors from the query's own problem.</dd></div><div><dt>Multi-Domain</dt><dd>Extrapolation with one model for all nine domains at once.</dd></div><div><dt>Cross-Domain</dt><dd>The protocol that trains on one domain and tests on another, with distractors from the target domain.</dd></div><div><dt>Leave-One-Out</dt><dd>The protocol that trains on eight domains and tests on the ninth, with distractors from the target domain.</dd></div><div><dt>Identity and Offset</dt><dd>Two non-learned floors that need no training. Identity predicts the current state's embedding, and Offset adds the mean training displacement of the ground action or of its schema.</dd></div><div><dt>Lifted STRIPS induction</dt><dd>A reference method that infers each action schema's add and delete effects from the symbolic facts of each state. It serves as an oracle upper bound.</dd></div><div><dt>Action disambiguation</dt><dd>A training objective that separates the effects of different actions on the same state: the true action must yield a prediction closer to the next state than up to 50 alternative actions applied to the same state.</dd></div></dl></section>
<section class="block text" id="questions"><h2>Questions</h2><div class="faq"><div><h3>What is EmbedPlan?</h3><p>EmbedPlan is a transition model for planning built on frozen text embeddings. Given a state and an action described in natural language, a frozen LLM embeds both, a lightweight learned network predicts the embedding of the next state, and the closest real state is returned. It is a transition component, not a complete planner: search composes many transitions, and the paper studies the single transition that search queries repeatedly.</p></div><div><h3>Why not let an LLM generate the next state?</h3><p>Because generating every next state token by token makes search slow and expensive. When an LLM serves as the world model of a planner, each next state is generated with a full forward pass per token, which makes multi-step lookahead and rollout-based search prohibitively expensive in latency and cost. EmbedPlan leaves the semantic encoding to the frozen LLM and learns only the cheap dynamics, so each prediction is one learned vector step followed by retrieval.</p></div><div><h3>Is EmbedPlan better than asking an LLM like GPT-5.4?</h3><p>Within observed problems, at picking the true next state, yes. Given the identical ranking task over the same 128 candidates, GPT-5.4 selects the true successor for 44% of Ferry and 92% of Logistics queries, against 99.0% and 96.3% for EmbedPlan. On Logistics the margin is small next to the spread across the paper's runs (91.2–96.3%). The comparison is matched in task, not in information: EmbedPlan has observed other transitions of the same problems, whereas the LLM is zero-shot, and on unseen problems its Ferry Hit@1 falls to 12.0%.</p></div><div><h3>How fast is EmbedPlan?</h3><p>About 0.17 ms per transition with cached state embeddings, 11,215× faster than generating the next state through an LLM API (1.9 s per transition). End to end, including encoding the query texts with BGE-M3, it takes 18.6 ms, 101× faster. Training is a one-time cost per domain, an embedding pass plus about 15 minutes of training, which breaks even against autoregressive generation after about 4,300 transitions with Llama-3.3-70B, so it pays off for a domain that is planned in repeatedly.</p></div><div><h3>What does Hit@5 measure, and what is chance?</h3><p>Hit@5 is the fraction of queries whose true next state ranks in the top 5 of a candidate pool. Unless stated otherwise the pool holds 128 states, the true successor and 127 distractors, so chance is 3.9% for Hit@5 and 0.8% for Hit@1. The paper emphasizes Hit@5 because it is robust to the pool size, and a top-5 list is also useful where a verifier can filter candidates.</p></div><div><h3>How well does EmbedPlan generalize to new problems and new domains?</h3><p>It depends on what the model has seen. With the Llama-3.3-70B encoder, Hit@5 is 99.7% on held-out transitions of observed problems, 54.6% on unseen problems of an observed domain, and 6.6% on unseen domains, near the 3.9% chance level. Training on the other eight domains (Leave-One-Out) reaches 9.2%. One model trained on all nine domains reaches 37.2% on their unseen problems.</p></div><div><h3>What limits transfer to unseen problems and domains?</h3><p>The paper traces the limit to the frozen state representation rather than to the learned transition. With the head, data and pool fixed, character 3–5 grams of the same text reach 79.4% Hit@5 on unseen problems of three domains, against 56.8% with Llama-3.3-70B embeddings. On unseen domains, the paper sees two plausible causes that its experiments do not separate: frozen embeddings organize states by surface form rather than by structural role, and the learned action semantics are domain-specific.</p></div><div><h3>Does EmbedPlan stay accurate over several steps?</h3><p>Yes, over short horizons, when each prediction is snapped to a real state. If each predicted state is replaced by the nearest real state before the next step, rollout keeps 92–99% of the accuracy obtained with the true state at each step, on Ferry, Logistics and Blocksworld within observed problems. Without snapping, the predicted embedding drifts off the real states. The test trajectories are short (mean 2.2 steps on Ferry), so this establishes stability over short horizons only.</p></div><div><h3>Does the true next state have to be among the candidates?</h3><p>Yes. Retrieval assumes a candidate pool that contains the true successor. The paper measures how accuracy degrades as the pool grows to every state observed in a domain. On Ferry (46,205 states), Hit@5 falls from 100.0% to 86.9%, and on Logistics (13,373 states) from 99.9% to 74.9%, while Hit@1 falls to 35.2% and 32.6%. The true successor thus stays in a short list against the full pool, though not reliably at rank one.</p></div><div><h3>Which domains and encoders does the paper use?</h3><p>Nine classical PDDL domains from ACPBench, with states rendered as natural language: Blocksworld, Depot, Ferry, Floortile, Goldminer, Grid, Logistics, Rovers and Satellite, 2.97M transitions over 67 problems. The frozen encoders are MPNet (all-mpnet-base-v2), BGE-M3, Qwen2.5-7B and Llama-3.3-70B. Larger encoders extrapolate better, from 26.8% Hit@5 on unseen problems with MPNet to 54.6% with Llama-3.3-70B, but none closes the gap.</p></div><div><h3>When is EmbedPlan worth using?</h3><p>When states and actions are available as text but no symbolic model is, so that the alternative is querying an LLM at every expansion. Action selection and search remain external to EmbedPlan. It must be trained per domain, with a one-time embedding pass plus about 15 minutes of training, so it pays off for a domain that is planned in repeatedly. The paper's classical domains have exact simulators, which makes them a controlled testbed.</p></div><div><h3>When EmbedPlan misses, how far off is its top guess?</h3><p>By a few facts. When the true successor misses the top five on unseen problems, the top-ranked state typically differs from it in one to three facts (median 2), usually involving object locations or holdings within the same problem instance. These errors suggest the model captures coarse transition structure (correct problem context, approximate state region) but struggles to resolve fine-grained predicate changes, particularly when multiple objects undergo similar transformations.</p></div><div><h3>Would fine-tuning the encoder help?</h3><p>Adapting the encoder already helps. Low-rank adaptation (LoRA) of BGE-M3 raises full-pool Hit@1 on unseen problems from 3.8% to 9.3% when the head is first trained with the encoder frozen, and to 6.2% with a cold joint start. Larger frozen encoders also extrapolate better. Because the encoder is a plug-in, the paper names a concrete target the framework can evaluate directly: an object-aware, order-invariant state encoder.</p></div><div><h3>Can EmbedPlan tell apart the effects of different actions?</h3><p>Within observed problems it mostly can, and on unseen problems far less often. The paper applies every action applicable in a state and ranks the actions by how close their predictions come to the true next state. Acc@k is the fraction of queries whose true action ranks in the top k. Mean Acc@5 is 85.3% within observed problems (Acc@1 32.0%) and 16.2% on unseen problems, where training without the action loss reaches only 4.8%.</p></div><div><h3>How often does EmbedPlan get a whole plan right?</h3><p>Under Plan-Variant, where the plans are unseen but their problems are observed, 22.6% of plans are correct at every step and mean per-step Hit@5 is 51.2%. Each step is ranked against successors along alternative optimal plans, the hardest distractors the paper uses: they are reachable and often differ from the target in a single fact. Plans of unseen problems fare far worse, at 10.1% per step and 3.3% of plans.</p></div></div></section>
<section class="block text" id="citation"><h2>Citation</h2><div class="cite"><button class="copy" type="button" id="copy-bib" data-goatcounter-click="bibtex-copy" hidden>Copy</button><pre id="bibtex">@article{shlomi2026textual,
title = {Textual Planning with Explicit Latent Transitions},
author = {Shlomi, Eliezer and Levy, Ido and Shapira, Eilam and Katz, Michael and Uziel, Guy and Shlomov, Segev and Mashkif, Nir and Reichart, Roi and Keren, Sarah},
journal = {arXiv preprint arXiv:2602.04557},
year = {2026}
}</pre></div></section>
</main>
<footer class="foot"><div class="row"><span>Eliezer Shlomi, Ido Levy, Eilam Shapira, Michael Katz, Guy Uziel, Segev Shlomov, Nir Mashkif, Roi Reichart, Sarah Keren</span><span><a href="https://github.com/embedplan/EmbedPlan">Code</a></span><span class="counted">Visits are counted without cookies (GoatCounter).</span></div></footer>
</div>
<script>
(function () {
var btn = document.getElementById("copy-bib");
var pre = document.getElementById("bibtex");
if (!btn || !pre || !navigator.clipboard) return;
btn.hidden = false;
btn.addEventListener("click", function () {
navigator.clipboard.writeText(pre.textContent).then(function () {
if (window.gtag) window.gtag("event", "peo_bibtex_copy");
btn.textContent = "Copied";
setTimeout(function () { btn.textContent = "Copy"; }, 1600);
}, function () {
var r = document.createRange(); r.selectNodeContents(pre);
var s = window.getSelection(); s.removeAllRanges(); s.addRange(r);
});
});
})();
</script>
<script data-goatcounter="https://embedplan.goatcounter.com/count" async src="//gc.zgo.at/count.js"></script>
</body>
</html>