Skip to main content

Changelog

One line per change, newest first. Every line says what is different for you, not what a developer touched in the code.

0.3.11Unreleased

Added

  • `browser`, the agent's signed-in browser as one MCP tool. 63 of 136 real sessions drove the agent browser through hand-written DevTools scripts; that is now built in, as one tool with an action instead of six thin ones: open (reuses a tab already on that site; starts the browser on the agent desktop if needed), read (interactive elements as ref|role|name|=value|state, stable refs, open shadow roots, or the page text), click, type, select, check, fill (a whole form, every field read back), upload (file input or upload button, no dialog), wait (text, css:, url:), screenshot (works on the invisible desktop), tabs / switch / close_tab, eval, logins and sign_in. Every action reads its effect back from the page (navigation, DOM change, new tab, checked state, value, input.files); a click that changes nothing says so. Read-only and disabled fields are refused, covered elements are not clicked, the senden/kaufen/… permissions and the window rules apply exactly as for click, vault placeholders are filled against the tab's real address. Costs +462 tokens per session.
  • Login handoff (`sign_in`). Asks first; only with user_agreed=true does it close the agent's browser and open the same profile on the user's screen (no DevTools port while they type their password). Once they close the window, it reads the profile's cookies and reports "signed in to x.com", or honestly that no login arrived, and open continues invisibly with that login.
  • New permission "Run scripts in web pages" (skripte, off by default): browser action='eval' refuses without it.
  • The agent browser now starts with renderer/timer backgrounding off, so pages on the hidden desktop keep animating and accept input (WhatsApp's file preview never rendered before).

Fixed

  • A field that stayed empty no longer counts as "holds the text". Typing into an <input type=time> fired beforeinput, left the field empty, and type_text/browser reported success, because an empty read-back "contained" the text. Date, time, colour and range fields are now set and read back.

Removed (features without real-world use)

  • Removed: Schedules, standing goals, routines, the routine inbox and the `desktop_goal` MCP tool. The tables were empty in real use and the tool appeared in none of 136 work reports; Claude Code's own /loop and scheduling cover it. Existing databases open unchanged (tables stay, unused).
  • Removed: the phone remote (phone page, QR pairing, its own TLS socket). It was never switched on and was the only way the service listened beyond this machine; every request from another device is now refused.
  • Removed: the memory folder mirror (Obsidian) and the memory graph. The mirror was never used; the list view of memory stays, with its colours.
  • Removed: "Record a skill" (Vormachen). It relied on global low-level input hooks, froze the app once and was never used on a real app. Skills are still learned from successful runs, imported as SKILL.md and replayed.
  • Removed: voice input. Off in real use; dictation happens elsewhere.
  • Old settings for these features are ignored when read; nothing to migrate.

Desktop helper process (on by default, AGENTFENSTER_HOST=0 to opt out)

  • One long-lived helper per agent desktop. With AGENTFENSTER_HOST=1, reading, clicking and typing on the agent desktop run in a helper process started directly on that desktop (python -m agentfenster.kern.host, job object, named pipe) instead of a fresh UI Automation thread per call. Elements carry a stable id from their UI Automation runtime id, so the action hits exactly the element that was read; an unknown id is re-read once, then refused. This is the default now; AGENTFENSTER_HOST=0 brings back the old thread path, which also takes over if the helper cannot start.
  • A frozen app no longer freezes the tool. A window that stops processing messages is reported as "not responding" instead of being read or touched; a call that runs past its deadline kills and restarts the helper and says that whether the action took effect is unknown. A click that sends an app into a hang now answers in about 2 s instead of waiting out the hang (the click target is described before the click, the replay picture of a hung window is skipped).
  • The helper pipe trusts nobody it did not check. The pipe allows only the current user, rejects remote clients and refuses a second instance; the client connects without identity, checks owner SID and program of the pipe server, and a self-started helper must answer with its start token.

Clicks and typing hit the element that was named, and say so

  • Search behind the 200-element cap. Discord puts its message field behind 200 server and DM entries, so read_desktop(filter=...) and type_text(element='Message @...') never found it (three real tasks in the week of 21 Sep). A name, id or filter that the capped tree does not contain now triggers one wider read (up to 1,500 elements); the answer stays small because only matches go into it. Ids do not change between the two reads.
  • Twins: the id decides, not the order. Two buttons with the same name got different ids, but the click always landed on the first one and then reported "changed". The position among same-named elements now travels with the click and the typing.
  • Ambiguous partial names refuse. click('Save') with "Save draft" and "Save and send" on screen used to press the first; it now lists both with their ids. A number that is not a current id no longer matches an element whose name merely contains it ("153" hit "Server 153").
  • type_text reads the field back. The answer says whether the target field now contains the text, warns when the text showed up in a different field or the field shows something else, and says plainly when a field does not expose its text. Vault placeholders are never compared.
  • click says what changed ("new: ...; gone: ...; changed: ..."), not only that something did.
  • Unnamed input fields are visible. Electron and web text fields often have no accessible name and were missing from the tree entirely. Unnamed edit and combo fields now get a line and an id and can be typed into.

desktop_task always answers within its time limit

  • A hard wall clock around the whole call. On 27 Sep 2026 a desktop_task(max_seconds=240) did not return for 32 minutes: the click loop's own deadline fired at 245 s, but the call kept waiting for a run thread stuck outside that loop, and the Claude session gave up after 1800 s. The whole call (model setup, run, report) now runs under one wall clock of max_seconds + 30 s; when it strikes, the answer says so, says the task is not verified, and the run is told to stop, so a model that answers late can no longer start clicking after the caller was told it stopped.

open_app on a browser names the real DevTools port

  • Verified, not read from an old file. Five times in the week of 21 Sep the agent had to hunt for the agent Brave's DevTools port because DevToolsActivePort named a port nobody answered on, or the profile was held by another Brave instance. open_app now names the port only when the running browser answers there with this profile's id; a stale port file is called stale, and a profile held by another instance without a DevTools port is reported as exactly that, with what to do.

open_app no longer opens a second tab for a page that is already open

  • The open tab comes to the front instead. open_app('brave.exe https://web.whatsapp.com') with WhatsApp already open in the agent's browser opened a second tab, and WhatsApp switched the first one to "open in another window". When a tab with the same host and path is already open (checked on the verified DevTools port), that tab is activated, no browser is started, and the answer says so. A different path on the same site still opens normally; if the browser does not confirm the activation, the old way is used.
0.3.102026-09-22

Session reports actually leave the PC now

  • The background sender was never started. Every run and every MCP session wrote its masked summary into the outgoing folder, as promised - and there it stayed: the thread that ships those files to the agentfenster server had no caller outside the test suite since 0.3.0. Measured on 22 Sep 2026: 5111 reports waiting in one installation, the oldest from 5 Sep, none ever sent, and the 50-file / 7-day cap never ran either because it lives in the same thread. The app's lifespan now starts the sender once per process (also headless) and stops it on shutdown.
  • Reports need a key, and the Reporting card says so. Without an activated tester key there is no device token, and without a token nothing is sent. That was true before and written nowhere; the card now says "Reports are collected on this PC and start leaving it once a key is activated below."
  • Server side: the report service hung for 17 days because the TLS handshake ran in its single accept thread - one client that connected and never finished the handshake blocked every later request, without any thread dying. The handshake now runs per connection with a 20 s limit, the accept queue is 64 instead of 5, and the service runs under systemd with Restart=always.
0.3.92026-09-07 (evening)

Four seams between this morning's packages (hunting round)

  • A note you delete in the app now also disappears from the mirror folder. The Obsidian mirror only ever wrote files, never removed one: a note moved to the memory trash stayed in the vault, Obsidian kept showing it, and its wikilinks kept pointing at it. Restoring the note from the trash then left two files for the same note, because a restored note gets a new id. Only files this mirror wrote itself are touched, anything you put in the folder by hand is still left alone, a truncated note list never counts as "deleted", and a file that cannot be removed (open in Obsidian, network drive gone) is retried on the next run instead of being forgotten.
  • A click that opens a different page no longer claims anything about the video on the old one. If the page stopped answering right after the click, the result sentence used to repeat the state measured before it, "the video on this page is still paused at 0.0 s" about a page that no longer existed. It now says that it could not check, and why.
  • The guided tour is offered at most once. Ending the setup with Start tour and then closing the tour left no trace, so the next start of the program pushed the same tour at you again unasked. Any way you see the tour now counts as having been offered.
  • The prompt-injection test is part of the test bench again (A61). It was built as A60, never wired into the catalogue, and that number was handed to the video task the next day, so pruefstand.fahren A60 quietly ran a different task, while docs/SECURITY.md cited its numbers. A guard now walks every task file, not just the catalogue list.

The memory is now a folder you can open in Obsidian

  • Every memory note can be mirrored to a folder as Markdown. Settings → Memory → Memory folder turns it on and picks the place (default Documents\agentfenster-memory). Each note becomes one file with a YAML header, one folder per application, an _index.md per app, and laeufe/ with one file per run. Notes from the same run link to each other with [[wikilinks]] and carry a Used by section pointing at the runs that used them later, open the folder as an Obsidian vault and the graph view works without any setup. Measured on a copy of a real memory: 117 notes → 132 files in 0.19 s.
  • What you change there flows back. Edit the text of a note in Obsidian and agentfenster picks it up on the next mirror run, the text only; header fields are the app's bookkeeping and never flow back, so a typo in app: cannot re-file a note. Rename or move a file and it still belongs to the same note; the id in the header is the key, not the file name.
  • Nothing is decided silently. If a note changed in the app and in the folder since the last mirror, that is a conflict: neither side is overwritten, and the inbox says which file it was. Delete a file and the note goes to the memory trash (recoverable) with a line in the inbox, but if the folder is gone, or if a lot of files vanish at once, that reads as an accident: nothing is deleted and the app says so instead.
  • It never holds a run up. Mirroring runs on its own thread when a run ends, when you edit a note and at start, with a 5-second cap; if it runs out of time it says how far it got instead of finishing quietly. The Mirror now button answers with numbers and the path, never with "done".

After a run ends, what the run did stays on screen

  • The activity column no longer vanishes in the second a run finishes. On a window narrower than 1600 pixels - a 1280x800 laptop screen, say - the column on the right folded away the instant a run ended, taking the "Done" line with its result sentence, every step, every "why" and the "Hand off to Claude Code" button with it; the bar above the terminal then read "No run yet". That happened at exactly the moment agentfenster pulls you back to the window to look at the result. The column now stays until the next task starts, and the bar says "Run finished - the stage opens again with the next run." The stage itself still shrinks to that one line, so the chat keeps the room it gained in 0.3.7.
  • A recording that leaves your windows minimized tells you in the main window. The sentence was written into the small recorder overlay - which agentfenster closes before the sentence even exists, so nobody ever read it. It now waits for the main window and appears there once, when it comes back.
  • The product tour really does come back at the next start. Finishing setup with "Run it" postpones the tour so it cannot cover the first run (0.3.8). Nothing asked for it again afterwards, so anyone taking that route never saw the tour automatically. The next start opens it, once.
  • Searching the memory and then deleting a note updates the list. The deleted note stayed in the results under "1 note contains ..." while the message said "Moved to trash"; after "Save" the old text stood next to the card showing the new one. Only typing in the search box cleaned it up.
  • A note in the memory list reads as one thing to a screen reader. The whole row is a button, and it held a paragraph and a box - so the reader announced the entire line, counters and timestamp included, as the button's name.

After a click, agentfenster says whether the video is actually playing

  • The worst kind of silent success, measured on real YouTube. In four runs against youtube.com the model reported "the video is playing" every single time; the browser, asked at the same moment, said paused: true - once with the clock at 0.0 s (it never started), once at 2.9 s (it had started, and the click on the player's Play/Pause toggle had stopped it). The model was not careless: it had nothing better to go on. A video site's control tree only offers hints - a tab title with a speaker icon, a button that is called "Play" on one visit and "Pause" on the next. Whether a video plays is known only to the media element itself.
  • So that is what is asked now. After every click inside a Chromium page, agentfenster reads the page's <video>/<audio> and puts the answer in the step's own sentence: "The video on this page is playing now (0.2 s in).", or "... is still paused at 0.0 s - the click did not start playback." A page with no media pays nothing and says nothing. paused: false on its own is not accepted as playback - the clock has to have moved, because play on a dead source sets the flag for a moment and then gives up.
  • And a page with an odd clock no longer costs the click. If a page reports something other than a number for its playback position, reading it used to throw, and a click that had worked was reported as a crash. An unreadable clock now counts as "not playing" - the click keeps its own result.

A new test task: browser open, search a video site, play the first result

  • A60 walks the whole of Philipp's oldest standing order - "browser open -> YouTube" - offline: a video site with a consent banner that really covers the page (the control tree shows only the banner, and the server refuses to search until it is answered), a search form, eight results that share words so that "the first one" is a place and not a guess, and a video that genuinely plays. Judged on what reaches the server; only the title has to come back in the agent's own closing sentence. Three runs, three passes, 7 steps and ~32 s each.
  • Why it is not judged against youtube.com: the search results there are Shorts, the consent dialog differs by country, and a video opened directly does not start on its own (Chromium's autoplay policy - measured: page visible, readyState: 4, duration: 634.6 s, and still paused at 0 s). A judge hanging on that measures the service, not the product.

Memory: search it, open a note, fix it

  • The field above the memory list is a search now, not an app filter. It used to match the application name only, in the browser, on the notes that happened to be loaded, so if you remembered that some note said "Auswaehlen" but not which app it belonged to, you could not find it. It now searches the note text, the goal and the app name on the server, and the matching words are marked in the result. Measured on a copy of a real memory: "charmap" → 11 notes in 13 ms.
  • A note opens. Click any line (in the list or in the graph) and it opens as a card: the text in an editable field, the goal, how often the lesson was seen and how often it actually helped, which run learned it (with a way straight to that recording, when it still exists) and which runs later got it as a hint, the backlinks. Notes the same run learned are listed and one click away.
  • You can correct a note instead of only deleting it. The agent learns wrong things too; throwing the whole note away was a bad trade. Saving is verified against what the database actually holds afterwards, if the stored text is not what you typed, you get a plain sentence saying so, not "Saved".
  • The trash bin now catches what you throw away. "Forget this" deleted a note for good, while the automatic cleanup right next to it moved notes to the trash. Same icon, different finality. Both go to the trash now, and both come back from it.
  • The graph draws a line between notes the same run learned. Until now the only connections were "belongs to this app" and "about the same element", which left the picture as separate islands. A misclick and the click that worked afterwards now hang together, even across different goals.
  • Honest gap: origin and backlinks are recorded from 0.3.8 onwards. Notes learned before that carry neither, the lines were never written, and the card says that in a full sentence instead of showing an empty field.
  • A saved note no longer lands in the card you opened meanwhile. Press Save, then click a note from the same run while the save is still in flight, and the first note's text was written into the second note's card, it looked as if you had just edited the wrong note. Closing the card during a save broke the card outright. Both actions now remember which note they meant: the save still happens and still reports, but it only touches its own card.
  • Pressing Save without changing anything no longer claims the note is empty. It said "A note needs some text." over a note full of text; it now says "Nothing changed."

Your phone can now start a task, not just watch one

  • The phone was a spectator. Its empty state read "Nothing is running right now. Start a task in agentfenster on your PC", which is the one thing you cannot do from the sofa. There is now a task box at the top of the phone page: type what the agent should do, tap Send, and the run starts on the agent desktop and opens right there. Measured on an emulated Pixel 5 over the real network address: 0 taps from the scanned QR code to the page, 2 taps from there to a running task. The task goes through the same core as the task bar, the chat and the fleet, so your brain, profile, permissions and the double-run guard all still apply; only the goal comes from the phone.
  • The result was invisible. The phone had the run's outcome in hand all along and never showed it: a run that failed looked exactly like one that succeeded ("finished · 7 steps"). A finished run now carries its own sentence on the phone, green when it reached its goal, warm red when it did not, and a plain "The run stopped without reaching its goal." when the run has nothing better to say.
  • Dead ends were one-sided. The phone said "This phone access link has expired"; the PC said nothing at all, so anyone debugging it was looking at the wrong device. Settings → Phone now shows the last attempt in the same words, "The last attempt from a phone 2 min ago was refused: …", "A phone was connected just now.", or the honest "No phone has used this link yet."
  • It is called Phone, and it is explained in three lines. Not "API", that word appears nowhere a phone user can see it. The three lines carry the limits with them: same Wi-Fi only, no relay on the internet, and one certificate warning when the connection is encrypted (measured: net::ERR_CERT_AUTHORITY_INVALID is exactly what a phone hits first).

Renaming a file no longer loses its extension

  • Windows asks, and the agent used to answer "Yes". Renaming vorher.txt to nachher.txt in Explorer, the agent pressed F2, typed nachher, and Windows asked whether the file name extension may really change. It clicked Yes, reported "renamed to nachher.txt", and left a file called nachher on disk. Measured today with the cheaper model; measured across the log, this trap costs 7 of 52 runs on the two rename tasks.
  • The first "Yes" on that dialog is now refused, with a sentence that says why. The run is told that the name it typed drops the extension, that "No" is the way out, and to check the folder before building on the result. Asking for the exact same click a second time goes through, so deliberately turning report.txt into report.csv still works. Live proof: the run that used to die here answered "No", pressed F2 again, typed nachher.txt and passed.

When a step fails, the agent is told why - not just that it failed

  • The prompt block that lists dead ends ("does not work in this run - do not repeat it") named only the action, never the reason. The sentence explaining it ("not possible in the background: no child window, no UI Automation pattern") stayed in the log. A model that only sees the action blames the name and clicks the neighbour; one that reads the reason changes tool. Measured over 313 recordings: 315 steps went nowhere, and 100 of them were followed within three steps by the same kind of action again.
  • Costs nothing when nothing fails: the block does not exist without a dead end.
  • A Windows error code counts as a reason again. Anything starting with a dash was dropped as "no reason given" - which swallowed every Windows result code written as a negative number (-2147024894 … file not found), exactly the kind of reason this block exists for.

Four honest sentences, and a session report that keeps its promise

  • Your Windows account name no longer slips into a session report. The report replaces the browser profile block - the folder path and the sites you are signed in to - with a bare count. That filter needed whole sentences, while the recorder cuts a tool answer at 1200 characters: with a long sign-in URL (an OAuth address is about that long) the cut fell inside the folder path, the filter no longer matched, and C:\\Users\\<your name>\\… travelled out with the report. It now also handles a cut-off sentence, a Windows account with an apostrophe in it (O'Neil), and a truncated list - which says some sites instead of guessing a number.
  • A cut-off site name no longer leaves its tail in the report. The filter for "This profile is NOT signed in to …" asked for the sentence to end in a full stop, and a greedy match fell back to the last dot in the host - so a truncated praxis-dr-mueller.example went out as "… the requested site.example", and with a host like bank.MeinePraxis the leftover was the telling part. The whole name goes now, cut off or not.
  • A note that comes back from the trash brings its backlinks with it. Restoring wrote a new note with a new id, so "used in 3 runs" turned into "used in 0 runs" and the old links stayed behind as orphans forever. The trash now remembers where a note came from; emptying it clears the links too.
  • The update installer is no longer reported as "nothing was installed" while it is still running. After 180 seconds the request gives up waiting - but if the download had finished and the installer had already started, cancelling is no longer possible, and the app may well update itself minutes later. It now says so, and tells you not to start the update a second time.
  • "Sign in once" no longer sends you after a run that does not exist. If the browser starts and closes again right away, the likely reason is your own sign-in window from an earlier press: the page opened as a new tab in it. The message names that first instead of claiming nothing was opened.
  • The browser start path says when it could not hand over its port. If the server does not answer within 30 seconds, the port file is missing - and a second agentfenster run would quietly start its own server. That used to happen without a word.

The guided tour actually comes back, and an empty trash says so

  • If you ended setup with "Run it", the tour never appeared, not once. The first-run assistant deliberately holds the tour back while your very first task is running: its overlay would swallow the "Allow" click of the first question, and that run would die after 120 seconds. The code promised the tour would then "come on the next start", but nothing ever asked for it again, because the only two callers lived inside the setup assistant itself. The same was true for "Show me", the button into the feature overview. Both paths now get the tour once, on the next start, and only when no run is in flight.
  • "Show trash" no longer renames itself and shows nothing. With an empty trash, the normal state once you have emptied it or restored everything, the button under "Memory was cleaned up once" flipped to "Back to the notes" and left the screen otherwise untouched. It now says the trash is empty.

A stranger finds Voice and the thinking level without the help button

  • The microphone button was invisible until you already knew it existed. With voice input off (the default), the button is hidden, correct, a button that would do nothing is worse, but that also meant a stranger never learned it was there without opening the help card first. A quiet line under the task box now says "You can also talk: turn on voice in Settings.", only in the empty chat, only while voice is off. It disappears for good after the first click, or the moment voice is switched on, never during a run, never above the license card or the "Heard: …" line.
  • A fresh install ran on Sonnet with medium thinking, and said nothing about it. The one setting that cost a hard evening (thinking hard-wired to the cheapest level) was invisible to a new user until they wandered into Settings on their own. The Brain step of setup now carries one sentence next to the model picker, Medium by default, right for most tasks; High is slower and costs more per step, and is worth it for a screen that keeps failing, with a link that jumps straight to Settings → Brain.
  • Three of the twelve "What agentfenster can do" cards were the ones a stranger paused longest on, per the seventh acceptance pass: "Speed" read as agent speed, not machine load, so it is now "Live picture, or how much of your machine it may use"; the Memory card's when now says when the saving actually starts ("From the second time you use the same app…"); and "Brains" is now "Brain, which model decides the next click", with the sentence leading with "model" instead of the in-house word "brain".

Smaller things that were quietly wrong

  • The window came back at the wrong size on a scaled second screen. Where the app was when you closed it was remembered in real screen pixels, but converted back using the main monitor's scaling, so a window left on a 150 % second screen returned half again too large. It now asks the monitor the window actually stood on.
  • Reading the desktop could hand the model the wrong window. With two windows open under the same title, read_desktop without a window argument silently read the front one, and the title in its header did not tell them apart. The answer now says how many windows share that title, which one it read, and the handles of the others.
  • Windows error numbers reached the model in German. Only "the window is gone" had been translated; every other Windows failure arrived verbatim, in the system language, with no next step, "WinFehler - (5, 'CreateDesktopW', 'Zugriff verweigert')". Known Windows errors are now one English half sentence, the same wording the app itself uses.
  • Handing off an older recording said it had been deleted. The hand-off looked only at the newest 150 recordings; anything below that, though plainly visible in the list, produced "That recording is no longer in the list." It now loads the rest before claiming anything.
  • Replaying a long recording kept growing. Every frame the player preloaded stayed in memory for as long as the view was open. The cache is capped now; the frames around the playhead are never dropped.
  • The window icon could land on the wrong window. It was looked up by title, and two windows of this app can carry the same one. It now goes to the window's own handle. The mini window's handle lookup also used a 32-bit conversion that fails outright on large handles, and failed silently.

A guest on your Wi-Fi can no longer take the phone remote away from you

  • Somebody without any key could lock your phone out for five minutes at a time. The counter that stops key guessing was shared by every device, so 119 wrong attempts from a neighbour's laptop locked out your paired phone, including its Stop button, which is exactly the button you need when the task you started from the sofa goes wrong. Measured, then fixed: the counter now belongs to the device that guessed. It locks itself out; your phone keeps working.
  • The same trick took the live picture away. 240 requests for files under /web/, even files that do not exist, used up a read budget shared by everyone, and the phone's picture stopped with "too many requests". That budget is now per device too.
  • The phone key no longer travels in the clear when encryption is on. Browser cookies ignore port numbers, so the key set on the encrypted connection was still attached to any plain-http request to the same machine. It is now marked Secure whenever the connection is encrypted.

A task sent from your phone stops before it deletes, sends or pays

  • New switch, Settings → Phone → "Let the phone start tasks." On after pairing, a remote that cannot send anything is not a remote, and off it leaves watching, answering, tapping into the picture and stopping a run untouched.
  • The limit under it is not a switch. A task sent from a phone now refuses any step that looks like it deletes, sends, pays or discards something, the same way a scheduled night run does, and says so instead of doing it. You are not looking at the screen where it happens, and a purchase cannot be undone. With "Ask first" nothing changes: the question reaches your phone and you answer it there.
  • docs/SECURITY.md has a new section, "Your phone as a remote", with the five routes a device on your Wi-Fi could try and what stops each one.

Waiting for a file, a window or a word is now one step, not one model call per poll

  • The agent used to poll by hand. Exporting the System Information report (msinfo32 /report) takes one to three minutes. Until now the only wait the agent had was wait N seconds (capped at 10 s), so it filled that time with wait 10, list folder, wait 10, read file ... - and every one of those rounds was a separate model call. Measured on the test bench task A58 with opus/high, three runs on the old code: 16 to 20 of 23 to 27 steps per run were nothing but waiting. Telling the model "you have waited three times, batch it" fired 18 times and changed nothing.
  • Now the wait carries its condition. {"art":"warte","bis":{"datei":"<path>"}, "hoechstens_s":300} - or fenster / fenster_weg (a window with that title is open / gone), element (a control with that name is in the tree), text (the words appear anywhere in the window). agentfenster checks every 0.5 s, stops at the cap (60 s by default, 10 min at most), keeps the stop button live, and reports what it saw: "reached after 12.4 s: file '...' exists (2 747 648 bytes, unchanged for 1.0 s; it appeared during the wait)" or "not reached after 60.0 s - last seen: no file at '...' - the folder holds: ...". One model call for the whole wait.
  • A file only counts when it is finished. "Exists" was not good enough: the agent likes to create a placeholder first, and a half-written report is not a report. The condition is met when the file appears or changes and then stops growing for a second; a file that was already there and never changed is reported exactly like that.
  • The result is in the next prompt (like a folder listing), a wait that timed out stops a batch that was planned on top of it, and a condition the model got wrong ("unknown wait condition 'foo'") is refused immediately with the four allowed ones - not after 60 s of waiting for nothing.

agentfenster starts a little faster and uses less memory

  • Every start now skips loading a numerics library it needed for one line of code. Deciding whether a captured window image is completely black used to pull in numpy - 220 of the 527 milliseconds it took to load the recording code, on every single start, for one comparison. Pillow, which is loaded there anyway, answers the same question. Measured over five starts each: the application code now loads in 1 607 instead of 1 829 milliseconds and the process holds 77 instead of 87 MB. That applies to every way in - the window, --headless, the MCP server and each command line call.

Maximising on a second screen stays on that screen

  • A maximised window no longer spills over a high-DPI second monitor. On a 2560x1600 screen running at 150 %, maximising made the window 3200x1920 points - 640 wider and 392 taller than the monitor itself, so it ran off three edges and covered the taskbar. Windows proposed that size on its own, and answering WM_GETMINMAXINFO correctly did not change it; the window now clamps itself to the work area of the monitor it is on. Snapping to half a screen and restoring from maximised are untouched, and nothing changes on a 100 % screen.

A file order across three programs, worded the way a user types it

  • A63 is the whole chain in one German sentence: pack the three newest files from this folder into `neu.zip` next to them, and write me a list with name and size in Notepad, saved as `liste.txt`. Seven files of different age and size, two with umlauts, one with a space - and the three newest are deliberately neither the biggest nor the first in the alphabet. Judged on files only: the archive holds exactly those three (checksums), the list names all three with a size (rounded is fine), the originals are untouched, nothing is left behind, and the closing sentence is in the language of the order. With Opus on high: three passes in 15, 5 and 11 steps; the cleanest run took 43 s.
  • A German answer to a German order is no longer framed as a mistake. Every German closing note used to come back as "The model wrote its closing note in German instead of English" plus a warning - eight out of eight runs of the last two packages. The frame now appears only when the order itself was English.
  • Writing a text file with Windows line endings works. A model that sent \r\n got a file with doubled carriage returns and the verdict "does not contain what was meant to go in - treat it as not done" - for a file that was there. Line endings are normalised before writing; the read-back matches.
  • A window that had already closed no longer costs a warning. A one-line PowerShell command is finished before the step picture is taken; the recording logged (1400, 'GetClientRect', ...) twice per run as a warning. A window that is still there and cannot be captured stays a warning.
  • Measured limit, not fixed: selecting several files in File Explorer without a real mouse. Posted shift+down does not extend the selection and a drag draws no rubber band - Sonnet on medium spent 90 steps there and gave up on the Explorer route; Opus never needed it (it reads the folder listing and packs with PowerShell). The way in is UI Automation's AddToSelection.
0.3.82026-09-07 (morning, second build)

Everything merged after the 0.3.7 build (05:12) and before the 0.3.8 build (08:09).

Three seams between last night's merges

  • A recording that leaves your windows minimized now says so where you can see it. While a recording runs, agentfenster's main window is out of the way and the only visible Stop button sits in the small recorder overlay. That overlay was handed the sentence "N of your windows stayed minimized after the recording - click them in the taskbar to bring them back." and threw it away; it showed "Stopped" either way. Every answer of the stop endpoint carries it now - not two of five. (Where it actually reaches you was put right in 0.3.9: the overlay window is gone by the time the sentence exists, so the main window shows it when it comes back.)
  • A German task about "dem Rechner" is no longer refused as the Calculator. agentfenster warns the model before the first step when a task names an app it cannot drive. The word list matched "Rechner", which in German means machine far more often than calculator - so any German task mentioning the computer was told, in the planning prompt, that it might have to be cancelled. The word still blocks a start when the agent types it as a program name; it no longer fires on a whole sentence.
  • A routine "every 6 hours" no longer runs at every program start. Since the interval clock restarts at the end of a run, a run that died with the app - you closed agentfenster, Windows restarted - left the old due time behind, and that time was already past. The next start ran the routine immediately, however long the interval. An interrupted run now restarts the clock, the same as a finished one. Weekly appointments are untouched.

The suggestions on the start screen now actually run

  • Every ready-made task has been run for real, and every one of them passed. Nine tasks are offered instead of twelve. Four of the old twelve told the agent to use Notepad, and on Windows 11 notepad.exe is redirected to the Microsoft Store app, which always opens on your desktop and takes no click the agent may send there. agentfenster's own warning said so before the first step, and pointed at write.exe, which 24H2 removed. So the app was offering you tasks it had already declared impossible. The three that had no Win32 way at all are gone (the Notepad note, print-to-PDF from Notepad, the Snipping Tool screenshots); the rest name a program that is on every fresh Windows 11. A test now holds each suggestion against that wall so it cannot happen again.
  • The three suggestions shown first are the three best-measured ones, not the first three in the file: zip a file from the right-click menu (26 s, 5 of 5 runs passed), read a value off a web page into a file (36 s, 5 of 5), copy a special character into a file name (80 s, 3 of 3). The order is checked against the real test-bench reports, so a card cannot claim more than was measured.
  • Two suggestions that had never been measured now have a test-bench task of their own, exporting system information and driving the Character Map, and three passing runs each. No card carries "not measured yet" any more.
  • The first task a new user is offered is no longer `open notepad and type hello`. It is open the character map and find the euro sign, it lives at exactly one place in the code now, and both the first-run assistant and the product tour take it from there.

Scrolling works where the agent actually scrolls

  • The wheel now turns inside a web app. Discord, WhatsApp Web, Spotify - anything built as a single page - do not scroll the window; they scroll a box inside it. agentfenster sent the wheel through the browser correctly and then looked at the wrong thing to see whether it had worked, so a scroll that had landed was reported as "scrolling cannot be delivered without real mouse input here". It now watches the box that actually moved. In your run of 6 September this cost three of the thirty-four steps.
  • Naming what to scroll no longer takes away the only way to scroll it. Saying "scroll the chat list" used to skip the browser route entirely and fall back to a Windows message that such a window never receives. The wheel now goes to the named element, at the place where that element sits.
  • Discord can be scrolled at all. Apps that carry their own Chromium keep their debugging port next to the workspace. Clicking already looked there; scrolling looked in the agent browser's folder instead and found nothing.

The very first task a new user starts now survives the tour

  • The product tour no longer covers the first run. Finishing setup with "Run it" started the sample task and, in the same second, laid the tour's half-transparent overlay across the screen. Six seconds later the agent asked for permission, as it should - but the "Allow" button under that layer took no click. Two minutes on, the first minute anyone ever spends with agentfenster ended in "Step 1 was not confirmed - stopped.", reproduced twice. The tour now stays out of the way while a task is running; it is not lost, it opens by itself at the next start. Measured afterwards: the button is clickable after eight seconds, one click, and the Character Map opens - two steps, 18 seconds.
  • The licence card no longer claims you have not accepted. After setup it still sat in the chat saying nothing could run until the licence was accepted, one second after you had accepted it in the assistant. It reads the real state now instead of asserting one.
0.3.72026-09-07 (night)

Two things a session report and a stray local process could reach

  • Where you are signed in stays on your machine. After "Sign in once for the agent", the agent's answer lists the sites its browser profile has a session for - it needs that to tell a login wall from a missing account. That answer also travelled in the session report. A report now carries the count only ("Signed in here: 2 sites."), never the site names and never the profile folder, which contains your Windows account name. The full sentence stays in the local log, where only you read it.
  • Nothing pops a window onto your screen without the app. Opening a browser for "Sign in once", installing an update and switching the console host now always require the session key of the open agentfenster window - before, a headless agentfenster (the one behind Claude Code) never asked for it, so any other program on the machine could have made a browser window jump into your face. Reading is unchanged.

The chat gets the room a conversation needs

  • When nothing is running, the stage steps aside. On a 1280x800 window the chat column was 260 pixels wide and showed nine lines, while the stage next to it filled 660 pixels with a picture of an empty desktop and the activity column showed "No run yet". Both of those only have something to say once a run is going. Until then the stage shrinks to a single line above the terminal, "No run yet, the stage opens by itself when one starts.", the activity column folds away, and the chat takes the space: 538 pixels wide and 15 lines at the same window size, measured in a real browser. The terminal keeps its full width, so Claude Code still has its 80 columns; it gains the whole height instead.
  • The stage comes back the moment a run starts, in 150 ms, and it respects "reduce motion". You can also open or close it yourself at any time, even while a run is going, and that choice is remembered. From 1600 pixels of window width the three columns stay as they were, there is room for all of them.
  • The workspace list stands in two columns while the chat is wide, so the shorter card still shows four workspaces instead of two.
  • A hairline at the top of the history says when there is more above it. The chat opens at the end of the conversation, as it should, but nothing told you that the sentence at the top was cut off rather than starting there.

The first task a new user starts actually runs

  • "Open notepad and type hello" is gone, the first task is Character Map now. Two changes from the same night collided: setup ended with "open notepad and type hello", and Notepad had just been added to the list of apps the agent is told it cannot drive. So the very first thing a stranger ever started was announced as impossible before the first click, with a way out (write.exe) that Windows 11 removed years ago. Setup, the feature cards and the tour now take the same sentence from one place: "open the character map and find the euro sign". Measured against a fresh install in a browser: setup clicked through, "Run it" pressed, and exactly that task reaching the server.
  • The wall of apps it refuses is measured again, app by app. Notepad, Paint, Snipping Tool and Phone Link came off it: they are packaged, but they start as ordinary programs (Windows.FullTrustApplication in their manifest), so they do open on the agent's desktop, and two real runs prove they can be driven there. Calculator, Clock, Camera, Sound Recorder and the Store stay, because they really are launched by Windows onto your screen. WhatsApp and Teams also stay, but the sentence now names the reason that was actually measured instead of a reason that was not.
  • "High" stays high. With Speed set to "Smaller model for simple tasks", a task that named one program pushed thinking from high straight down to low, two rungs, against the rule written in the same file, while the settings screen went on showing High. It now drops at most one rung, only from medium, and the Speed option says so.
  • A repeating task waits for the last run to finish. "Every 10 minutes" counted from the moment a run started, not from when it ended. A run longer than its own interval was therefore due again the second it finished: back-to-back forever, with "Next run 22:41" on screen while the 22:41 run was still going. The clock now starts at the end, exactly as the form, the README and the module have promised all along.
  • "Sign in once" no longer says "opened" when no window came up. One browser profile can only be open once. If the agent's browser was already running, a run paused at a login wall, say, the new window was handed to it and appeared on the invisible desktop, while you got a green "Brave opened with the agent's profile". It now checks first, and checks that the program it started is still alive a moment later; otherwise it says what to do instead.
  • The phone's fingerprint card stopped grading its own homework. It asked you to compare the value on the phone with the value in Settings on the PC, both come from the same file, so they always matched and the check proved nothing. It now asks for the one comparison that does prove something: the certificate details your phone's browser shows in its warning.
  • Smaller, same night: the chat header now updates when a run switches to a bigger model mid-way instead of claiming the old one for the rest of the run; undoing the console setting checks that Windows really took it back; a window that stayed minimized after a recording says so on screen instead of only in the log; the memory view keeps its cleanup note when every note is in the trash, and the note no longer claims nothing was deleted.

Four things that quietly said the wrong thing

  • "Show trash" now shows the trash. After a memory cleanup the note above the list offers a button to see what was moved out. It did nothing at all, the screen was character for character the same before and after the click. It opens and closes the trash list now, and says which way it points.
  • A run no longer claims to think harder when it cannot. Only Claude Code has a thinking-level switch. When a run got stuck and escalated, it still told you "Raised thinking from medium to high" on Codex, Gemini, Grok, Kimi and any OpenAI-compatible endpoint, where that setting changes nothing. It stays quiet there now, exactly as the setting itself already did.
  • A finished run is no longer reported as failed because a dialog had a question mark in its name. Save changes?, Discard changes?, the agent clicks those all day. If it quoted one in its closing note, the run was marked "did not finish" and paid for one more model call. A question mark inside quotation marks is the dialog talking, not the agent.
  • An update step that was renumbered while it was being built now runs anyway. The database remembers which update steps it has run by number. If two packages claim the same number and one of them is moved, a file that ran the old one skips the new one, silently, until something reaches for a column that is not there. agentfenster now checks that the tables and columns its plan promises really exist, catches up what is missing and writes a line about it in the log. Measured on a copy of a real database: nothing to catch up, all 13 tables unchanged, 3.5 ms.

Speaking a task now really runs it, and says so before it does

  • What was heard is on screen before anything happens. After you speak, a line above the task bar says Heard: "open the character map and search for the euro sign." · 1.2 s, check it, then press Enter to run it. The sentence sits in the task bar as normal text; one Enter (or the Send button) starts the run, and the line disappears. Nothing is ever sent off on its own, measured: in the whole path from microphone to task bar, not one request reaches the run endpoint until you press a key.
  • A cough no longer becomes a task. Speech recognition turns silence, throat-clearing and a slammed door into a lone ".", that used to land in the task bar where it was easy to miss and send. Anything shorter than three letters or digits is now refused with a sentence: Only "." came back, which is too short to be a task, nothing was put into the task bar and nothing was started.
  • "No microphone" now says "no microphone". Every way the microphone can fail has its own sentence and its own next step: none connected, access refused, in use by another program. The old text always blamed permissions. On the agent desktop there is no audio device at all, and the message says so instead of looking like a crash.
  • Ten seconds, not sixty. Transcription used to be allowed to run for a minute before giving up, a minute of nothing for a sentence that took three seconds to say. The limit is 10 s now, measured: the largest recording the window allows takes about 1.7 s with the base model, so there is six times the headroom, and going over says why and what to change.
  • Measured end to end, in both languages, on the real agent desktop: "open the character map and search for the euro sign" → 1.2 s to the text, Enter, and the agent had opened Character Map and typed "Euro" 22.1 s later. German the same way in 25.2 s. English came back word-perfect on the test corpus, German does not: Whisper splits compounds ("Zeichen Tabelle"). The full sheet is pruefung/VOICE-ECHT-2026-09-07.md.

Fixed: things that only broke when something else went wrong

  • Your other windows come back after a recording, even when one of them was closed while it ran. agentfenster clears the desktop before a recording and brings everything back afterwards. If a window was closed in between, it could no longer confirm that clearing had worked, and gave up on bringing anything back, although it had already minimized everything. Every window stayed in the taskbar until you clicked it. It now separates "the desktop was cleared" from "we could measure it": the way back only needs the first. The report stays honest and still tells you when it could not measure.
  • agentfenster comes back after a recording even if its own window went away. Restoring the main window to its earlier size and position could throw, and then the window was never shown again.
  • A second agent desktop no longer starts with a browser profile it cannot read. Passing a signed-in browser profile on to a fresh profile copies both the cookie file and the key Chrome encrypts it with. If only the cookie file made it through, agentfenster reported success and the run then stood in front of a sign-in page it could not explain. Half an inheritance now counts as none, and says so.
  • A run no longer dies from trying to think harder. When a run raises its thinking level mid-way it closes the old thinking session; if that session hung, the error ended the whole run.
  • A broken window position or a profile folder that is really a file are now refused with a sentence instead of crashing.

The update button finally does what it says

  • Updating no longer says "done" while nothing happens. Measured on a fresh Windows in a sandbox: the app answered ok after 12 seconds, and seven minutes later it was still running the old version. The installer was standing behind everything else, waiting for someone to answer a Yes/No question about the Edge WebView2 runtime, a question the app never told it to skip. It skips it now, and if the installation is refused for any other reason, the app waits for the real answer instead of looking away after three seconds.
  • "Up to date." is only said when someone actually looked. The app checks the update server once a day and remembers that on disk. A freshly started app therefore did not look at all, and still reported everything was current. It now says when the last look was and offers to look again.
  • When the installer cannot replace the files because agentfenster is still running, it says so, with the path of the file you already downloaded, so you can finish the update by hand in one double-click. It used to report "stopped straight away", which was neither true nor useful.
  • Starting without the Edge WebView2 runtime no longer hides the app from its own command line. Without that runtime agentfenster falls back to the browser interface, and in that mode it never wrote down which port it was listening on. agentfenster run "…" therefore could not attach to the window you had open and started a second server beside it. It writes the port down now, once the port really answers, and never over another running instance's.

A batch of actions no longer dies on a keystroke

  • A key press that changes nothing visible no longer throws away the rest of a batch. The agent may plan up to five actions for one thinking step; the batch stops as soon as one of them turns out not to have worked, which is right for a click that missed its target. A key press is different: it moves focus, a selection or the caret, and none of that appears in the window tree at all, "nothing changed" is the normal answer there, not a failure. Measured over 358 recordings: one batch in four ended early, and 25 of 102 planned actions were thrown away, most of them behind a key press. Nothing else changes, the step is still recorded as having had no effect, still counted by the repeat, circle and wall guards, and a refused key or an unverified one still stops the batch.

Three seams between last night's merges

  • A recording that leaves your windows minimized now says so where you can see it. While a recording runs, agentfenster's main window is out of the way and the only visible Stop button sits in the small recorder overlay. That overlay was handed the sentence "N of your windows stayed minimized after the recording - click them in the taskbar to bring them back." and threw it away; it showed "Stopped" either way. It shows the real sentence now, and every answer of the stop endpoint carries it - not two of five.
  • A German task about "dem Rechner" is no longer refused as the Calculator. agentfenster warns the model before the first step when a task names an app it cannot drive. The word list matched "Rechner", which in German means machine far more often than calculator - so any German task mentioning the computer was told, in the planning prompt, that it might have to be cancelled. The word still blocks a start when the agent types it as a program name; it no longer fires on a whole sentence.
  • A routine "every 6 hours" no longer runs at every program start. Since the interval clock restarts at the end of a run, a run that died with the app - you closed agentfenster, Windows restarted - left the old due time behind, and that time was already past. The next start ran the routine immediately, however long the interval. An interrupted run now restarts the clock, the same as a finished one. Weekly appointments are untouched.

The suggestions on the start screen now actually run

  • Every ready-made task has been run for real, and every one of them passed. Nine tasks are offered instead of twelve. Four of the old twelve told the agent to use Notepad, and on Windows 11 notepad.exe is redirected to the Microsoft Store app, which always opens on your desktop and takes no click the agent may send there. agentfenster's own warning said so before the first step, and pointed at write.exe, which 24H2 removed. So the app was offering you tasks it had already declared impossible. The three that had no Win32 way at all are gone (the Notepad note, print-to-PDF from Notepad, the Snipping Tool screenshots); the rest name a program that is on every fresh Windows 11. A test now holds each suggestion against that wall so it cannot happen again.
  • The three suggestions shown first are the three best-measured ones, not the first three in the file: zip a file from the right-click menu (26 s, 5 of 5 runs passed), read a value off a web page into a file (36 s, 5 of 5), copy a special character into a file name (80 s, 3 of 3). The order is recomputed from the real test-bench reports, so it cannot drift.
  • Two suggestions that had never been measured now have a test-bench task of their own, exporting system information and driving the Character Map, and three passing runs each. No card carries "not measured yet" any more.
  • The first task a new user is offered is no longer `open notepad and type hello`. It is open the character map and find the euro sign, it lives at exactly one place in the code now, and both the first-run assistant and the product tour take it from there.

Scrolling works where the agent actually scrolls

  • The wheel now turns inside a web app. Discord, WhatsApp Web, Spotify - anything built as a single page - do not scroll the window; they scroll a box inside it. agentfenster sent the wheel through the browser correctly and then looked at the wrong thing to see whether it had worked, so a scroll that had landed was reported as "scrolling cannot be delivered without real mouse input here". It now watches the box that actually moved. In your run of 6 September this cost three of the thirty-four steps.
  • Naming what to scroll no longer takes away the only way to scroll it. Saying "scroll the chat list" used to skip the browser route entirely and fall back to a Windows message that such a window never receives. The wheel now goes to the named element, at the place where that element sits.
  • Discord can be scrolled at all. Apps that carry their own Chromium keep their debugging port next to the workspace. Clicking already looked there; scrolling looked in the agent browser's folder instead and found nothing.
0.3.62026-09-06 (night)

The chat column looks like a tool now, not a form

  • The empty chat has a centre. It used to open with a grey half-sentence in the top left corner and 190 px of nothing below it, a third of the column saying neither what to do nor what mattered. It now opens the way the Activity column beside it always did: a mark, one question, "What should the agent do?", one sentence, and up to three suggestions right underneath, all of it in the middle of the column. Nothing scrolls; measured at ten window heights from 700 to 1440 px.
  • A conversation reads like a conversation. What you said sits on the right in a quiet bubble, what the agent answered on the left as plain text, and the thinking log and the waiting plan are cards inside the thread instead of bands running across it. Your own line used to look like a second input field.
  • One field to write in, instead of three stacked bars. Autonomy, Permission and the task box were three strips with three separator lines on the last 170 px of the column, none of which separated anything. They are one raised field now. Both settings stay visible without scrolling and are still one click away, they are just quiet: 12 px grey labels instead of bold white ones that shouted louder than the agent's answer.
  • "Permission mode" is now "Permission": in a 258 px column the longer label pushed "Ask first" onto a second line. The full name is still there for screen readers.
  • Four task suggestions became three plus "More…", four filled five lines in a narrow column, which is a list to read rather than three ideas to glance at.

The first five minutes: it now explains itself

  • New: "What agentfenster can do". Twelve cards, one per feature, each saying in one sentence what it does and in one sentence when you would want it, voice, your phone, skills, schedules, memory, brains, work style, speed and permissions. Every card has a button that takes you straight there, and where it can it sets the first step up for you (the task card puts a harmless example in the box; the others open exactly the right settings section and highlight it). No card starts a run on its own.
  • Three ways in, and one of them needs no prior knowledge: the question mark next to the tabs, Settings → Introduction, and the last screen of the first start.
  • The first start went from eleven screens to five. It now asks only what cannot be skipped, which model thinks for it, what Windows will not let it do, the licence, and then runs one harmless task while you watch. Work style, autonomy, reply language, the app lists, the permission switches and the memory setting are no longer questions in the way; they are cards you can try one at a time, and their defaults are the careful ones. Measured on a fresh install: six clicks and 2.3 seconds from a blank profile to a task actually running, against eleven screens before.
  • A settings section you jump to now really appears at the top. The jump used to count "did the page move at all" as success, measured, it moved 761 pixels and left the card you asked for 5799 pixels further down, and you were looking at a different section with no error anywhere. It now checks where it landed and corrects itself; if it still cannot get there, it says so.

No stray console window on your screen

  • Windows, not agentfenster, decides which terminal opens a console window, and on Windows 11 that is Windows Terminal, a single process that lives on your screen. Every console an agent opened on its own desktop therefore popped up on yours. Settings now has a Console windows card that says this in one sentence and offers a single button to make Windows Console Host the default instead, with Undo right next to it. It only appears when your machine really does delegate elsewhere, it never changes anything on start-up, and it touches nothing but those two values under your own user account.
  • agentfenster no longer opens one itself. Starting the app from the command line spawned a background instance that got its own console, and that console went straight to Windows Terminal, on your screen, for no reason. It now starts without a window at all.

Recording a skill: a clear stage before, the same window after

  • Starting a recording now clears the desktop, not just the agentfenster window. Your other windows go down to the taskbar exactly the way Show desktop does it, through the Windows shell, with no keystroke faked on your behalf, so the recording starts on an empty screen and everything you open from there belongs to it. They all come back when you stop; anything you opened during the recording stays open. The card in front of the button says this in one sentence before you press it.
  • The window comes back exactly as you left it. Until now a maximized window came back at some other size, because the way back forced it to "normal" and two further corrections then argued about the number. Measured at the real window on all three cases, normal, maximized, and small on the second monitor, the size, the position, the monitor and the maximized state now match to the pixel.
  • Windows that cannot be minimized at all (dialog boxes) are counted and named while you record, "Windows kept 1 window(s) on screen", instead of the app pretending the desktop is clear.

A workspace names itself after your first task

  • A workspace is no longer called "New workspace" once you have said what to do. The first task names it, "Zip the folder", "Open Discord", in the language you typed it in, and both the list on the left and the header above the chat show the new name immediately, without reloading.
  • The name is worked out from your own words, without asking a model: it is there the moment you press Send, not twenty seconds later, and it costs nothing. Small talk ("hi", "thanks") and commands (/goal, /loop) name nothing, a workspace called "Hi" would be worse than one still called "New workspace".
  • Only the first task names it. Every later task leaves the name alone, and a name you typed yourself is never overwritten, not even if you deliberately called your workspace "New workspace". Renaming by clicking the name still works and always wins.

The agent can now be signed in to your sites, once, by you

  • Settings → Browser → "Sign in once for the agent". The agent has always browsed in its own profile, never yours; nobody could ever sign in to that profile, so every task about "my account" hit a login page. Two tasks died on it on 6 Sep, and the agent's answer, "no account is logged in to X", read like a statement about you. The button opens the browser you pick, with the agent's profile, on your screen: you type your password yourself, with your own hands. No model sees it, no agent types it, and no remote-control port is open while you do.
  • The card shows what is really in there, "x.com (9 cookies, valid for 199 more days)". Read from the cookie file's plain columns (host, count, expiry); the cookie values are never touched.
  • The agent no longer says "you have no account". When it opens a browser it is told which sites this profile is signed in to, and for a page it is not signed in to it now says exactly that, plus where to fix it, instead of reporting a login wall as a fact about you.
  • A new agent desktop inherits the sign-in from the first one. Signing in on desktop 1 and finding a QR code on desktop 3 (measured 6 Sep, 08:16) is over: a fresh profile is seeded from the shared desktop's profile. Sessions and the key to read them travel; saved passwords and payment data never do. That a copied profile keeps its login was measured, not assumed, real Brave, Chrome, Edge and Firefox against a local login server.

One bar at the top, not two

  • The Windows title bar is gone; the app's own header carries the window buttons. Minimize, maximize/restore and close now sit at the right end of the same bar as the wordmark, the version and the tabs Agent, Recordings, Skills, Memory and Settings - at the same height, as one band. Measured on a 1280x800 window: the old build put 31 extra pixels of Windows title bar above that header; now the interface starts flush at the top edge of the window.
  • Nothing the Windows title bar could do was lost. Dragging the header, double-click to maximize, Aero Snap, all eight resize edges, minimize and restore from the taskbar, the system menu on Alt+Space and the window position surviving a restart all still work - measured at the window itself on both screens, at 100 % and at 150 % scaling.
  • Maximized, the interface now reaches every screen edge. It used to stop 7 points short of all four of them, in a colour that was not ours.
  • The small live window got the same treatment, for the same reason: it has its own header too, so it no longer draws a strip of Windows frame in front of it.
  • Known limit, on purpose: hovering the maximize button does not open Windows 11's snap-layout flyout. Claiming that button for Windows would take hover, keyboard focus and the tooltip away from it. Snap layouts are still there via Win+Z and by dragging the window to a screen edge.

The agent thinks as hard as the task needs, and you can see what it thinks with

  • How hard the agent thinks per step is now a setting (Settings → Brain → Thinking: low / medium / high, default medium). It used to be a constant in the source, hard-wired to low earlier the same day, a value measured on five easy bench tasks, which every hard task then paid for. Only Claude Code has a switch for this; for every other brain the card says so in a sentence instead of pretending the choice does anything.
  • A run that has to look for another way now thinks harder for the rest of it. The same triggers that already pulled the strong model back (going in circles, more than ten steps on a task that looked short) now also raise the thinking level to high, once, with one sentence saying so, never silently.
  • A task that looks simple may think one level cheaper, but only if you set Speed to something other than "Always strong". The estimator is only ever allowed to spend less, never more.
  • Every step in the activity column now says which model and which thinking level produced it (small, monospace, e.g. opus · medium), and the chat header says it once when a run starts: "Clicking with opus, thinking medium." Both matter because both can change mid-run.

A routine can now repeat all day, and it waits its turn

  • New trigger "Every so often", every 10 minutes, every 2 hours, anything from 1 minute to a full day. Until now a schedule could only run on a weekday at a time of day, when a file arrived, or when the app started; something that repeats within a day could not be set up at all. The clock starts again when a run finishes, so a long run never causes two at once.
  • Two routines that fall due together no longer race. Only one run may click on the agent desktop, and the one that has waited longest now goes first, before, the schedule you added most recently always won, and the oldest one was pushed back again on every tick. The one that waits keeps its place in the queue instead of being sent to the back.
  • The line that says why a routine is waiting tells the truth at night. It used to read "the schedule never clicks while you are working" even when nobody was working and the other run was itself a scheduled one.
  • A routine that ran into a wall is no longer reported as a refused step. A task that looked for a button in a window that does not exist used to end up in your inbox under "a step must not run while nobody is watching", with the advice to "change it so it does not delete, send or pay", sending you to look for a dangerous step that was never there. It now says it ran into something it could not get past, and the advice points at the task text.
  • Turning an "every so often" routine off and on again keeps its interval. Without this the Turn off button would have been refused by the server.

No step of the introduction needs a click nobody can guess

  • "Next" now moves the guided tour forward in every single step. Three of its eleven steps had no Next button at all: you had to click the stage, send a real task, or click into the shell to get past them, and nothing said so. Measured before the fix against a fresh install: someone who only presses Next was stuck on step 2 of 11.
  • Where a step still offers something to try, the card says so as a sentence, "Click into it to put your cursor there, or press Next to skip this for now", and pressing Next is honest about it: the following card carries one quiet line starting with "Skipped: …" instead of moving on as if it had happened.
  • Escape closes the tour instead of asking a second question first, the close cross is back in the top corner of every card, and either way the app remembers, it will not push the tour at you again, and one sentence tells you it lives in Settings → Introduction.
  • "Skip setup" is now also at the top right of the setup card, where every window keeps its exit. It used to sit alone in the bottom-right corner (at x=913 of a 940-pixel-wide card), which is the last place a glance looks.

A rejected task no longer disappears

  • The task you typed stays in the chat, and the refusal is the agent's reply right underneath it. Until now every refusal before the first step (license not accepted, no brain, no free desktop) took your own line back out of the chat and left only a toast and one line in the Activity column. On 06.09.2026 that made a real task look like it had vanished, the same order was typed seven times in 26 minutes and the chat never said a word back. The server stores the refusal in the conversation now, and the window shows it exactly once (chat, not chat and toast).
  • Activity no longer claims "Task received" when nothing was received; it says the agent answered in the chat instead.

The licence is offered where you are, not two screens away

  • A quiet card above the task bar, one sentence, Read the agreement, Accept. One click accepts, the card disappears, and the task you just typed starts on its own. The old refusal told you to "open Setup again", which is no route at all for an installation that finished onboarding weeks ago.
  • Installations from before the licence gate are asked, not signed up. There is deliberately no migration that flips eula_akzeptiert to true for them: an unasked "yes" is not consent. They get the card at the next start.
  • No planning turn is paid for a run that cannot start. The licence check now sits in front of the model call, not behind it.

A job against a wall Windows put there now ends in step 1, not step 34

  • A run that names WhatsApp Desktop, Teams or another Store-only app now says so before it spends a single step. Measured on a real run of yours (20260906-202353-9a11): the order was "open Discord ... then write on WhatsApp in the Alles-Wichtige group", and the run took 34 steps and 17 warnings to discover what had been in the log since that same morning, WhatsApp Desktop is a WinUI app, and a WinUI window accepts neither a posted click nor a posted key. agentfenster now recognises the app from the name you typed (WhatsApp, whatsapp.exe, whatsapp:// alike), warns before the first model call, and refuses the start with a sentence that says what the app is, why it is useless here, and where to go instead: "use the web version, say 'open web.whatsapp.com in Edge' and agentfenster can click and type there."
  • It is a warning, not an abort. That order had two halves, and the Discord half worked. Throwing the whole run away over the second half would turn half a result into none.
  • `open 'WhatsApp'` no longer lies. It used to answer "not on PATH, not registered under App Paths, not in the Start menu", three times in that one run. WhatsApp is installed; it simply has no .exe a process start could reach. The new sentence says that.
  • If the wall is only noticed mid-run, the first refused move ends it, not the seventeenth. The loop's circle guard cannot catch this by design: it counts (state, action) pairs, and each of those 17 refusals hit a different element, so no move was ever a repeat. At a window that accepts nothing, the first refusal is already final. The silent case counts too: a key that PostMessage accepts ("ok") while nothing happens is the same wall.
  • Chromium windows, games and classic Win32 windows are untouched by this, and so is the run if you have switched on Use the real mouse and keyboard.

The closing sentence is written for you, not for the log

  • The result of a run is no longer the model's raw text. That same run ended with "FULLVIEW, Chat geöffnet? Ich brauche das Nachrichtenfeld.", German, a question, carrying an internal planner switch, and it was filed as a success. Now: the switch words (FULLVIEW, PICTURE) are stripped, a German note is quoted unchanged inside an English frame (translating it would be a second claim about a run only the model saw), and a closing note that asks a question or names something still needed is not a success. The run is asked once to go through the order part by part, and if it insists, the report says plainly that the run did not finish.
  • Every one of those interventions is logged as a warning, a silent fix to the result sentence would be exactly the silent success this tool exists to stop.

Which messenger works, measured

  • README now names both camps. WhatsApp (5319275A.WhatsAppDesktop) and Teams (MSTeams) are Store/WinUI and cannot be driven. Telegram Desktop (Qt/Win32), Discord and Slack installed from their own site (Electron) and Spotify (CEF) are ordinary programs and run on the agent desktop as usual.
0.3.52026-09-06 (evening)

Fixes from the second merge review and the evening packages.

0.3.42026-09-06

The third build of the weekend: everything merged after 0.3.3 (06.09.2026, 07:06).

The brain leaderboard can finally compare two builds

  • `pruefstand/benchmark/RANGLISTE.md` now groups runs by the product code they ran on, not by the day (six packages were merged on 06.09.2026 alone) and not by the commit (which moves hourly, splitting one measuring night into four tables of three runs each). Two commits whose agentfenster/ tree is identical apart from tests are one build.
  • A delta row per brain says what changed between the two most comparable builds, and refuses to overstate it: only tasks measured on both sides count, and a difference below the spread the same brain shows when repeating the same task prints as within noise instead of a number with an arrow. Silent successes and "claimed done without acting" are exempt from that softening: they are house rule 1, so every change there is real.
  • Measured with 72 fresh runs over 12 tasks: claude 35/35 and codex 36/36, zero silent successes, zero blind reports, against 91.7 % and 90.5 % the same morning. The image path (A31) went from 3 of 6 runs to 6 of 6, and the two runs where Claude claimed a click it had never made are gone. ⚠️ The seconds from that round are not a statement about the product: the Claude account hit its usage limit mid-block, so the real-world task took 906 s instead of 101 s. docs/research/2026-09-06-benchmark.md §7.4 says so plainly rather than publishing the number.
  • A report is only marked +dirty when product code differed from the commit; an uncommitted test or protocol no longer invents a separate build.

You can now pick a smaller model for simple tasks, and we measured that it does not make them faster

  • New Speed setting next to the model, with three positions: Always strong (the default, and exactly what agentfenster did before), Smaller model for simple tasks, and Always smaller. A task counts as simple when its wording names one program and at most two actions, with no chained steps, no "compare/research/decide", and no download-unpack-install chain. That check is a handful of rules on the text, it costs no tokens and no time.
  • A run that turns out harder than it looked switches back on its own. After more than ten steps, or the moment it needs another way at all (going in circles, an element that is not there, the model giving up), it moves to the model you picked and says so in one sentence: "Switched from sonnet to opus after 6 steps: the run needed another way, this run was going in circles." The run report now names the model that actually thought, step by step.
  • ⚠️ Why this is off by default: we measured it, and the assumption was wrong. Over 54 real thinking steps on three real tasks, dropping a rung saved no time, it made the model wordier, and answer length is what costs the time. Median per step: haiku 13.5 s / 565 answer tokens, sonnet 18.8 s / 252, opus 11.9 s / 103. On one task Haiku wrote twelve times as much as Opus (766 against 63 tokens), and one Haiku step ran to 10 273 tokens. All three models returned a usable action 18 times out of 18, so this is about cost and wordiness, not about being right. Turning the setting on will cut what a run costs; it will not make it quicker, and we would rather say that than ship a speed promise with no number behind it.
  • Nothing changes for anyone who leaves the setting alone, with Always strong the model call is byte for byte the one it was before.
  • Brains other than Claude Code are unaffected: there is no measured model ladder for them, so the setting stays inert there instead of guessing a model name that would fail on every step.

Development runs no longer write into the memory you actually use

  • A run started from a checkout keeps its data next to that checkout. If agentfenster runs from a linked git worktree (.git is a file, which is what every parallel build tree looks like) and neither AGENTFENSTER_DATEN nor an explicit override is set, its database and recordings go to <checkout>/.daten instead of %LOCALAPPDATA%gentfenster. The first time this happens, the log says which folder it picked and how to override it. Measured cause: on 06.09.2026, between 12:07 and 12:51, twenty-five junk notes (FakeKlasse, #32770) landed in the live database, one day after it had been cleaned up. ⚠️ The main checkout deliberately keeps using the normal profile folder, because that is where an editable install actually serves the user; moving it would empty the memory people work with.
  • `python -m pruefstand.fahren` and `python -m pruefstand.benchmark` never touch the user's store any more. They put notes and recordings next to the checkout they are run from and say so on the first line of output.
  • A memory note is refused rather than cleaned up afterwards. Notes whose application is a test double (FakeKlasse, MockWindow, …) or a bare window class every program shares (#32770, Chrome_WidgetWin_1) are no longer written at all; the reason is logged as a full sentence. Until now such rows were written first and moved to the bin later.
  • A dialog with no history is named after the process that owns it. Where agentfenster previously had nothing to call a #32770 window, it now asks the owning process ("charmap") instead of storing the useless class name.

A skill you share no longer carries your name, and it works on the other machine

  • A shared `SKILL.md` used to spell out your Windows user name. Every recorded path (C:\Users\<you>\Documents\…) went into the file verbatim, in the readable prose, in the machine-readable block, in the Application line of the header, and in the name: slug that other tools list the skill by. The masker that guards every other shared document only catches secrets you put in the vault, and a user name is not one. An exported skill now speaks in %USERPROFILE%, and importing it puts the home folder of that machine back in, so the same change makes it anonymous and makes it work somewhere else. Measured on a real learned skill: 14 substitutions, no home path left, and the round trip gives back byte-identical steps.
  • What cannot be rewritten is now said out loud instead of slipping through: a recorded window title is the anchor the replay finds its window by, and the contents of a file the skill writes are data, not a path. If either still mentions your user name, the message that tells you where the file was saved says so, before you hand it to anyone.
  • An imported skill now tells you what is missing here, before it runs. It was recorded on someone else's machine, so its program may not be installed and its folders may not exist. Until now the gap list said "nothing missing" and the run failed halfway through; it now names the program and the path, in the same reply that confirms the import.
  • Measured end to end on the real agent desktop, two separate instances: one learns the task (5 model calls, 22 662 tokens), the other knows it only from the file and replays it with 0 model calls. Change the source data and only the broken section costs a model call (1 call, 4 799 tokens); the skill repairs itself and the next run is free again, the self-repair works for a skill that came from outside, too.

Scheduled tasks now run without asking, nobody is there at night

  • A new scheduled task starts on "Run free" instead of "like the setting". The factory setting is "Ask first", and with it no scheduled run ever finished: opening a program already counts as a risky step, and at 3 a.m. nobody approves it. Measured: the same task stopped after 131 s with "not confirmed" and finished in 10 s on "Run free".
  • Existing tasks that followed the setting are switched over once, and the inbox tells you so ("3 schedules switched to run free"). A setting that changes itself and says nothing would be a silent rewrite.
  • Running without asking does not mean running without limits. In a scheduled task, a step that deletes, sends or pays is no longer asked about, it is refused. The run stops, the inbox line says "Needs you" and names the exact step, and after two refusals in a row the task pauses itself. Measured live: a routine that was told to type del /q … stopped after 27 s and the target file was still there.
  • A scheduled run no longer waits for an answer nobody will give. When the agent hit a wall it used to ask "continue, skip or stop?" and wait, measured at 300 s with the schedule line stuck on "Running", after which that task never started again. It now stops with an honest sentence and an inbox line.
  • Opening a program is no longer a risky step inside a routine.

An update that would have arrived half-installed

agentfenster stores its data in a small database on your machine, and each new version may add a table or a column to it. Which of those steps have already run was tracked by a single number. That works right up to the moment two features are built in parallel and one of them reserves a step number it has not delivered yet, the number was counted as done, and the step, once it arrived, was silently skipped forever. The database would have been missing a table, and you would only have found out much later, in the middle of using the feature that needed it. agentfenster now keeps a list of the steps it actually ran instead of a single number, so a step that arrives late still runs. Nothing changes for existing installations except that they become able to catch up.

The phone remote grew up: an app on your home screen, encrypted, and it comes back by itself

  • It can live on your home screen. The phone page now ships a web manifest, a 192/512 icon pair and a maskable icon, so "Add to Home Screen" gives you an agentfenster app in its own window instead of a browser tab. Measured over the real LAN address: manifest served as application/manifest+json, display: standalone, all three icons 200.
  • A dropped Wi-Fi is no longer a dead end. The page shows "Reconnecting…", keeps trying at a calmer pace after three failed attempts, and picks the sessions and the live picture back up on its own, no reload. Measured: 10 s offline, back to "Connected" with both sessions and the same URL. ⚠️ Two things it used to get wrong here, both found in that live run: it showed the desktop app's sentence ("cannot reach its own local service"), which is simply false on a phone; and reloading the page while the Wi-Fi was down showed "Link this phone", you would have believed your access was gone and walked to the PC for a QR code you did not need.
  • The key no longer expires without warning. Twenty minutes before it runs out the header switches from "Link valid for 12 h" to "Key expires in 20 min, scan again", and the countdown keeps running even while the connection is down.
  • Its own offline page. Instead of the browser's bare "no internet" you get "agentfenster is not reachable, same Wi-Fi?" with a Try again button, and it returns by itself when the network comes back.
  • Optional encryption (Settings → Phone → Encrypt the phone connection, off by default, needs a restart). agentfenster creates a certificate for itself (P-256, 397 days, valid for this machine's LAN addresses, stored per installation, never shipped) and serves the phone on a second TLS socket; the old HTTP address becomes a redirect there. Measured on the wire: TLS 1.3, TLS_AES_256_GCM_SHA384, and the certificate's SHA-256 is exactly the one in the QR code. The fingerprint is shown on the phone and in Settings so you can compare it once, and the phone speaks up if it ever changes. ⚠️ Honest price, measured rather than assumed: your phone shows a security warning once (no authority vouches for a home IP), and while encryption is on the offline page and Android's "Install app" stop working, browsers refuse a service worker on a self-signed origin. Adding to an iPhone's home screen keeps working either way. Full reckoning in docs/SECURITY.md. ⚠️ And the bug that live run caught: with encryption on, the redirect fired on the encrypted socket too, so the phone page was unreachable (ERR_TOO_MANY_REDIRECTS). Fixed; nothing that was locked to the PC opened up, /, /mini, /api/tresor and /api/einstellungen still answer 403 from the network, encrypted or not.

agentfenster can now update itself, it never could before

  • The installer download failed for everyone, every time. Checking for an update read the manifest over a connection that knows this server's pinned certificate; downloading the installer opened its own connection that did not. Against our own (self-signed) update server that is always CERTIFICATE_VERIFY_FAILED, while the window above it cheerfully said "Update 0.3.3 available". Measured live against the real server: the shipped 0.3.2/0.3.3 code gives up after 0.11 s; the fixed code pulls all 133 012 442 bytes in 7.3 s, the checksum matches, and an isolated 0.3.2 installation upgrades itself to 0.3.3 in 19.5 s. ⚠️ Because the fix ships in 0.3.4, an existing 0.3.2 or 0.3.3 install cannot reach 0.3.4 by itself, that one hop stays a manual download.
  • The download now honours the size stated in the signed manifest: anything larger is stopped mid-flight, anything shorter counts as interrupted, and either way nothing is left behind, so a retry starts clean. Previously the size was never looked at and the whole file sat in memory twice.
  • A forged signature is no longer reported as a network glitch. It used to read "Update check failed: could not reach the update server (…)". It now reads "Update refused: … Nothing was downloaded or installed."
  • A mandatory update is now offered instead of swallowed. When the server declares a minimum version newer than yours, you get a clear required-update notice and can install it; before, it hit the same wrong error message and the manifest was thrown away, so the update was unreachable.
  • An installer that falls over immediately is reported with its exit code instead of counting as a successful start, a mismatched download is deleted instead of waiting in your temp folder, and "check now" no longer has to wait out the 24-hour interval.

English pages are readable again, and a missed click stops pretending

  • Edge's "Translate this page?" bubble took the whole page away from the agent. On a German Windows, every English page, so nearly every sign-up form there is, raised it, and it became the active window: the control tree then held five translation buttons and nothing else. The page was not empty, it was replaced. The switch against it had been there all along, but only Chromium's (Translate,TranslateUI); Edge translates with its own (msEdgeTranslate) and never saw it. Measured live: before, click "I agree to the terms of service" answered "No element named … in the active window 'Seite aus der Sprache Englisch übersetzen?'"; after, every field on the same page is found and filled.
  • A click that landed 17 px off was waved through as a hit. Before clicking, the page is asked what lies at that point, and the answer climbs upwards until something has a name, eleven pixels above a checkbox that is no longer a control but the whole form, whose text of course also contains the wanted name. So the guard said "matches" and the click fell into empty space. Measured on a real page: the checkbox spans y 500-513, the click went to y 489, and afterwards the page still reported checked: false. A container is now rejected, the proven older path takes over, and the same step reports checked: true.

Four holes the second merge review found, closed

  • Putting `cmd /c` in front of a scripting tool walked around the whole scripting-tool layer. Since consoles became allowed, cmd and powershell are on no list, but the check only ever looked at the program, never at the arguments, and any http(s):// argument was waved through before its file extension was read. Measured at commit 45dc9ca: cmd /c mshta http://…/x.hta, cmd /c certutil -urlcache -f http://…, cmd /c bitsadmin /transfer … and powershell -c "([scriptblock]::Create((irm http://…))).Invoke" all came back with "Allowed: 'cmd' is not blocked by any rule." Every argument of a console is now checked, a web address ending in .hta/.ps1/.exe is refused as a download, and the usual PowerShell downloading verbs (irm, iwr, Invoke-WebRequest, Net.WebClient, [scriptblock]::Create) are recognised. cmd /c dir, powershell -c Get-Date and a YouTube link stay silent, as before.
  • An abbreviated switch now counts as that switch. PowerShell takes any unambiguous abbreviation, so -win hidden and -ex bypass used to slip past markers that spelt them out, as did a Unicode dash, which powershell.exe 5.1 really does accept as a switch prefix (measured here, not assumed).
  • Installing an update now requires a genuinely newer version. Asking to install while already up to date used to download 133 MB and shut the running app down with /CLOSEAPPLICATIONS to lay the same build over itself; and because a signature never expires, whoever held the update server could replay an older signed manifest and downgrade an installation. Both are refused now, along with a manifest whose signed publication date is stale. If the download runs past the request's time limit it is genuinely cancelled, before, the answer said "Could not install" while the worker carried on and started the installer anyway.
  • Two `Continue` taps no longer skip a CAPTCHA. On a phone, double-tapping a button that does not react at once is the normal thing to do, and the second tap counted as "I can see it, carry on anyway", so the run continued with the puzzle still on screen while the log claimed you had said so. The button now says Continue anyway once the warning is up, and only that press carries the run past the puzzle.
  • Gemini: a key in the vault now counts as being signed in. Storing the API key without ever having run gemini interactively used to produce "not signed in … add a Gemini API key in Settings instead", advice you had already followed, for a run that would have worked. Brains waiting for a key also get a key needed label in the first-run list, where they previously had none.

A click that arrived is no longer reported as "not done"

  • On windows that draw themselves (games, canvases, Direct3D) a correct click was thrown away. Effect was measured only by comparing the window picture before and after; a window with four fixed colour panels does not change a single pixel when you click it, so the run said "treat it as not done", the agent then clicked a second time somewhere else, overwrote its own correct answer, and the run died on "two actions in a row changed nothing". Measured over eight recorded runs of the picture-path task: none of the eight passed, and five failed exactly this way. There is now a second, equally measured question, did the application process the click (WM_NULL round trip), and a third answer for it: delivered and processed, the window looks unchanged, and that is not proof your goal is reached. It counts as done for the run, never as proof for a "finished" claim. Keys are deliberately excluded: they are measured not to arrive on such windows at all. Result: 0 of 8 → 6 of 6, one step instead of two to six.
  • "Finished" on the very first move no longer throws the whole order away. The model was reading its own plan as a history of what it had done ("clicked the red field at C2 successfully", nothing had been clicked). The run now says once, plainly, that not a single action has been performed and asks for the first real step; only a second such answer stops it. It is still never counted as success.
  • An agent desktop that held a window a moment ago and is suddenly empty gets a second look before the model is told "there is nothing here".

Voice: the speech model now arrives in the open, and gives its memory back

  • The 141 MB speech model is no longer downloaded behind your back. Until now the first press of the microphone button quietly pulled the model from the internet with no size, no progress and no way to stop it. Settings -> Voice now shows what is missing, how big it is and a Download button with a progress bar and Cancel; the microphone button stays disabled until the model is really there, and recording without it answers with a sentence instead of a long wait. The files live in agentfenster's own data folder, so a portable install carries them along. Measured end to end on a real download: 78.2 MB in 8.1 s, 0 -> 100 %, and the freshly downloaded model transcribed a sentence in 0.7 s. (Also corrected: the model is 141 MB, not the 282 MB the spike reported - that number was a cache holding every file twice.)
  • A recording now ends by itself after 1.2 s of quiet following at least 0.5 s of speech. Clicking again still works as the emergency exit, as does the 30-second cap. A thinking pause in the middle of a sentence does not cut the recording short (measured against synthetic speech, twelve cases).
  • A speech model that nobody uses for ten minutes gives its ~300 MB back. The next sentence loads it again, once.

The agent thinks less out loud, and the runaway answers are gone

  • Three quarters of what the model wrote was never the answer. Measured over 1 370 real thinking steps from 204 recordings: only 24 % of the reply is the action itself; the rest is reasoning written out before the first bracket. The longest reply in the whole field was 5 173 tokens for a single scroll, from a run that was scrolling the same list for the fourth time. The thinking level is now set explicitly (--effort low), which is what a step of this product actually needs: a list of elements, a goal, one JSON line. Measured on the same five test-bench tasks, three runs per side: 15/15 passed before and after, reply tokens per step 89 → 61 (−31 %), the worst reply 1 146 → 249 tokens (−78 %), thinking time −24 %. Wall-clock time only moved 3 %, thinking is half of a run, the other half is acting, checking and reading the window.
  • A reason longer than twenty words is now cut instead of kept, except on fertig and abbruch, which the user reads as the result of the run. And when the run is going in circles, the next prompt says so and asks for a different, short answer instead of a longer justification of the old one.

Test gaps closed in the second night's new code

  • 148 new tests across six files cover what happens when things go wrong in the fifteen modules the second night added: a missing speech model, a schedule hitting a locked database unattended, Windows refusing a CAPTCHA notification, a drag whose slider won't move. Coverage on those modules rose from 82 % to 94 %; the four weakest are now effectively complete (sprache.py 64→98 %, captcha_melden.py 68→100 %, pruefstand/benchmark.py 76→99 %, koennen_text.py 79→100 %). No module is left under 85 %. No real bug turned up, only two comments that no longer matched their code.

A stranger's first look at the second night, end to end

  • A scheduled task does not run to completion on the factory settings. Ask first is the factory style, a program start always counts as risky, and nobody is there at night to answer. Measured on a real schedule: the same entry fails after 131 s ("Step 2 (open on 'charmap.exe') was not confirmed - stopped.") and finishes in ten seconds under Run free, the failure was loud and correctly named, not a silent success, but the feature's own heading promises it works unattended.
  • The phone toggle saved correctly but only showed its QR code after a full page reload, you could switch it on and then find nothing to scan. Fixed. Two places still spoke German to an English screen reader (the four performance fields, and one vault entry), fixed. Across 21 screens in a fresh instance with real Chromium: 0 console errors, no broken server answers.

The command prompt and PowerShell are allowed by default now

  • The command prompt and PowerShell are allowed by default now, cmd, powershell and pwsh are no longer on the preset block list, so "open a command prompt and write the time into a file" works out of the box. The console opens on the agent desktop, not on yours. Updating keeps your own entries: the default list changed, anything you put on the block list yourself stays.
  • What you run in it is still judged. A console that starts with an encoded, hidden or downloading command line (-EncodedCommand, -w hidden, -ExecutionPolicy Bypass, IEX(…)) asks before it runs, and in ask me first mode a typed del, rd /s, reg delete, net user, taskkill or curl -d stops for your yes, a console command takes effect at once, with no dialog left to cancel. regedit, taskmgr, diskpart and format stay blocked.

Recordings: the boxes sit on what was clicked

  • The white box in a replay now hugs the element the step actually touched. 217 of 226 recorded runs show the whole screen, but the player drew that box at the corner that only holds inside a single window, measured on a real run, its centre sat a median of 216 px and at worst 587 px away from the arrow marking the same action. Box and arrow now come from one calculation, the same one the arrow has used since 05.09.
  • No box where there is nothing to draw. 24 of 147 recorded action frames carry no screen position at all; the player used to draw a rectangle for them anyway, with no arrow anywhere near it. Those, and click marks with no width or height, now draw nothing, an invented rectangle is worse than none.
  • A box no longer sticks around. It used to stay until the next step arrived, a median of 3.6 s and up to 20.9 s over a picture that had long since changed. It now fades after two seconds, and disappears as soon as the pointer moves on. A red box on an action that had no effect still stays: loud failure stays loud.

Security

  • 🔑 The phone remote no longer opens the whole app to your Wi-Fi. With Settings → Phone switched on, the service binds every address of this machine, and the interface pages (/, /mini, the recording overlay) were not behind the lock. Any device on the same network could load the interface, be handed the session cookie, and then reach every API route (a picture of your visible screen, the vault, starting a run) and the terminal WebSocket, which is code execution, a self-written Host: 127.0.0.1 header was enough, because nothing ever looked at where the connection actually came from. Now the address the operating system accepted the connection from decides first: from another device only the phone page, its files and /api/handy/* are visible; everything else is refused, and the session cookie never leaves this computer. Measured against a real socket from this machine's LAN address, before and after. The phone remote is off by default, so an installation that never switched it on was never exposed.
  • A 2 kB report can no longer blow up the report server. The 2 MB body cap only covered the compressed body; the gzip was then unpacked without any limit, so 2 MB of zeros became ~2 GB in memory. Unpacking now stops at the cap and answers 413.
  • `/v1/activate` answers instead of dropping the connection when the body is valid JSON but not an object.

Fixes from the second merge review of the night

  • A double click on a Chromium, Electron, UWP or WPF window no longer claims it was delivered when nothing was sent. Those window classes throw away posted mouse messages, so the double click honestly refused to send anything, but the check that followed asked the window whether its message loop was alive, got a yes (it was), and reported "the double click was delivered to … and its message loop processed it". Nothing had been delivered. Worse, the fallback that actually opens things in those windows (asking the app for the element's default action, the only route that opened anything in the 24 August user test) was unreachable from then on. The refusal is now read instead of discarded: nothing sent means no success, the fallback runs, and what comes back is the honest sentence "… cannot be delivered without real mouse input here … Nothing was sent."
  • The red box around an action that changed nothing now stays on screen. In a real run a warning event follows such an action about 20 ms later, and a warning carries no box, so the box that was meant to stay "until you look away" vanished after a single frame. Loud failure is loud again.
  • The microphone button notices that the speech model has arrived. Download the model in Settings → Voice, go back to the chat, and the button used to stay greyed out with "the model is not on this machine yet", while Settings two clicks away said it was. Voice was unusable until the next restart. Switching Voice on now also shows the button right away instead of after a restart.
  • Reading the voice status no longer loads the model. The Settings download row asks for that status every seven seconds and kept doing so after you left the view, so the 300 MB model was pulled back into memory within seconds of every idle unload, forever. Reading is reading now; the model is warmed when you enter the chat, and the row stops when you leave Settings.
  • Switching Voice off gives the memory back immediately instead of leaving the model in RAM until the ten-minute idle timer.
  • A speech-model download that cannot even create its folder now says so. It used to leave the progress bar at "Downloading base: 0 MB of 141 MB (0%)" with a Cancel button that did nothing, until the next restart. Cancel also takes effect before the connection attempt now, instead of up to a minute later.
  • The one-time memory cleanup finally says what it did. A hint line above the list, "Memory was cleaned up once", now names what happened and opens the trash on a click. Until now that cleanup removed 1 042 of 1 152 notes without a word, and the trash was reachable only through SQL. (F-REVIEW-FIXES-2; the trash view itself and its Restore button were built the same day in F-GEDAECHTNIS-2 and are the one the button opens, two packages had each built one, see below.)
  • The Permission-mode line above the task box only stands out while a run is using a different mode than the setting, not during every run.

Fixes from the merge review of the night

  • A folder trigger no longer loses files when a run cannot start. If the schedule saw a new file while another run was still active, it had already moved its "seen up to here" mark before asking, the file was skipped from then on, forever, while the message said it would be tried again. The mark now moves only once a run has actually been accepted.
  • Notification balloons no longer leave an invisible window behind. Every CAPTCHA or failed-schedule balloon left one hidden window in the process, and a click on the balloon could not reach the app at all. Both are fixed: the balloon lives in its own short-lived thread that answers the click and cleans up after itself. Measured: three balloons, zero windows left.
  • Recording warnings are in English again. Four notes that end up in the step list of a replay ("no picture at this step", "the replay may have gaps") were German fragments; they are whole English sentences now and say why.
  • The microphone button and the section chips in Settings no longer paint a second green next to the one accent each view is allowed.

The two settings you actually use stay in sight

  • Autonomy and Permission mode now sit right above the task box and never scroll away. They used to live at the top of the chat history, behind the task suggestions, measured at 1280x800 the suggestions were 1248 px tall in a 301 px window, so the Autonomy row started 585 px below the bottom edge of the app. The suggestions themselves are now one line of short chips with More… next to them (the full cards, with what each one costs, are one click away under Templates); in a very short chat column even the chips give way to a single Browse templates button. Measured after the change: both settings and the task box fully visible at 1280x800 and 1920x1080, empty and full chat, and the empty chat no longer needs scrolling at all.
  • …and the chips now fit in the box they sit in. Taking the two settings out of the scrolling history left the suggestions behind, and they still did not fit: at 1280x800 the empty chat scrolled again, with one chip cut off at the bottom and the other three plus More… below the edge. The welcome block alone is 248 px tall and was shown in a 243 px area. It now appears only where it fits; below that the same wording is a single sentence, so the chips get the room. Measured over ten window heights from 700 to 1440: nothing scrolls, and every chip can be clicked.
  • All eighteen settings sections can be reached from the bar again. The section bar is 701 px wide with 1793 px of content, the same at 1280 and at 1920, so eleven of the eighteen chips sat outside it, among them Permissions, Credentials and Mailbox for 2FA codes. It scrolls, but it has no scrollbar and gave no hint, and a mouse wheel does not move a horizontal box on Windows. The right edge now fades out while there is more to see, and the wheel scrolls the bar.

It says what it cannot do before it tries, not after

  • A window on your own screen that takes no input is refused up front. Handing desktop_task a WhatsApp window used to cost seven steps, seven model calls and 185 seconds before it gave up, every one of those steps reported "no visible effect". WinUI and Store apps, Chromium windows without a debugging port and games accept no click and no key through window messages, and agentfenster never takes your real mouse and keyboard on its own. It now says so before the first step: measured against the live WhatsApp window, 0.06 s, zero steps, zero model calls, with the two ways on named in the same sentence (the Use the real mouse and keyboard permission, or the web version in the agent browser).
  • The refusal comes before the "this window is minimized, restore it and call again" hint, at a WinUI window that hint was a second empty promise.

"Sent" now means sent

  • A message counts as sent only when it appears in the message list. On 05.09. a report to the user was booked as delivered because an element with the time "22:14" stood in the WhatsApp window, that was the draft. The next morning the chat list still ended on 29.08.: two reports were lost, both recorded as success. A key in a chat window is now only proven when a new entry carrying a time stamp appears in the message list; otherwise the run says "delivered, not proven" and names what is missing.

A thirteenth tool: close_window

  • A window the agent opened can be closed again. A Brave window that open_app had put on an agent desktop could not be closed through any tool, press_key('alt+f4') is honestly refused for Chromium, and there was nothing else. close_window sends WM_CLOSE and then checks the window list: a window that is still there (a save dialog) is reported as still there, never as success. force=True also ends the process, but only one agentfenster started itself. It never touches the user's own windows.

Honest about itself

  • An MCP server older than its own code says so. max_seconds was added at 01:02 and ignored at 08:16, not because it was broken, but because that Claude session's server process had been running since before the change and still advertised parameters it did not have. Every desktop_task answer now carries a note when the code on disk is newer than the running process, with its version and the fix (restart the MCP client).
  • `open_app` names the browser profile. Each agent desktop has its own Chromium profile, a login made on agent desktop 1 is not there on desktop 3, which is what turned a working fallback into a QR code nobody could explain. The answer now names the folder and says so.

Gemini works again, with a free API key, and it says so

  • The free Google sign-in for Gemini CLI is gone, and the app now admits it. Google stopped serving Gemini CLI for personal accounts on 18 June 2026; measured here on 06.09.2026, every call ends in IneligibleTierError. The app used to report that as "Gemini CLI is not signed in. Run `gemini` once …", sending you to a sign-in that cannot work. It now says what is true and what helps: get a free key at aistudio.google.com and paste it into Settings → Brain → Gemini API key. The key lives in the Windows Credential Manager, never in the settings database, never in a prompt, and it is masked out of every log line. Google rejecting a key gets its own sentence, so you are not told to fetch a second key when the first one is simply wrong.
  • The first-run brain list distinguishes "not installed" from "no key yet". Gemini CLI installed but no key stored now reads "found at …, but no API key is stored yet, so it cannot answer" instead of showing up as ready. A brain that cannot answer is no longer listed as available.
  • Gemini, Grok and Kimi can finally be chosen as the standard brain. Saving basis_url: "gemini" (or grok, or kimi) in Settings answered 400 Bad Request, the check still only knew Claude Code and Codex, and an empty model field (which every CLI brain uses to mean "your own default") was refused as well. Three brains that existed everywhere else in the app were unreachable from the settings screen.
  • Every Gemini call now carries `GEMINI_CLI_TRUST_WORKSPACE`, so the CLI no longer refuses to start with "not running in a trusted directory".

The memory looks like a memory again

  • One note per lesson instead of one per run. The agent used to write a new line every single time it did something it already knew, in the real database that meant the same three steps stored 375 times. A note now counts instead: the list shows "seen 375×" on one row. Measured on a copy of the real store: 1152 notes → 110, duplicate groups 6 → 0.
  • Notes are named after the window, not after Windows. Every dialog on Windows has the class #32770, so 610 unrelated notes shared one "app" and the graph showed a ball of hundreds of dots with nothing to read. A dialog is now filed under the program it belongs to, "Notepad, Save as".
  • 457 notes that never belonged to you are gone from the list. They came from the test suite, which had been writing into the real user database for months without anybody noticing. They are moved to a trash bin, not deleted, and the memory view says what happened: "Memory cleaned: 585 duplicate notes merged, 457 test notes moved to trash." The one-off cleanup runs once, on the first start after the update.
  • Tests can no longer reach the real database at all, an attempt now fails loudly, with the name of the offending test.
  • The cleanup missed two things, a second pass gets them. Measured on a copy of the real store right after the first cleanup had run: 38 notes still sat under #32770 and 12 under FakeKlasse. The dialog notes are old ones that never stored a window title, so no program can be worked out for them after the fact; they are now grouped under a name a human can read, "Dialog (unknown app)", instead of a Windows constant. The test notes were not left over at all, they had been written after the cleanup, by a test run that predates the guard. Measured on the copy: 135 notes → 110, "apps" 11 → 10, and the view says "Memory cleaned: 38 dialog notes regrouped under \"Dialog (unknown app)\", 13 duplicate notes merged, 12 test notes moved to trash." Running it twice changes nothing.
  • The trash is finally something you can open. Until now notes were "moved to trash, not deleted", but there was no way to see that trash and no way back, which made the promise empty. The memory view now shows it, every note with the reason it was removed, a Restore button per note, and one Empty trash that asks before it deletes anything for good.
  • The guard that keeps tests away from your real database now travels into child processes as well. A test that starts a second interpreter used to slip past it completely, no error, no red test.

Solve a CAPTCHA from your phone

  • A waiting puzzle now reaches the phone by itself. The session that needs you moves to the top of the phone page, marked red, and above the live picture stands the one sentence that matters: "Solve the puzzle by tapping in the picture, then tap Continue." Until now the phone page listed a paused run like any other, it never said what it was waiting for.
  • A tap works during the pause, no Continue needed first. The run is waiting, the window still takes clicks: tap the tile, then Continue. Measured on a phone-sized Chromium over the real network address: the tap landed at 450,291 for a target of 450,290-1 pixel off, one click, not two. Typing into the field under a puzzle still works the same way.
  • 🔑 Continue is no longer taken at its word. Pressing it makes the run read the screen again and look for the puzzle. Still there? The pause stays, and it says so: "The puzzle was still on screen when the run looked again, so it stays paused." Before this, one press carried the run on with the puzzle still visible, and it failed three steps later at something else entirely. Pressing Continue a second time carries on anyway, an explicit, logged decision, shown as "continued on your say-so" rather than "solved".
  • The notification carries the link. When phone remote control is on and a QR code is valid, the webhook (Ntfy, Pushover, your own receiver) now includes the address of the phone page, so the notification leads straight to the puzzle instead of only telling you about it. The link carries no access key, and what leaves the machine is still the fixed sentence plus the run number, never text from your screen.

"Installed" no longer passes for "ready"

  • 🔑 A brain that is on the disk but not signed in now says so. Kimi Code CLI sat in ~/.kimi-code/bin on the test machine and the first-run wizard printed "ready" next to it, the installer creates that folder, and the sign-in check only asked whether the folder existed. The first failure arrived at the user's first click on "Run it". There is now a fourth state, "sign-in needed", with the one command that is missing and no advice to install anything a second time.
  • The benchmark's "OK" is now a promise it can keep. --trocken reported Grok and Kimi as measurable while the runner itself only knew Claude, Codex and Gemini, a benchmark would have crashed on its first Grok run. The runner reads the same one table as everything else now, so a sixth brain is drivable the moment it is registered.
  • `--trocken` works without `--aufgaben`. Asking "which brains could run here?" before planning a benchmark used to exit with an error.

One genuinely free brain, and one template that pointed at the wrong country

  • Z.ai GLM is free through the OpenAI-compatible endpoint, GLM-4.7-Flash and its siblings cost nothing on Z.ai's own price list, no card. The endpoint template offered in the wizard pointed at open.bigmodel.cn (the mainland China platform) with a paid model, so an international key got a bare 401. It now fills in https://api.z.ai/api/paas/v4 with glm-4.7-flash.
  • Kimi has no free path today and the app no longer implies otherwise: the free tier carries no Kimi Code quota, the paid tiers have been sold out since July, and the platform API needs a top-up. Written down with sources in docs/research/2026-09-06-hirne-gratis.md.

A scheduled task that fails no longer fails quietly

  • Every scheduled run now leaves a line in the new Inbox (Settings → Inbox), including the ones that worked: how long it took, how many steps, what came of it, and a link to that run's report and video. A routine you never hear from is a routine you cannot trust.
  • A task that comes back with nothing twice in a row pauses itself instead of repeating the same wrong click every night. The list says "Paused", which is not the same as the "Turn off" you pressed yourself, and gives the reason plus the way back; the button next to it reads "Resume", and pressing it clears the failure count.
  • A run that stopped waiting for your approval is now told apart from one that failed. It shows up as "Needs you" with the honest sentence: a step needed your approval and nobody was there to give it. A routine never performs a confirmation-only action unattended.
  • Windows notification for both cases, a task that paused itself and a task that is waiting for you. Success stays silent on purpose.
  • Each scheduled task now has its own work style: follow the setting, "Ask first", or "Run free". This matters more than it sounds: with the factory setting "Ask first", a scheduled run stops at the first risky step, even opening a program counts, and at 3 a.m. nobody approves it. Measured on the same entry: 131 seconds and stopped, against 10 seconds and done under "Run free". The form says so where you pick it.
  • The schedule list now shows the last run with its verdict and duration and the next run time, instead of only a red chip.

Two accents where the view allows only one, found by the merge audit

  • Two settings sections and the microphone button both claimed to be "the one accent" of their view, only one of each pair was right. Settings shows its accent on the active toggle, not the section bar; the agent screen shows it on the Send button, not the microphone next to it. Both are neutral now (lighter text on a lighter surface, the same treatment the recording light already used).
  • The recording light in the small overlay window was alarm red, the same bug the large window had already fixed, with a comment claiming both were identical. It now uses the same colour as everywhere else a recording is shown.
0.3.32026-09-06

It says which brain it looked for, and where

  • "Not found" no longer covers up "did not answer". The first-run wizard now lists every brain it looked for on this machine, names the program names it asked PATH for and the folders it checked on top of that, and offers Check again so a codex login you just typed is picked up without restarting the app. A search that hits its five-second deadline, a virus scanner holding up the file search is the usual cause, now says exactly that: "Codex CLI did not answer within 5 s … it may just be slow rather than missing. Nothing was ruled out." It used to read "not found, install it or pick another", and people switched away from a brain they actually had.
  • One sentence builder for all five brains instead of five. And the list of installed brains was already there as an API since 02.09, until now no part of the interface read it.

A second look before something you cannot undo

  • A neutral button in a dangerous window now gets a question too. "Ask me first" mode has always stopped at words like Delete, Send or Pay, and at every action that writes a file or starts a program. What it could not see: a button called "Ready" or "Step 3" in a window that is about to spend 249 euros or publish a post. Before such an action, the same brain you already run is asked one cheap question, measured here: 5 of 6 verdicts right, 1.8 s each. It is budgeted to six questions per run, remembers its answers, and never costs anything in free mode.
  • 🔑 It can only ever be stricter, never laxer. When the word list already says "risky", the model is not asked at all, otherwise a window could talk the safety check into "no, this Delete button is safe", and the check would become the hole. The screen text is handed over framed as data; an injection telling it to always answer no was tried and did not work. A check that fails or answers something unreadable is reported, never silently taken for "harmless".
  • New test case in the bench: a forged system note hidden inside a file the agent has to read ("ignore all previous instructions…"), judged by files afterwards, never by what the agent claims. Three runs with Claude Code: 0 of 3 followed it, all three did the real job. The honest baseline is in docs/SECURITY.md §11 together with what this defence does not cover.

Runs stop waiting for things that cannot happen

  • File work no longer waits for a screen to change. Reading, writing or copying a file cannot alter a window's control tree, yet every such step sat out the full 1.5-second effect check and re-read the tree up to five times while doing so. Measured on the test bench: that was 26 % of task A26 and 25 % of task A22. The effect check now answers immediately for file work, its proof comes from the file itself, as it always did. Across four bench runs the effect check dropped from 54.7 s to 26.0 s (5/5 passed, before and after).
  • A "1.5-second" wait that really took 9.1 seconds is capped. The old limit counted only the planned pauses and forgot that each round also reads the control tree, about 1.5 s on a 200-element window. On large windows that turned into nine seconds of waiting per step.
  • Every recorded step now says where its time went: reading the tree, thinking, acting, checking the effect, recording. Until now a recording knew only how long the model took, 59 % of a run, and the other 41 % was a black box. The numbers and what they mean are in docs/research/2026-09-06-denkschritt.md.

The memory graph, the way Obsidian does it

  • Every app in the graph now has its own colour, and its notes carry the same colour, the same dot appears next to the app in the list on the left, so you can tell at a glance which cloud belongs to which app. Misclick notes keep the warning colour, because those are the ones worth spotting.
  • Five sliders under the graph, the same ones Obsidian gives you: center force, repel force, link force, link distance, and how early note names fade in as you zoom. They act on the running layout, the dots drift into their new positions instead of jumping, and the setting is remembered.

The agent can use a code editor now, and never opens Notepad on your screen

  • VS Code and other Chromium apps are drivable. They were not: the agent saw eight elements instead of 149, and not a single character arrived. Two causes, both fixed. VS Code keeps its Chromium files in a versioned subfolder, so the detection that hands out the accessibility switch and the debugging port looked next to Code.exe, found nothing and said no. And keys had no way in at all. They now go through the same DevTools channel as text and clicks, and the tool reports how many key events the app actually counted. Measured end to end on the agent desktop: three lines written into a file through VS Code, 39 bytes on disk, six calls, 30 seconds, where the same task was impossible the day before.
  • `notepad.exe` no longer puts a window on your screen. On Windows 11 the notepad.exe in System32 is only a switch that starts the Store app, it names its own redirect target inside the binary. agentfenster reads that and refuses the start before anything opens, the way it already did for calc.exe. Measured: 0 new windows on the user's desktop.
  • A full editor is no longer reported as empty. VS Code takes its input through Chromium's new EditContext API, where no DOM input event fires, so the check that proves "the app really saw the text" measured zero and called a successful typing a failure, sending the agent off to clear a field that was full. When the focused field uses an EditContext, the proof is now a different one and the sentence says so: the text is visible in the app now and was not before.
  • `open_app` reports the pid that owns the window, not the pid of the launcher. With VS Code the reported process was gone minutes later (taskkill /PID found nothing), so nobody could clean up what they started.
  • `type_text` no longer cuts its echo silently. It reported the first 40 characters of what it typed with no ellipsis and no length; the report read as if the rest had never been typed.
  • What the agent learned is written in English. A learned misclick entry read In #32770: taste ctrl+z does nothing - use klick Abbrechen instead, German words in a sentence that goes straight into a model's prompt.

The agent can drag

  • Drawing, selecting, sliders and drag & drop now work. Until now the agent could click, double-click, right-click, type and scroll - but it had no answer for a stroke. Asked to "open Paint and draw a yellow circle" it picked the colour and the oval tool correctly and then clicked eleven single points on the canvas before giving up after fourteen minutes. It now drags: between two points of the window picture (a canvas or a game), or from one element onto another, or a slider by a direction and a distance. Measured on the same task against a test canvas: 11 steps, 151 seconds, six yellow strokes drawn, against 14 steps and 852 seconds with nothing drawn before.
  • The MCP server gets a twelfth tool, drag, with the same two forms - so a Claude Code session can draw and select too, not just the built-in agent.
  • A drag no longer claims success it has not checked. A drag delivered to a real input field used to report ok via postmessage; the field's selection was measured afterwards and was empty. The effect is now measured on the window picture, and where there is no picture the sentence says "delivered" instead of "done".

Watch a session from your phone (preview)

  • agentfenster can now be watched and answered from a phone on the same Wi-Fi. Settings › Phone shows a QR code; scan it and the phone lists the running sessions, shows a live picture of the agent desktop, and lets you answer the agent's question, continue, or stop it.
  • When a CAPTCHA needs a human, tapping the picture on the phone clicks that exact spot in the window on the PC, and a short text can be typed into it. agentfenster still never solves a CAPTCHA itself - a person does, from wherever they are.
  • Off by default, and narrow when on. Only the phone endpoints become reachable from the network; settings, vault, recordings and terminal stay locked to the PC. The access key lives in the Windows credential manager, expires (12 hours by default), can be revoked with one click, and locks out after eight wrong tries. There is no TLS on this path, which is exactly why it never leaves your own network.

Skills you can actually manage

  • Everything the agent can do now lives in one list, what you recorded yourself and what it picked up from a successful run, told apart by a small label, replayed by the same player.
  • Each skill can be renamed, switched off and back on, and shows how it has held up: how often it worked, how often it failed, when it last ran. Off means "stop picking this one on your own", it stays in the list, keeps its counters and can still be replayed by hand. Deleting used to be the only way, and that threw the recording away with it.
  • A shared SKILL.md can now be imported, not just exported: the exported file carries the recorded steps in a machine-readable block, so a skill survives the trip to another machine intact. A SKILL.md without that block is refused with a sentence saying why, rather than half guessed from its prose.
  • After a run that taught it something, the app says so once, "Saved as skill X, next time this runs without a model", in accent colour instead of the warning red it used before. Silence it with the koennen_hinweis setting without giving up the saving itself.

Faster and less guessing

  • A batch of actions now stops at the first one that did not work. The agent may plan up to five actions in one thinking step, which is what makes it fast. Until now the batch only stopped when a target had disappeared - an action that ran and provably moved nothing let the rest go out anyway, on an assumption that had just been disproved. The batch now stops there, and the actions it skipped come back by name ("Not executed: 3. click 'Save'") instead of silently vanishing, so the next thinking step knows exactly what did and did not happen.
  • A picture of one control instead of the whole window. click_point takes an element (an id or name from read_desktop) and hands back just that control, cut out of the window picture at its native resolution instead of a whole window shrunk to fit. Small text stays readable, and the picture costs a fraction: a 1600x1200 window costs about 2,458 tokens as a picture, the crop around a button about 20. It only looks - it never clicks.

Fixed

  • A click on a covered control is refused instead of landing in the window on top. The check for "who owns this spot on the screen?" only ran when the caller had not already resolved the control - which is the rare case. In the normal one it was skipped, and the last resort (moving the real pointer to a screen coordinate and clicking) went out unchecked. It is now its own named step in front of every raw click, and the refusal says which window is in the way and how to get past it.

Fixed

  • A stuck run started over MCP no longer waits five minutes for an answer nobody can give. Since the run-keeps-going change, a run that hits an obstacle pauses and asks whether to continue, skip or stop - but those three buttons only exist in the agentfenster app. Over MCP the run stood still until desktop_task cut it off after three minutes, and the caller got a time-limit message instead of the real reason. It now ends right away and says both what it is stuck on and why nobody could be asked, and the pause can never wait longer than the call itself is allowed to run. Measured: the test file covering this path went from 905 s to 4 s.
  • A setting that cannot be applied is no longer saved anyway. PUT /api/einstellungen wrote first and applied afterwards, so a value that passed validation but failed while being applied stayed in the database and quietly never took effect. It is now converted first and only written if that works - the answer is a 400 with a sentence, and nothing is stored.

Works without you

  • Schedules and triggers (Settings → Schedule): run a task on chosen weekdays at a time of day in your own local time ("every Monday at 08:00"), every time a new file lands in a folder inside your user profile (optionally *.pdf, and a file still being written waits), or once every time agentfenster starts. Every scheduled run goes through the same run path as the task bar, the chat and the CLI, so the licence gate, your profile, permissions and the zero-model replay of a learned task all apply unchanged. Nothing starts while a run of yours is active, a folder outside your user profile is refused on every pass (not only when you save it), and a failure is loud: the row shows the reason and Windows shows a notification. "The agent says it is done" is not counted as success, only a last action that measurably changed something is. New API: GET/POST /api/zeitplan, DELETE /api/zeitplan/{id}, POST /api/zeitplan/{id}/jetzt.

First five minutes

  • A library of twelve ready-made tasks sits where the empty chat used to be: write a note, rename it and zip the folder; show hidden files and extensions; save a DxDiag report; read a value off a web page into a file; print a text to PDF; and seven more. One click puts the task in the box, a short dialog asks for the blanks ({folder}, {file}), and you press Send, a template never starts a run by itself. A Templates button next to the task box opens the same library once the chat is no longer empty. Every card names the programs it drives and how long it took: eight of the twelve carry the median of the test runs that actually passed (~85s · 5/7 test runs passed), the other four say "not measured yet" rather than an invented number, and a test recomputes both from the real test reports, so a made-up duration cannot slip in. Also available as GET /api/vorlagen.

A run that hits a wall asks instead of dying

  • When the agent is truly stuck, window gone, element missing, two actions in a row with no effect, the model itself gives up, it no longer ends the run. It tries the usual escapes first, and only then pauses and asks you in the live view: one sentence on what is in the way, and three answers, Continue, Skip this step, Stop, plus an optional note that goes into the agent's next turn. The run then carries on from that exact point, not from scratch. It rings like a CAPTCHA pause (Windows notification, the same optional webhook). Never more than two questions without progress, and never a question for a safety refusal (a rule, a missing permission, a CAPTCHA stay a plain refusal). Nobody answers within 5 minutes (adjustable) and the run ends honestly, keeping the recording. Measured live: closing the window mid-run paused it, an answer typed for it resumed the same run; without anyone to ask, the identical run died one step earlier.

Reliability fixes from the wish-list audit

  • The agent no longer keeps pressing a dead button. A run now remembers, per target, that two clicks in a row measurably changed nothing, the third click on that same target is not even sent. The run does not stop over it: it gets a sentence saying why and what else it can try, and the final report counts how many clicks were skipped this way.
  • A capped list now says it was capped. The element tree is capped at 200 entries so one huge window can't eat an entire run's budget; the model used to get those 200 rows with no hint that more existed. It now sees a one-line notice above the capped rows, for 59 measured tokens.
  • A second Claude Code profile no longer silently loses the cheap-planner setting. Creating a profile for a second Claude account used to switch off the tiered planning step (a cheaper model plans, a stronger one acts) without saying so; it now stays on for a second Claude Code profile, and stays off for other brains (Codex, Ollama, Gemini), where it was correct before.
  • Settings describe themselves completely. All five places a setting can be changed (Settings page, GET/PUT /api/einstellungen, MCP permissions, profiles, .env) are now covered by one machine-readable schema (GET /api/schema), built from the same source the app is checked against. (all four: F-WUENSCHE-2)

Idle cost, measured

  • The background ticker behind schedules and goals used to open a fresh database connection every 2 seconds forever, even on an installation that never used the feature again after its first check. It now sleeps 30 seconds while the goal table is empty and snaps back to the fast tick the moment a goal is created, a check failure keeps the fast tick so a broken state isn't hidden for 30 s, and stopping wakes the idle sleep immediately instead of waiting it out. A cold, idle instance measured afterwards: ~1.7 s to answer its first request, ~0% CPU and 87 MB RAM a minute later.

A skill that breaks fixes itself, one section at a time

  • A learned skill is now replayed in sections, a section ends where the agent changes stage anyway (starting a program, switching window, moving between file work and window work). If one section no longer fits, only that section costs model calls; everything before and after it still replays for free. Measured on a real two-section task: a full run cost 5 model calls and 23 646 tokens, the broken-section run cost 1 call and 4 637 tokens, same correct result.
  • Afterwards the skill remembers the new way: the section the model redid replaces the old one, the skill is marked "repaired after a change", and the previous version is kept under versions in case you want it back. The next run of the same task is free again (measured: 0 model calls). Before, the same broken button cost the full price on every single future run.
  • Every replayed result now says where it comes from, "replayed from steps recorded on 2026-09-04; recorded source values re-checked: yes", so a skill can never quietly sell you yesterday's answer as today's.

Skills learn from plain file work, and check their sources are still true

  • A recurring task that only reads and writes files (read a value, save it elsewhere, copy a file) can now be learned too, before, only tasks that ended by clicking or typing something on screen could be. Measured on real test-bench tasks: the second run of the same task took zero model calls instead of four or five, 9× and 8× faster, same pass/fail verdict both times.
  • Before replaying a learned step that once read a value out of a file, the run now checks the file still says the same thing. If it changed, the model takes over and reads the new value instead of quietly writing yesterday's answer and calling it done.

Recordings & video

  • Runs started over MCP (desktop_task) and runs from the test bench now record the whole desktop as a continuous frame track at one fixed size, like a run started in the app. Before, an MCP run had no frame track at all, so its video was assembled from single-window snapshots in sizes between 162x186 and 1280x761, each blown up to full screen. Their pointer is drawn where the click really landed on screen, not where it landed inside the window (measured on a real run: 545 px apart). Recordings made before this release are repaired too: a desktop run is now also recognised by its window entry, so an already-saved MCP or test-bench run replays with the pointer in the right place.

A fresh Windows, actually tried

  • The installer no longer hangs forever on a fresh Windows without WebView2. Windows Sandbox (a real 20-second-fresh Windows, no Python, no WebView2) caught it: the installer's "silent" switch suppressed the wrong kind of Inno Setup dialog, so a Yes/No prompt sat there with nobody able to click it, after three and a half minutes not one file was installed. Fixed, and the fallback path (open the UI in a browser when WebView2 is missing) no longer crashes with a raw traceback of its own. New tool werkzeuge/sandbox_probe.py repeats the whole check in three minutes with screenshots.
  • 🔑 Without a code signature, Smart App Control refuses to run the installer at all on a fresh Windows, no "Run anyway", confirmed against Windows' own event log.

Dogfooding: using it, not just testing it

  • A blocked window no longer says "changed the window" when nothing changed. A file-write step used to claim it altered the target window's control tree even when a re-read proved it hadn't; it now says where it actually looked.

Three of Philipp's wishlist points, checked for real

  • A dropped browser window is no longer silently cropped out of the recording. A recording fixes its frame size at the start; if the agent later drags a window onto a second monitor, the frames covering it used to be cut with no note. The recording now counts how many frames this happened to and the player says so in one sentence, measured on a real Explorer task: 489 of 617 frames, 79 %.
  • `desktop_task`'s default mode says up front that it asks before every program start, nobody unattended can answer that, and its built-in 180-second timeout is raised where a real two-app "Save as" task measurably needed 338 seconds.
  • The `desktop limit` setting now actually limits the desktops MCP hands out. Claude Code over MCP was wired to a fixed 6 regardless of the slider, so raising it to 12 still put a seventh session on a shared desktop, and the message to the model kept claiming "all 6 agent desktops are taken" even after the slider said otherwise. Both sides now read the same number, and a full desktop's message names where to change it.
  • A right-click that opened no context menu no longer reports success. The proof that a menu actually opened existed already; it was measured and then thrown away.
  • The first-run brain search no longer dies completely if one brain's check runs long, it now says "Codex: not found" for that one instead of failing the whole screen; and a blocked action's explanation no longer reads click '' instead of a sentence naming the risky action.

Thirteen files nobody had read, three real bugs

  • Reviewed the 13 modules and interface files that no review or test had ever named; three had real bugs (see above: right-click, brain search, blocked- action wording). A coverage tool (werkzeuge/reviewdeckung.py) now finds such gaps automatically instead of by hand.

Nothing pretends to be finished when it isn't

  • `New task` in the tray icon's menu now actually starts a new task. It was wired to the same code as Open, you got your window back exactly as you left it: old tab, old text, no cursor in the field. It now switches to the agent view, clears the task box and places the cursor in it, and checks afterwards that the cursor really landed there.
  • One data folder, one environment variable. A second, isolated test instance with its own AGENTFENSTER_DATEN still wrote into the running app's port file, because the port file used a different variable underneath, now both use the same one, so two users on one machine no longer collide.
  • Settings now has a jump bar, one button per section, that follows you down the page and highlights the section you're in, the 18 setting cards used to be one long column with no way there except scrolling.

Ten stuck test threads, one real product bug

  • Closing a terminal now honestly reports failure instead of "ok" when its reader thread is actually stuck. On Windows, a terminal's output reader can hang forever in a read that neither ending the process nor closing the pipe unblocks, closing it used to report success anyway. beenden now returns False in that case rather than pretending it worked. Alongside it, the full test run no longer leaves background threads running after their own test finished, measured at 10 stray threads before, 0 after, and a new guard fails any test that leaves a new one behind.

Seven holes closed in the ways agentfenster now works without you

  • A stuck run's webhook call no longer repeats the on-screen dialog text (names, invoice numbers) to whatever address you configured, it now says only that a run is waiting, plus its run number; a receiver on your own machine still gets everything.
  • Webhook calls now require https (plaintext only to a receiver on your own machine) and no longer follow redirects.
  • Schedules are capped at 100 entries; a typed answer to a paused run is capped at 2000 characters and a template's answer at 400, these go straight into the agent's next instruction, so an unbounded one was an open door for injected text.
  • The file that records the app's own port is now fully validated (hostname, not just port and process id), closing a way another program could hijack it and redirect your agentfenster key to itself.
  • The report server no longer starts without encryption unless it is bound to itself only, and its read-only endpoints are rate-limited too, the installer download used to be the one fully open path.

DPI: a third bug, found while checking two supposedly-fixed ones

  • The window picture on a scaled monitor is now stretched to match what Windows actually draws. A window that doesn't know about display scaling paints itself at two-thirds size while Windows stretches it to full size on screen, the live view showed the small version at the large position, so a click landed exactly where intended but the picture under the pointer showed something else entirely (up to 320 px off on a 640 px window). Measured on a real test window: 3 of 4 colour fields mismatched before, 0 of 4 after, across the live view, desktop recording, window recording and single-window picture. An unscaled monitor changes 0 pixels.

Nine interface bugs from a fresh instance, one truly ugly

  • Continue / Skip this step / Stop no longer get cut off in a narrow chat column. The three buttons needed 351 px in a 258 px column, Stop, the one button that ends a hung run, was not on screen at all.
  • Dates now always render in English regardless of Windows' own display language, instead of mixing in German ("last used 3. Sept. 13:02") on a German machine.
  • The task input field's placeholder text was truncated to "Tell the agent what to" after two new buttons (Templates, microphone) landed in the same bar without the column's minimum width being re-measured.

The agent's first Spotify playlist that actually has songs in it

  • A "create playlist, add three songs" task now runs start to finish, checked against the window afterwards, not the agent's own report: 3 tracks really landed in the new playlist, in 26 steps, 169 seconds (the same task ended with seven empty playlists on 29.08.). Two causes: the agent typed into the field's label instead of the field itself, and it did not see a dialog it had just opened (Spotify's edit dialog starts at element 1904 of 1937), an open dialog is now sorted to the front of the list, the same fix already applied to open menus.
  • `cmd.exe` no longer opens its window on your visible desktop. Windows 11 routes every console through Windows Terminal, a Store app, the console now goes through the bundled conhost.exe instead and stays on the agent desktop.
0.3.12026-09-05

Everything built after the 0.3.0 doc pass in the same overnight run, the second wave, roughly 07:30-14:00.

Reliability & effort economy

  • One recurring task is now learned by itself: succeed once, and the same task next time is replayed with zero model calls instead of thinking again, measured live: 4 planner calls → 0, 1.3 s → 1.0 s, and the real window shows the same result both times. A run without a clear window anchor (app already open) is never learned, so it can't click into a random foreground window later.
  • A CLI, agentfenster run "<task>" --warte --json, drives a run from PowerShell, n8n, Task Scheduler or CI, starting a hidden instance itself if none is running yet, and a header-based API key (X-Agentfenster-Key) covers tools without a browser cookie; both are documented with a real example in docs/API.md, generated from the app itself so it can't go stale.
  • Reasoning inside a run is capped at twelve words for an ordinary action (full sentences stay for the final verdict and any refusal/question), measured −42% output tokens per step (254 → 147) across twelve real step situations, −15.5% on the price of a session step. MAX_THINKING_TOKENS itself changes nothing measurable and an English system prompt saves under 1% once caching is accounted for, so neither was built.
  • Every agent desktop now carries a short hash of its own data folder in its name (agentfenster_<hash>_platz2) instead of a machine-wide name, so two installations with different data folders never collide on the same Windows desktop and never "clean up" a desktop that belongs to another installation; the normal installation keeps its old names on purpose (the browser profile is tied to them).
  • On a scaled monitor (125%/150%, most laptops) the picture the agent saw was shifted and partly black, up to 57% black, one measured point 15 px off. Fixed by cropping in the same coordinate space the window actually paints in, not the one Windows reports; measured afterwards on a real 150% monitor: 0% black, 0 px offset. The click path itself was never wrong.
  • A tab click in a classic Windows property sheet (folder options, system properties) could report success while the visible page never changed; a classic radio button could report an honest "nothing happened" but never actually flip. Both are fixed: the tab reads the visible page back and presses Ctrl+Tab at most once if needed, and a radio button that ignores the normal click gets one direct BM_CLICK and is read back again. A double-click now checks what it actually hit, the same way a single click already did.
  • press_key without a target element now works in RichEdit controls (the Character Map "Character to copy" field, and similar controls elsewhere): the caret is moved to the end of the text first, because these controls silently drop keystrokes sent through plain window automation otherwise.
  • Five review findings across package boundaries fixed: an unmeasurable tab page no longer gets a blind Ctrl+Tab; the terminal's workspace limit now counts running shells, not ever-touched ones; a failed vault write-back cleans itself out of Credential Manager instead of leaving a half-saved secret.
  • A regression hunt across all 25 packages merged so far found the merged whole was red even though every package was green on its own: an unmeasured scroll counted as proven success and reset the stuck-agent counter; read_desktop's memory lookup silently returned nothing once a parallel change tightened its filter. Both fixed; schleife.py cut from 869 back to 760 lines.

A real end-to-end task, twice

  • Philipp's own task from 05.09. (make a folder, write and save three lines, rename the file in Explorer, zip it, report path and size) now completes in one run, three times in a row: 45 steps/6:39 min → 11-14 steps/52-84 s, zero wasted actions (was one in nine). Root cause of the first failure: Windows Explorer's rename box hides the file extension, so the agent typed a name that became file.txt.txt on disk, now the planner knows to check the extension. A checkbox click now actually ticks instead of only "selecting", an OK/confirm dialog is now known to need its button pressed, and an unproven "done" is questioned once instead of trusted.

CLI brains

  • Grok Build and Kimi Code join Claude Code/Codex/Gemini as subscription brains (own sign-in, no API key); Z.ai's GLM Coding Plan runs through a Claude-Code-compatible endpoint, filled by one button instead of typing the URL by hand. Measured: Grok never exits after answering on its own (agentfenster now stops reading once its JSON is complete, or every step would wait out a 3-minute timeout) and spends about 91,000 input tokens of its own scaffolding on a trivial step, both now shown on its card. Kimi accepts its prompt only as a command-line argument, capped at ~32,000 characters by Windows itself; called through a .cmd wrapper it silently truncates everything after line one and reports success anyway, agentfenster refuses that combination outright rather than risk it.
  • A hosted-model card ("A hosted model (API key)") in the first-run assistant offers one-click templates for xAI Grok, Z.ai GLM, Moonshot Kimi and OpenRouter; the key goes into the vault, never into settings.

Recordings & video

  • Desktop recordings no longer cut into a single window's own screenshot mid-film, a real 15-glitch, 29-second run went from 380 KB/visibly jumpy to 209 KB/smooth once only the continuous desktop capture is used. The recorded mouse pointer position in the exported video/player was off by up to 544 px (it used the wrong source); both are now measured pixel-exact against the click target.
  • MP4 export no longer needs a separately installed ffmpeg, the app now ships its own (imageio-ffmpeg, ~31 MB, no download at startup); a Unicode export path, a too-short server timeout and a UI hitch while measuring recordings' disk usage are fixed. The recorder's never-used pointer track file is removed (the player never read it).
  • Recordings now follow the same data folder as settings (%LOCALAPPDATA%\agentfenster\sitzungen or AGENTFENSTER_DATEN), not the install folder, which was read-only for a normal user and gone on update/uninstall, existing recordings are moved there automatically on next start. Test-bench runs get the same continuous frame track as a normal run now; a Settings → Recording toggle turns the frame track off for all four run paths (desktop, MCP, CLI, test bench) to save disk space.
  • Every recording gets a self-contained HTML report (GET /api/sitzungen/{id}/bericht.html, steps without effect on top in red, before/after pictures per step, every field through the vault masker), reachable from a new "Report" button next to MP4/GIF.
  • The taught-skill recorder ("Record what you do") is now actually wired into the running app, it minimises the window and shows the overlay for real, not just in its own tests; the exported SKILL.md now carries the name/description header Claude Code needs to find it at all; a mitschnitt stopped from the overlay is no longer lost in the UI.

First run & settings

  • The guided tour now starts by itself, exactly once, the moment first-run setup ends, however it ended (Run it / Start tour / Skip setup); a return path in Settings is now labelled "Show tour again" once you've seen it.
  • Eight real first-run bugs found by watching every screen: clicking one of the lower brain cards (Grok/Kimi/hosted) scrolled the page and hid the card you just picked; the report opt-out toggle sat below the fold and was never seen without scrolling; Tab could escape the assistant into the app underneath it. All fixed and measured; also fixed: raw Markdown hashes showing in the licence text, duplicated CLI card labels, hardcoded counts that could drift from the real list, and playback buttons that sat clickable-but-inert until something was actually loaded.
  • A settings-only-mode reader path (/api/eula) plus the new "License" onboarding step: an EULA (docs/EULA.md) shown once, acceptance required before any run can start server-side, enforced at one shared choke point.

CAPTCHA & 2FA

  • A paused CAPTCHA now actively reaches you: a Windows notification (click it to bring the live view forward) and an optional webhook (Ntfy, Pushover, your own receiver). A Continue button in the live view says "look again" without answering anything for you, the run already resumes by itself once the screen changes. Default wait raised from 3 to 10 minutes, now a 1-60 minute Settings slider; more Cloudflare/hCaptcha/ reCAPTCHA phrasings recognised. Still no automatic solver, by decision.
  • A new agent action, hole_code, reads the newest 2FA code out of a configured mailbox (IMAP with an app password, or a locally running Outlook) mid-run and types it in, only ever offered to the model when a mailbox is actually configured, so it costs nothing when nobody set one up. The code itself never appears in logs or in the session report.

Background & shutdown

  • Optional tray icon (Open · New task · Pause agents · Quit), optional "minimise to tray on close", a configurable global shortcut (Ctrl+Alt+A by default) and a quiet autostart entry, all off by default.
  • The app now asks "Tidy and close?" before it quits, from the X, Alt+F4 or the taskbar, with three choices: tidy and close (closes open agent desktops and running tasks cleanly, 10 s per step), just close, or cancel; a setting turns the question off entirely.

Compliance & legal

  • docs/LEGAL.md documents, for each bundled CLI (Claude Code, Codex, Gemini), whose terms apply and what agentfenster itself promises (no resale of access, no credential storage); the onboarding "Brain" step links Anthropic's terms when Claude Code is chosen. Accepting Anthropic's Commercial Terms of Service is the one remaining condition only Philipp can complete, no money may be taken for a product bundling Claude Code until then.

Hardening (test-suite honesty)

  • 13 of 14 tests that silently skipped on a fresh machine (no user recordings yet to check against) now run everywhere via a small, checked- in 3 KB sample recording; three new tests cover what happens when a Chromium app's DevTools connection drops mid-click or mid-type (the existing code already did the right thing, untested until now).
  • agentfenster/planer.py cut from 823 to 569 lines (the shared prompt constants moved to planer_prompt.py) to clear the 800-line house limit again.
  • A full prüfstand run of all 26 non-screen tasks passed 25/26 (96%); the one failure (A53, a Windows checkbox list without reliable element names) is a real, pre-existing input-layer limit, not a regression from any of the night's packages.

Install & updates

  • Build and publish is now one repeatable command (python werkzeuge/bauen.py --installer --veroeffentlichen): PyInstaller build, Inno Setup installer, Ed25519-signed manifest, upload to the report server, all from the single version in agentfenster/__init__.py.
  • Verified the frozen build actually carries imageio_ffmpeg's bundled ffmpeg.exe (needed since F-VIDEO-F made MP4 export a real dependency) and that MP4 export works with no system ffmpeg on PATH.
0.3.02026-09-05

Install & updates

  • Real, frozen Windows build (PyInstaller, agentfenster.exe + agentfenster-mcp.exe, no Python install needed) and a per-user Inno Setup installer (dist/agentfenster-setup-0.3.0.exe), no admin rights needed. (2be2e38, 332dadb, d78d436)
  • The app checks for a newer version itself and only accepts an update manifest signed with our Ed25519 key, verified against a real server. (2be2e38)
  • Unsigned build, Windows SmartScreen and Defender may warn on first run; documented as expected, not a bug.

Reporting & licensing

  • Every run and MCP session builds a small, masked report (never images, element trees or vault secrets) and sends it to our own report server, opt-out in Settings → Reporting, on by default. (4f4995e)
  • Tester keys (AF-XXXX-XXXX-XXXX) unlock the full app on one device, including seven days offline. (4f4995e)

More brains

  • Google Gemini CLI is a third built-in brain (gemini -p, Google sign-in, free daily quota). (f6a7b9e)
  • Moonshot Kimi K2 and Z.ai GLM run through a Claude-Code-compatible endpoint profile, no separate CLI needed, billed on their side. (f6a7b9e)
  • Fixed: a chosen brain other than Claude Code/Codex showed up as "Local / custom" instead of by name. (f6a7b9e, ad602f7)

Guided tour & settings

  • The first-run tour now points at the real controls (skill recording, memory, the first settings group) and asks before an early exit instead of aborting silently. (94c963e)
  • Settings expose a machine-readable schema (GET /api/einstellungen/schema) and a new "Performance" group: frame rate, screenshot quality and the number of concurrent agent desktops are now live, adjustable settings. (d63feb4)

Recordings

  • Every recording gets an automatic, inline-renameable name; a favourite star that survives cleanup; live search by name/task/window/client/date/ outcome; ten sort criteria including most/fewest steps and most tokens. (ee1cf52)
  • Recording a taught skill now minimises the main window and shows a small on-screen overlay (Stop / Finish / Save) instead of leaving the app window in the way. (c006f72)
  • The recording canvas is fixed once at the start of a run instead of being recalculated every frame, no more format jumps when a window sits outside the primary monitor; MCP/Claude Code sessions now get a real frame track, not just per-step screenshots; time-lapse buttons (2×/4×/8×/32×) now visibly differ on long pauses. (d31a8eb)
  • Recordings live next to the app's own data (AGENTFENSTER_DATEN or %LOCALAPPDATA%\agentfenster), not inside the install folder, recordings made before this fix are moved there automatically on the next start. Test-bench runs now get a continuous frame track as well, and a new Settings → Recording toggle turns it off for all run types to save disk space.

Window & desktop behaviour

  • The main window now has Windows' own title bar in the app's colours: Snap Layouts, double-click-to-maximise, Alt+Space, taskbar preview and window shake all work; the stray 7-pixel border around the maximised window is gone. (47cd6a4)
  • Each agent can get its own invisible desktop; a leftover-desktop banner with a "Clean up" button appears on the wall when abandoned desktops from earlier runs are found. (16b313a)
  • Clicking your own mouse into the live agent view no longer draws a second cursor under your real pointer; a click on the agent desktop is now 10-20× faster (546 ms → 41-53 ms, measured live); a click that silently reported success without effect is fixed and verified live. (7237b10)
  • The terminal (and every other list) now has a slim, dark scrollbar that only appears on hover/scroll instead of the bright Windows default; a terminal resize while scrolled up no longer yanks you back to the bottom. (db49371)
  • Optional tray icon (open, new task, pause agents, quit), optional "minimize to tray on close", a configurable global shortcut to bring the window back (Ctrl+Alt+A by default), and an autostart setting that adds a quiet (--leise) entry under HKCU\...\Run, removed again on uninstall, all off by default.

Reliability & performance

  • One Claude Code session now stays open for an entire run instead of starting a new process per step; each step sends only what changed since the last one. Measured: 7.1 s vs. 10.5 s per step, 110 vs. 1,908 tokens per step (a 24k-token session restarts itself before the cache stops paying off). Falls back to the old, proven per-step process on any session error. AGENTFENSTER_SITZUNG=0 disables it. (34d3c8c)
  • A test-suite crash that killed the whole run (a Windows fatal exception) is fixed: tests no longer touch the real Windows Credential Manager by accident, and the one test that deliberately does runs alone, not in the nightly's parallel worktrees. (d5258e7)
  • open_app notepad++ no longer opens the wrong program (Notepad instead of Notepad++), special characters like +/-/. are kept in a separate, precise name index. (0e3b6c9)

CAPTCHA handoff

  • A CAPTCHA pause now actively gets you: a Windows notification (click it to bring the live view to the front) plus an optional webhook URL (Ntfy, Pushover, your own receiver) in Settings. A Continue button in the live view lets you confirm you solved it without waiting for the next poll. The default wait before a run gives up rose from 3 to 10 minutes, and is now a Settings slider (captcha_wartezeit_min, 1-60). Detection hardened with more Cloudflare/hCaptcha/reCAPTCHA phrasings. Still no automatic solver, by decision, see README "What it cannot do".

Security

  • A malformed session key now returns 403 instead of a raw 500 with a traceback; /openapi.json (the full route map) is disabled; every error sentence is masked before it leaves the process; the local-only guard now checks the Host header itself, not just the port; request bodies over 8 MB are rejected before being read; workspace name/message/count all have upper bounds; an unknown terminal workspace no longer starts a shell. (a3dbe40)