mahbot 0.7.3

An autonomous agentic engineering system that manages software development through role separation, subagents, and deterministic diagnostics.
Read file contents with line numbers. Preferred over shell `cat` for reading files — it provides line numbers for reference, supports offset/limit for partial reads, and can list code structure via AST symbols.

{{path_policy}}

When the path is a directory, the tool lists its contents instead of returning an error. The directory listing groups subdirectories and files with sizes and an extension summary. Note that `mode`, `offset`, and `limit` parameters only apply to file reads — they are silently ignored when a directory is passed.

Modes:
- `content` (default): Outputs a file or a range (using `offset`+`limit`) with line numbers. Large outputs are truncated to a small budget (~5 KB) — for big files, read in slices with `offset`/`limit`, or navigate via `symbols`/`zoom`. Handles any file type; binary files are read with lossy UTF-8 conversion, except the document containers converted below and a background shell session's own output file (read as a program's output). When the path is a directory, lists the directory contents.
- `symbols`: Lists all AST-level symbols (functions, structs, impl blocks, etc.) with line ranges — the quickest way to map a large file's structure without reading it whole. Works for supported code formats only (Rust, JS/TS, Python, Go, C, Ruby, SQL, Markdown, JSON, TOML, CSS, HTML, shell), not arbitrary formats.
- `zoom`: Extract a single symbol's full source by name (requires `symbol` parameter). Same supported-formats limit as `symbols`. Use with the output of `symbols` mode to drill into specific definitions.

When a content-mode path is a raster image (PNG, JPEG, or WebP), the tool reads and attaches it to the conversation as a native image the model can inspect, instead of returning lossy text. Other binary formats that cannot be decoded (e.g. GIF, BMP, HEIC) are reported as unsupported. Reading the same image twice returns a reference to the already-attached image rather than adding it again.

When a content-mode path is a document container — a PDF, a Word `.doc`/`.docx`/`.docm` document, an Excel `.xls`/`.xlsx`/`.xlsm` table, or a PowerPoint `.ppt`/`.pptx`/`.pptm` presentation (an encrypted one is reported as password-protected instead) — the tool converts it the same way an inbound chat attachment is converted, instead of returning its raw bytes. The extracted text is returned; when it is too long to inline, the tool writes it out to a file and reports that path. Every reading names the content it met and did not show, one line per kind right after the text ({{unshown_report}}), and it is silent when nothing was left out. An image the reading meant to deliver keeps the note that says why it is not there rather than a line of the report, and something a reading never looks at is not in the report either; formatting, themes and the service parts of a file stay out of the text, and what lies out of each format's reach is stated below. An old format is read only — this tool never edits one — and it comes back in the shape the same family's reading is described by below, as far as its own reader reaches; text is all it reads, so hidden sheets and hidden slides, merged ranges, comments, hyperlinks, defined names and formula text are neither read nor passed on. A legacy `.xls` is `Sheet "<name>":` blocks of one two-space-indented line per valued cell with its address (`A1: 42`), each value as Excel displayed it — a date as a date, a percentage as a percentage — a break in a value kept as `\n`, while its formula text is not recovered and nothing hidden, merged or annotated is marked. A legacy `.ppt` is a {{ppt_slide_label}} block per slide holding that slide's text lines, with its speaker notes under {{ppt_notes_label}}, and no mark for a title, a hidden slide or a diagram. A legacy `.doc` is its body text followed by whatever stories its reader finds, each as a labelled block — `{{doc_headers_footers}}:`, `{{doc_footnotes}}:`, `{{doc_endnotes}}:` and `{{doc_comments}}:`, one block per story rather than one per note or comment, with no comment author — and a `{{text_box_mark}}` block per text-box story, and then the names of the embedded objects it recognises. The images an old format embeds — and, for a `.xls`, the text its charts carry (chart titles, series and trendline names, and axis titles), and for a `.ppt`, the objects the deck holds — are never extracted, and the answer names what was left out instead of dropping it. A new-format presentation (`.pptx`/`.pptm`) yields each slide's text plus its speaker notes: a title the presentation declares for the slide is marked `{{ppt_title_mark}}`, a slide it does not show is marked `{{ppt_hidden_slide_mark}}` and keeps its own number, and the text a diagram on the slide shows follows under `{{ppt_diagram_text_mark}}` — a diagram whose text could not be read is named `{{ppt_diagram_text_lost_mark}}` instead of silently dropped, and a title merely written as ordinary text is not marked. A slide's layout is not read — the sample text a layout and its placeholders carry is not shown — and neither is the deck's theme. For a PDF the text comes page by page — a `Page <n>:` block for every page that produced any text at all, in page order and under the page's own physical number (the first page is 1), with even the shortest text kept, while a page that produced none gets no block and is still delivered as an image instead — followed, when the document carries marks, by an `Annotations:` section with one entry per mark as `<kind>[ (<author>)] on page <n>`, the kinds being `note`, `free text`, `highlight`, `underline`, `squiggly underline`, `strikeout`, `ink`, `square`, `circle`, `polygon`, `polyline`, `line`, `caret`, `stamp` and `signature`, with the mark's own text indented under its line and a mark that carries no text of its own still named by its kind and page — and followed, when the form holds values, by a `Form fields:` section with one line per field as `<field name>[ on page <n>]: <value>` (the name `pdf_form_fill` fills by), a text field showing its text, a checkbox or radio group showing `checked` or `not checked` plus the chosen option's name in parentheses where the file names one, a choice field showing the selected option, and a form whose values could not be read saying so in a note rather than looking like an empty one. Pages without a usable text layer, pages whose text could not be read — the answer names those pages, so an unreadable page is never mistaken for one with nothing to read — and images embedded in a page, are attached to the conversation as images when there are at most 5 of them. With more, all of them are saved in a folder and listed by path, so you can read them individually; when even that listing would not fit the answer, the folder alone is named — list it to find them. A page whose text was read is not drawn, so drawings, shapes and charts on it are outside what this tool sees: a PDF page gives its text, its embedded images, and the annotations and form values named below. `offset`/`limit` do not apply to a converted document. A document that yields neither text nor images, or whose conversion cannot run at all, is answered with a plain note saying so — never its raw bytes.

An Excel `.xlsx`/`.xlsm` table comes back as `Sheet "<name>":` blocks in workbook order, one two-space-indented line per valued cell with its address (`A1: 42`), a cell whose formula has no stored value shown as its formula text (`A2: =SUM(B1:B2)`) and one holding both as `A2: 42 (=SUM(B1:B2))`; a value whose own text holds a line break keeps it as `\n`, so every valued cell is one line whatever the workbook's text carries (`A3: two\nlines`); a sheet with no valued cell at all says so instead of standing empty (`Sheet "Empty": (no values)`), and a sheet that is one chart over the whole tab rather than cells says `{{chart_sheet_mark}}` — the sheet is that chart, so the report names it as a chart sheet and no reading draws it. Every value is the one Excel displays in the cell's own number format — a date as a date, always in the same unambiguous shape (`YYYY-MM-DD`, with the clock the value holds), a percentage as a percentage, a currency amount with its symbol and its thousands separators — so the displayed value and the stored one can differ: a percentage is stored as its fraction, a date as a serial number. The one exception is a cell that keeps its date as ISO text rather than as a serial: that text is shown as the workbook wrote it instead of through the cell's format. A displayed value copied out of this reading does not carry the stored number it was rendered from (a percentage reads as `12.5%` while the cell holds `0.125`), so a value written back must be the number the cell is meant to store, never the displayed text. A numeric cell whose format this reader cannot read, cannot reproduce for that value, or that shows nothing at all for it, has its stored value shown instead, marked `(stored number, format not shown)`; a cell with no format of its own (`General`) shows the number as the cell stores it. Hidden rows, columns and sheets are still read — nothing hidden is dropped — but never passed off as ordinary: a hidden sheet's header says `Sheet "<name>" (hidden):`, a very hidden one says `Sheet "<name>" (very hidden):`, and a sheet's hidden geometry is named once by range at the head of its block as `(hidden rows: 5, 12-14)` or `(hidden columns: C, F:G)`. A merged range is marked rather than repeated: the cell holding the value says `(merged A1:C1)`, a covered cell that kept a value of its own says `(in merged range A1:C1)`, and a covered cell holding nothing prints no line at all. A cell's comments follow the sheet's cells, one line each as `A1 comment (Ivan Petrov): …` or, when its author is not known, as `A1 comment: …`, whether Excel stored them as an older comment or as a modern threaded discussion. A hyperlink is marked on the cell it covers as `(link: https://…)`, with the caption the link carries named in front of the target where the cell's own value is not already what it says (`A1: Docs (link: Документы -> https://…)`); a link range no cell line covers — one over empty cells — gets a line of its own naming the cell, the caption the link carries when it has one, and the target (`C1: Go (link: Sheet2!A1)`, `A1: (link: https://…)`); and the workbook's defined names follow the last sheet as a `Defined names:` block, one `Rate: Data!$B$1` line per name, a sheet-scoped one naming its sheet in front of the name as `Порог (sheet "Служебный"): Служебный!$B$1`, with Excel's own service names (`_xlnm.…`) left out. What a workbook shows but its sheets do not hold is never read, so it is not in the report either: the values a pivot table displays, and the text a shape or a text box carries.

A Word `.docx`/`.docm` is read beyond its body. A table comes back as a `Table N:` block — one line per row under its `R1:` row prefix, one `C2:` token per cell — and a merged cell neither loses nor repeats its value: the cell that holds it says `(spans C1-C2)` and each continuation of the merge says `(part of R1C1)`. The headers and footers follow as `Header (default):`, `Header 2 (first page):`, `Footer (even pages):` — one block per part, labelled with every variant that uses it — then the footnotes, the endnotes and the comments as `Footnote 1:`, `Endnote 1:`, `Comment 1 (Ivan Petrov):`, each with its text indented under it — a definition nothing in the text refers to says `(not referenced in the text)` beside its author — while `[footnote 1]`, `[endnote 1]`, `[comment 1]` — or `[comment 1 starts]` … `[comment 1 ends]` for the range a comment covers, and `[footnote ?]` for a reference the package defines nothing for — mark the place in the text that note or comment belongs to. Word's own separator notes are not passed off as footnote text, and a page number or a date a field produces is reported as that field (`[field PAGE]`, `[field DATE \@ "dd.MM.yyyy"]`) rather than as ordinary text.

Unaccepted tracked changes are never passed off as the final text — in the body, in a table, in a header, a footnote or a comment alike: `[ins]…[/ins]` was inserted, `[del]…[/del]` was deleted, `[moved-here]…[/moved-here]` and `[moved-away]…[/moved-away]` were moved, `[ins ¶]`, `[del ¶]`, `[moved-here ¶]` and `[moved-away ¶]` revise a paragraph mark; a row or a cell that was inserted or deleted says `(inserted)` / `(deleted)`, a whole row or cell added or removed rather than a formatting change; and `[fmt]`, `[fmt ¶]`, `[fmt section]`, `(merge revised)` and `(formatting revised)` on a table, a row or a cell revise formatting alone. A document carrying any of them opens with an `Unaccepted tracked changes: …` line counting them and naming their authors, so a document still under revision is never taken for an agreed one. Text a text box holds is shown once, as an indented `{{text_box_mark}}` block.

Reading a credential-bearing file — `.env*`, a `.pem`/`.cer`/`.crt` certificate, `.key`/`.p12`/`.pfx`, or a known credential config (`.netrc`, `.npmrc`, `.pypirc`, `.git-credentials`, Maven `settings.xml`/`settings-security.xml`, Gradle `gradle.properties`/`init.gradle*`, cargo `credentials.toml`) — returns its content with the credential values masked. Every other file comes back as it is, and a secret stored under a name the masking does not recognise is returned whole.

Files larger than 10 MB are rejected, except document containers, which are accepted up to 50 MB.