406 Commits
Author SHA1 Message Date
afourney b8f79c57eb Bump versions or markitdown and markitdown-ocr (#2545) v0.1.8 2026-09-21 14:02:08 -07:00
Nithin JambulaandAdam Fourney d51937de3a fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support (#2407)
* fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support
* Add 3.14 to test matrix, and limit advertised compatibility.

---------

Co-authored-by: Adam Fourney <adamfo@microsoft.com>
2026-09-21 09:47:07 -07:00
afourney 945314a45d Refactor Office OCR converters to reuse core conversion pipelines (#2506)
* Add _image_to_html helper for docx ocr conversion.
* Ensure compatibility with core markitdown version
* Have pptx, xlsx and docx expose an _image_to_html semi-private method, overridable by plugins.
* Fix conditional check for llm_description
* Update relationships function calls for drawings and images
* Forward stream_info to LLMVisionOCRService
2026-09-16 10:23:08 -07:00
afourney eb31b5c945 bump_versions (#2490) v0.1.8b2 2026-09-14 09:48:43 -07:00
afourney cc0ca9edd8 Preserve whitepace underlines (#2477)
* Preserve whitespace underlines
2026-09-12 13:25:31 -07:00
afourney 3a6ce44358 Throw an error when llm client fails with image converter. (#2476) 2026-09-12 12:11:30 -07:00
afourney 5640da7145 fix(youtube): fall back to HTML when no video content is extracted (#2469) 2026-09-11 15:52:59 -07:00
afourney 485a7b93b6 fix(docx): preserve namespaces when repairing stylesheets (#2467)
* fix(docx): preserve namespaces when repairing stylesheets
* Fixed d:strike as well
2026-09-11 15:05:05 -07:00
afourney 75e7114598 fix: avoid splitting UTF-8 characters during charset detection (#2466) 2026-09-11 14:35:07 -07:00
afourney 2277fff52d Improve efficiency of escaping pipes in CSVs (#2464) 2026-09-11 12:24:26 -07:00
dickbown f0c5d01fbe perf(csv): Batch-remove blank rows to eliminate the quadratic overhead in table conversion 2026-09-11 11:12:55 -07:00
afourney 40658cdcc5 Fix ANSI Outlook MSG decoding for Japanese code pages and padded strings (#2462) 2026-09-11 11:05:33 -07:00
afourney 9edea6f5f6 Pin cryptography to 46.0.3 (in the CI, only) to temporarily enable ARM tests of the rest of the suite (#2460) 2026-09-11 10:22:13 -07:00
afourney 8c11f50275 Add arm to matrix. Simplify matrix. (#2459)
* Add arm to matrix. Simplify matrix.
* Fixed a broken powershell CLI, which caused the test suite to exit without running tests.
2026-09-11 10:11:29 -07:00
afourney 9480644d9c Added file paths tests. (#2454) 2026-09-10 15:52:54 -07:00
afourney 73a26dac09 Expand test matrix to include windows-latest (#2453) 2026-09-10 15:36:19 -07:00
afourney e99a726687 Reject UNC and Windows device paths (similar to how netlocs are already rejected). 2026-09-10 15:02:35 -07:00
Çağdaş Yürekli 57a481e1aa fix(mcp): migrate to MCP SDK 2.x so 2026-07-28 clients can connect (#2363) 2026-09-10 12:20:11 -07:00
kevin 075c05bd4d fix(pptx): do not emit a heading for a slide with an empty title (#2442)
* fix(pptx): do not emit a heading for a slide with an empty title
2026-09-10 09:46:52 -07:00
liyrds 6270920c28 fix(zip): preserve content of entries with duplicate filenames (#2434) 2026-09-09 23:54:16 -07:00
Machen John b59e0643dc fix: prefer data-src over placeholder data URI in img src (#2417)
* fix: prefer data-src over placeholder data URI in img src
* Keep data URIs is keep_data_uris is True
2026-09-09 23:30:55 -07:00
freetg71527152-ui f217a288a7 Fix --list-plugins help text to reference correct --use-plugins flag (#2381) 2026-09-09 22:49:11 -07:00
Sushant Lokhande fef3a5de16 fix(epub): resolve percent-encoded manifest hrefs to ZIP entries (#2413) 2026-09-09 22:37:57 -07:00
SpongeBob f9042b4227 Fix PPTX shape sorting treating top-zero as missing (#2408) 2026-09-09 16:34:11 -07:00
kevinandafourney 24b9e122ea Do not convert an undecodable file to the word "None" (#2418)
* Do not convert an undecodable file to the word "None"
* Inline the fallback conversions.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-09 16:28:27 -07:00
SpongeBob 0504d7d50f Fix Zip converter forwarding kwargs to nested conversions (#2409)
* Fix Zip converter forwarding kwargs to nested conversions
* Add regression test for Zip converter kwargs forwarding
* Fixed args construction.
* test(zip): cover DOCX member conversion with and without an archive URL
2026-09-09 16:08:01 -07:00
Erwin Kersten 83c7730093 fix(docker): move base image off EOL bullseye (#2421) 2026-09-09 15:32:41 -07:00
kevin 5dc7bafcaf fix(epub): forward conversion options to the HTML converter (#2426) 2026-09-09 15:30:01 -07:00
afourney 75cc34a364 fix(rss): preserve complete feed content and resolve relative link (#2432)
* fix(rss): preserve complete feed content and resolve relative link
* Fix extra whitespace, failing tests.
2026-09-09 15:20:57 -07:00
kevin 2e71e117ea fix(ipynb): strip the UTF-8 BOM so a notebook is not emitted as raw JSON (#2425) 2026-09-09 11:33:59 -07:00
kevin 481f22a9a9 fix(pptx): do not emit a Notes heading for a slide with no notes (#2427) 2026-09-09 11:29:47 -07:00
kevin a04fd8bd19 fix(rss): drop layout whitespace from feed and entry titles (#2428)
* fix(rss): drop layout whitespace from feed and entry titles
* Refactor return value handling in RSS converter
* Do not flatten descriptions.
2026-09-09 11:23:30 -07:00
Xing Zheng cb785cb2f7 fix(rss): support namespace-prefixed Atom feeds (#2429) 2026-09-09 09:49:30 -07:00
dependabot[bot] a2a7a5293d chore(deps): bump actions/setup-python from 5.6.0 to 7.0.0 (#2431)
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 5.6.0 to 7.0.0.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](https://github.com/actions/setup-python/compare/v5.6.0...5fda3b95a4ea91299a34e894583c3862153e4b97)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
2026-09-09 09:40:50 -07:00
Kim Do Yeon ab6792caca Fix CSV parsing for CR-only line endings (#2412)
Enable universal newline handling without translation in the CSV text stream. Add public API regression coverage for LF, CRLF, and CR records, optional final newlines, and quoted multiline cells.

Fixes #2411.
2026-09-09 09:36:37 -07:00
afourney 21887b3bf0 Preserve all CSV columns by matching the widest row (#2430) 2026-09-09 08:40:53 -07:00
afourney b6e8bbdce6 Updated README (#2393)
* MarkItDown is now a standalone project, and the README is updated to reflect this.
2026-09-06 21:58:12 -07:00
kapil971390 4459ed0115 fix(pptx): chart_title.text_frame is never None -- use has_text_frame (#2385)
The #2194 guard (chart.chart_title.text_frame is not None) doesn't
check what it's meant to. python-pptx's ChartTitle.text_frame is
documented as destructive -- it creates a text frame if one isn't
already present -- so it can never return None against the real,
installed python-pptx (1.0.2, unpinned in pyproject.toml).

has_text_frame is the property that actually reflects presence,
without the auto-creating side effect. Swaps both converters to use
it, updates the existing mock-based tests to match the real API
contract (has_text_frame=False instead of text_frame=None), and adds
a positive-path test for each so both branches are covered.

Verified against the real, installed python-pptx (not just source
reading): with has_title=True and no title text ever set,
chart_title.text_frame is None is False and .text_frame.text is ''
-- confirming the existing guard is a no-op.
2026-09-04 10:04:57 -07:00
afourney a035350629 Remove uv lockfile that was accidentally committed. (#2380) 2026-09-03 20:57:50 -07:00
afourney e153a92a18 Be sure to run the proper Python version. (#2379) v0.1.8b1 2026-09-03 20:36:11 -07:00
afourney bc90c5d7a5 Test markitdown-ocr in CI against the sibling markitdown (#2378)
* Test markitdown-ocr in CI against the sibling markitdown
* Run the full matrix.
2026-09-03 20:28:53 -07:00
afourney 2dffd8bb6e Bump versions. (#2377) 2026-09-03 17:21:53 -07:00
afourney 0e9ede3e05 Update README to help scope PRs. (#2376)
* Update README to help scope PRs.
2026-09-03 16:26:27 -07:00
Nefelibataandafourney 58783e3272 fix(docx): handle unknown math functions in OMML converter without crashing (#2268)
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter

Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.

Fixes #1904

* fix(docx): handle unknown math functions in OMML converter without crashing

Add common standard math functions to FUNC dictionary in latex_dict.py and fallback to \operatorname{} for unknown functions in omml.py instead of throwing NotImplementedError.

Fixes #1982

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 15:50:58 -07:00
hanhan761 dd2df72913 fix: ImageConverter gracefully handles LLM API failures (fixes #1942) (#1948) 2026-09-03 15:33:16 -07:00
Lazizbek Ergashevandafourney 697a2a0e0b fix(rss): treat Atom text content and summary as plain text (#2374)
* fix(rss): treat Atom text content and summary as plain text
* Parse content types more carefully.
* Don't pass text through the html converter.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 15:01:34 -07:00
hanhan761andafourney 94a84834a2 chore: remove unused mammoth import from PlainTextConverter (#1951) (#1953)
* fix: handle LLM API errors gracefully in ImageConverter (#1942)

If the OpenAI API call fails (auth error, rate limit, network timeout, etc.), the exception propagates and crashes the entire conversion. Wrap the call in try/except to gracefully skip the LLM caption, matching the pattern already used by PptxConverter.

* chore: remove unused mammoth import from PlainTextConverter (#1951)

The mammoth import was copied from DocxConverter but PlainTextConverter does not need it. Remove the dead code and unused _dependency_exc_info variable.

* Revert unrelated changes.

---------

Co-authored-by: afourney <adamfo@microsoft.com>
2026-09-03 12:58:47 -07:00
c870d49284 fix: suppress pydub RuntimeWarning when ffmpeg is missing (fixes #1685) (#1985)
* fix: suppress pydub RuntimeWarning when ffmpeg is missing

Closes #1685

Importing pydub emits a RuntimeWarning when ffmpeg/avconv is not
installed on the system. This warning appears even when audio
transcription is not actively being used.

_transcribe_audio.py already wraps the imports in a
warnings.catch_warnings block that suppresses DeprecationWarning and
SyntaxWarning. Add a filter for the ffmpeg RuntimeWarning as well.

* Modify RuntimeWarning filter for ffmpeg

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: afourney <adamfo@microsoft.com>
Co-authored-by: afourney <adam.fourney@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-03 12:41:55 -07:00
Mukunda Rao Katta 7de6ce4943 Support short YouTube URLs in converter (#1882) 2026-09-03 12:22:34 -07:00
a5ddc8efd5 fix: handle DOCX files with inconsistent ZIP filename casing (#1812) (#2016)
* fix: handle DOCX files with inconsistent ZIP filename casing (#1812)

Some document generators (e.g. certain Microsoft Word versions, legal
document systems) produce .docx files where the central directory records
one casing (e.g. 'customXml/item2.xml') but the local file headers record
another (e.g. 'customXML/item2.xml'). Python's zipfile module raises
BadZipFile when reading such files.

Add _fix_zip_filename_casing() to patch local file header filenames to
match the central directory before any ZIP processing occurs.

* test: register DOCX zip casing test in __main__ runner list

* fix: encode central directory filenames with the ZIP's actual charset

ZIP filenames are stored as cp437 unless general purpose flag bit 11
(0x800) marks them UTF-8, which is how zipfile decoded orig_filename.
Re-encoding unconditionally as UTF-8 produced the wrong bytes for
non-ASCII cp437 names, so a local header could be skipped or patched
with mismatched bytes.

* fix(ocr): pre-process the DOCX before reading it in DocxConverterWithOCR

convert() read the caller's stream twice before repairing it: once for the
embedded style map, and once to extract images for OCR. A .docx whose ZIP
local file headers disagree with the central directory on casing would then
fail those reads. The image extraction swallows every exception, so the
result was an empty OCR map and silently missing OCR blocks -- output
identical to the no-OCR path, with no error.

Call pre_process_docx() once at the top and feed the repaired stream to
every read, matching the ordering the base DocxConverter already uses. This
also drops a redundant second rebuild of the archive on both branches.

* Improve name comparison for casing mismatch

* test: assert non-case ZIP filename mismatches are still rejected

Equal encoded length does not imply two filenames differ only in case, so
the local file header repair must not wave through archives whose local and
central directory names genuinely disagree. Cover an unrelated same-length
name, a single differing character, and a difference outside the cased
characters, alongside the case-only mismatch that should still be repaired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: lyydsheep <lyydsheep@lyydsheepdeMacBook-Pro.local>
Co-authored-by: Adam Fourney <adamfo@microsoft.com>
2026-09-03 12:18:46 -07:00