* fix: bump youtube-transcript-api to >=1.2.3 for Python 3.14 support
* Add 3.14 to test matrix, and limit advertised compatibility.
---------
Co-authored-by: Adam Fourney <adamfo@microsoft.com>
* Add _image_to_html helper for docx ocr conversion.
* Ensure compatibility with core markitdown version
* Have pptx, xlsx and docx expose an _image_to_html semi-private method, overridable by plugins.
* Fix conditional check for llm_description
* Update relationships function calls for drawings and images
* Forward stream_info to LLMVisionOCRService
* Fix Zip converter forwarding kwargs to nested conversions
* Add regression test for Zip converter kwargs forwarding
* Fixed args construction.
* test(zip): cover DOCX member conversion with and without an archive URL
Enable universal newline handling without translation in the CSV text stream. Add public API regression coverage for LF, CRLF, and CR records, optional final newlines, and quoted multiline cells.
Fixes#2411.
The #2194 guard (chart.chart_title.text_frame is not None) doesn't
check what it's meant to. python-pptx's ChartTitle.text_frame is
documented as destructive -- it creates a text frame if one isn't
already present -- so it can never return None against the real,
installed python-pptx (1.0.2, unpinned in pyproject.toml).
has_text_frame is the property that actually reflects presence,
without the auto-creating side effect. Swaps both converters to use
it, updates the existing mock-based tests to match the real API
contract (has_text_frame=False instead of text_frame=None), and adds
a positive-path test for each so both branches are covered.
Verified against the real, installed python-pptx (not just source
reading): with has_title=True and no title text ever set,
chart_title.text_frame is None is False and .text_frame.text is ''
-- confirming the existing guard is a no-op.
* fix(doc-intel): default api_version to None in DocumentIntelligenceConverter
Do not set a default api_version string in DocumentIntelligenceConverter. If api_version is omitted or None, avoid passing api_version kwarg to DocumentIntelligenceClient so Azure SDK uses its native default version.
Fixes#1904
* fix(docx): handle unknown math functions in OMML converter without crashing
Add common standard math functions to FUNC dictionary in latex_dict.py and fallback to \operatorname{} for unknown functions in omml.py instead of throwing NotImplementedError.
Fixes#1982
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix(rss): treat Atom text content and summary as plain text
* Parse content types more carefully.
* Don't pass text through the html converter.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: handle LLM API errors gracefully in ImageConverter (#1942)
If the OpenAI API call fails (auth error, rate limit, network timeout, etc.), the exception propagates and crashes the entire conversion. Wrap the call in try/except to gracefully skip the LLM caption, matching the pattern already used by PptxConverter.
* chore: remove unused mammoth import from PlainTextConverter (#1951)
The mammoth import was copied from DocxConverter but PlainTextConverter does not need it. Remove the dead code and unused _dependency_exc_info variable.
* Revert unrelated changes.
---------
Co-authored-by: afourney <adamfo@microsoft.com>
* fix: suppress pydub RuntimeWarning when ffmpeg is missing
Closes#1685
Importing pydub emits a RuntimeWarning when ffmpeg/avconv is not
installed on the system. This warning appears even when audio
transcription is not actively being used.
_transcribe_audio.py already wraps the imports in a
warnings.catch_warnings block that suppresses DeprecationWarning and
SyntaxWarning. Add a filter for the ffmpeg RuntimeWarning as well.
* Modify RuntimeWarning filter for ffmpeg
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: afourney <adamfo@microsoft.com>
Co-authored-by: afourney <adam.fourney@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix: handle DOCX files with inconsistent ZIP filename casing (#1812)
Some document generators (e.g. certain Microsoft Word versions, legal
document systems) produce .docx files where the central directory records
one casing (e.g. 'customXml/item2.xml') but the local file headers record
another (e.g. 'customXML/item2.xml'). Python's zipfile module raises
BadZipFile when reading such files.
Add _fix_zip_filename_casing() to patch local file header filenames to
match the central directory before any ZIP processing occurs.
* test: register DOCX zip casing test in __main__ runner list
* fix: encode central directory filenames with the ZIP's actual charset
ZIP filenames are stored as cp437 unless general purpose flag bit 11
(0x800) marks them UTF-8, which is how zipfile decoded orig_filename.
Re-encoding unconditionally as UTF-8 produced the wrong bytes for
non-ASCII cp437 names, so a local header could be skipped or patched
with mismatched bytes.
* fix(ocr): pre-process the DOCX before reading it in DocxConverterWithOCR
convert() read the caller's stream twice before repairing it: once for the
embedded style map, and once to extract images for OCR. A .docx whose ZIP
local file headers disagree with the central directory on casing would then
fail those reads. The image extraction swallows every exception, so the
result was an empty OCR map and silently missing OCR blocks -- output
identical to the no-OCR path, with no error.
Call pre_process_docx() once at the top and feed the repaired stream to
every read, matching the ordering the base DocxConverter already uses. This
also drops a redundant second rebuild of the archive on both branches.
* Improve name comparison for casing mismatch
* test: assert non-case ZIP filename mismatches are still rejected
Equal encoded length does not imply two filenames differ only in case, so
the local file header repair must not wave through archives whose local and
central directory names genuinely disagree. Cover an unrelated same-length
name, a single differing character, and a difference outside the cased
characters, alongside the case-only mismatch that should still be repaired.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: lyydsheep <lyydsheep@lyydsheepdeMacBook-Pro.local>
Co-authored-by: Adam Fourney <adamfo@microsoft.com>