fix: Boyer-Moore bad character shift was dead code in for-loop (#14770)

* fix: Boyer-Moore bad character shift was dead code in for-loop

The bad_character_heuristic() method used a for-loop with an
assignment to the loop variable i, which was immediately
overwritten by the next iteration. This caused the algorithm
to degrade from O(n/m) to O(n*m) naive search.

Changed to a while-loop so the shift actually takes effect.
Added max(i+1, shift) guard to prevent backward skips when
the mismatched character appears to the right of the mismatch
in the pattern. Added edge case doctests.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Update boyer_moore_search.py

* Correct documentation for positions in BoyerMooreSearch

Fixes the documentation to clarify that 'positions' contains the locations where the pattern was matched.

* Fix grammatical issues in Boyer-Moore search comments

Corrected grammatical errors in comments and docstrings.

---------

Co-authored-by: Christian Clauss <cclauss@me.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
This commit is contained in:
liu xiaochen
2026-09-16 22:44:47 +02:00
committed by GitHub
co-authored by Christian Clauss pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
parent 351805cd69
commit 681929de5d
+19 -5
View File
@@ -1,5 +1,5 @@
"""
The algorithm finds the pattern in given text using following rule.
Find the pattern in the given text using the following rule.
The bad-character rule considers the mismatched character in Text.
The next occurrence of that character to the left in Pattern is found,
@@ -11,7 +11,7 @@ If the mismatched character does not occur to the left in Pattern,
a shift is proposed that moves the entirety of Pattern past
the point of mismatch in the text.
If there is no mismatch then the pattern matches with text block.
If there is no mismatch, then the pattern matches the text block.
Time Complexity : O(n/m) average case with bad character heuristic
n=length of main string
@@ -29,7 +29,7 @@ class BoyerMooreSearch:
bms = BoyerMooreSearch(text="ABAABA", pattern="AB")
positions = bms.bad_character_heuristic()
where 'positions' contain the locations where the pattern was matched.
where 'positions' contains the locations where the pattern was matched.
"""
def __init__(self, text: str, pattern: str) -> None:
@@ -59,8 +59,8 @@ class BoyerMooreSearch:
def mismatch_in_text(self, current_pos: int) -> int:
"""
Find the index of mis-matched character in text when compared with pattern
from last.
Find the index of the mismatched character in text when compared with pattern
from the last.
Parameters :
current_pos (int): current index position of text
@@ -89,9 +89,23 @@ class BoyerMooreSearch:
>>> bms = BoyerMooreSearch(text="ABAABA", pattern="AB")
>>> bms.bad_character_heuristic()
[0, 3]
>>> bms = BoyerMooreSearch(text="AAAAA", pattern="AB")
>>> bms.bad_character_heuristic()
[]
>>> bms = BoyerMooreSearch(text="ABABAB", pattern="ABA")
>>> bms.bad_character_heuristic()
[0, 2]
>>> bms = BoyerMooreSearch(text="", pattern="AB")
>>> bms.bad_character_heuristic()
[]
>>> bms2 = BoyerMooreSearch(text="AAAAAA", pattern="AA")
>>> bms2.bad_character_heuristic()
[0, 1, 2, 3, 4]
>>> bms3 = BoyerMooreSearch(text="ABCDEF", pattern="XY")
>>> bms3.bad_character_heuristic()
[]