fix: raise ValueError in encode() for non-lowercase input (#14936)

* fix: raise ValueError in encode() for non-lowercase input

encode() previously accepted uppercase letters, digits, and other
non-lowercase characters silently, producing incorrect/out-of-range
values (e.g. negative numbers for uppercase letters) instead of
failing. Add input validation using str.islower() and str.isalpha()
to raise a ValueError when the input isn't purely lowercase a-z.
Added a doctest covering the new error case.

* resolved doctest

* Fix Ruff 0.16 lint failures

* Enhance encode function error handling examples

Update error handling in encode function to include examples for mixed case and invalid characters.

* Fix indentation in test_cancer_data function

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: Christian Clauss <cclauss@me.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
This commit is contained in:
Ali Satwat Khan
2026-07-31 18:07:20 +02:00
committed by GitHub
co-authored by Christian Clauss pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
parent 758d487394
commit eea1bacfe0
2 changed files with 15 additions and 1 deletions
+14
View File
@@ -13,7 +13,21 @@ def encode(plain: str) -> list[int]:
"""
>>> encode("myname")
[13, 25, 14, 1, 13, 5]
>>> encode("abCd")
Traceback (most recent call last):
...
ValueError: plain must contain only lowercase letters (a-z)
>>> encode("n0w")
Traceback (most recent call last):
...
ValueError: plain must contain only lowercase letters (a-z)
>>> encode("later!")
Traceback (most recent call last):
...
ValueError: plain must contain only lowercase letters (a-z)
"""
if not plain.islower() or not plain.isalpha():
raise ValueError("plain must contain only lowercase letters (a-z)")
return [ord(elem) - 96 for elem in plain]
@@ -451,7 +451,7 @@ def test_cancer_data():
print("Hello!\nStart test SVM using the SMO algorithm!")
# 0: download dataset and load into pandas' dataframe
if not os.path.exists(r"cancer_data.csv"):
request = urllib.request.Request( # noqa: S310
request = urllib.request.Request(
CANCER_DATASET_URL,
headers={"User-Agent": "Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)"},
)