Details
-
Bug
-
Status: Open (View Workflow)
-
Major
-
Resolution: Unresolved
-
10.11, 13.1
-
None
-
None
Description
This problem was found by Claude:
- Important [sql/item_timefunc.cc:1824] format_length() walks the format one byte at a time (const char *ptr ... ptr++, *ptr != '%') instead of one character at a time, so for utf16/utf32 format strings it never sees %M/%W/%a/%b as specifiers and underestimates the length. Verified with -
cursor-protocol: DATE_FORMAT('2004-09-04',_utf32 0x000000250000004D) returns Septemb, the %W case returns Wednesd, _utf16 0x0025004D returns Sep, and the ru_RU result is сентябр, all silently cut. The commit says it covers "utf32 result character sets", but the utf32 cases only pass because byte scanning overcounts %h. The loop should decode with format>charset()>cset>mb_wc() the same way make_date_time does (line 496), so both functions parse the format identically. (cited from correctness-and-security.md:Buffer / length validation, S6) — by silent-failure-hunter
- Important [sql/item_timefunc.cc:1873] %T still reserves 8 characters, but make_date_time prints %02d:%02d:%02d with TIME hours up to 838. With --cursor-protocol, TIME_FORMAT(TIME'838:59:59','%T') returns 838:59:5 and TIME_FORMAT(TIME'-838:59:59','%T') returns -838:59: on the pre-fix build. After this patch the positive case fits only by accident (it uses the new sign slot), and the negative case is still cut to -838:59:5. So
MDEV-34215's "negative TIME truncated in cursor protocol" bug remains for %T; the reservation should be size += 9 (hhh:mm:ss). (cited from correctness-and-security.md:Buffer / length validation, S4) — by silent-failure-hunter
This script demonstrates the problem that format_length() returns a wrong value for utf32:
SET NAMES utf8mb3; |
CREATE OR REPLACE TABLE t1 AS |
SELECT
|
DATE_FORMAT('1997-10-04 22:23:00', '%T') AS c1, |
DATE_FORMAT('1997-10-04 22:23:00', CONVERT('%T' USING utf32)) AS c2; |
DESCRIBE t1;
|
+-------+-------------+------+-----+---------+-------+
|
| Field | Type | Null | Key | Default | Extra |
|
+-------+-------------+------+-----+---------+-------+
|
| c1 | varchar(8) | YES | | NULL | |
|
| c2 | varchar(80) | YES | | NULL | |
|
+-------+-------------+------+-----+---------+-------+
|
The expected result is two varchar columns of the same size for c1 and c2.
For multi-byte character sets other than ucs2/utf16/utf32 it also does not work correct:
SET NAMES utf8mb3; |
CREATE OR REPLACE TABLE t1 AS |
SELECT
|
DATE_FORMAT('1997-10-04 22:23:00', 'I') AS c1 , -- just one letter, expect varchar(1) |
DATE_FORMAT('1997-10-04 22:23:00','ß') AS c2; -- also just one letter, expect varchar(1) |
|
|
DESCRIBE t1;
|
+-------+------------+------+-----+---------+-------+
|
| Field | Type | Null | Key | Default | Extra |
|
+-------+------------+------+-----+---------+-------+
|
| c1 | varchar(1) | YES | | NULL | |
|
| c2 | varchar(2) | YES | | NULL | | -- Oops. Expected varchar(1)
|
+-------+------------+------+-----+---------+-------+
|