【 使用环境 】生产环境
【 OB 】
【 使用版本 】4.2.1
【问题描述】该问题为企业版生产环境OCP metadb集群的告警。已通过工单咨询,原厂工程师大概解释为磁盘性能/繁忙问题引起。再次发帖是想和大家再讨论一下,请教一下我不太理解的地方。
因日志比较多,按照时间顺序摘出来了日志简要,大体如下,事务2106353433长时间未完成,阻塞了其他的事务的执行,告警出来IO异常。最后事务2106353433 commit失败。报错transaction log sync use too much time。主要不理解的点是最后这个报错,说事务日志同步消耗了太长时间。这里指的是副本日志流之间同步的时间吗。
日志大体如下:
[2026-07-29 11:55:26.463379] errcode=-6004 lock_for_read need retry block_txid:2106353433 request_txid:2106353437
[2026-07-29 11:55:27.464216] errcode=-6004 lock_for_read need retry block_txid:2106353433 request_txid:2106353440
[2026-07-29 11:55:28.464726] errcode=-6004 lock_for_read need retry block_txid:2106353433 request_txid:2106353441
[2026-07-29 11:55:29.464945] errcode=-6004 lock_for_read need retry block_txid:2106353433 request_txid:2106353438
[2026-07-29 11:55:30.229261] handle_timeout txid:2106353433
[2026-07-29 11:55:30.229366] handle_timeout trx is waiting log_cb txid:2106353433
[2026-07-29 11:55:30.465682] errcode=-6004 lock_for_read need retry block_txid:2106353433 request_txid:2106353438
[2026-07-29 11:55:30.529435] handle_timeout txid:2106353433
[2026-07-29 11:55:30.529539] handle_timeout trx is waiting log_cb txid:2106353433
[2026-07-29 11:55:30.660044] check_clog_disk_full_ errcode=-4009 unexpected io error block_txid:2106353433 request_txid:2106353434
[2026-07-29 11:55:30.660209] check_with_tx_data errcode=-4009 do data check function fail txid:2106353433
[2026-07-29 11:55:30.660238] check_with_tx_data errcode=-4009 check with tx data failed ret=“OB_IO_ERROR” txid:2106353433
[2026-07-29 11:55:30.660246] check_with_tx_data errcode=-4009 check tx data in tables failed ret=“OB_IO_ERROR” txid:2106353433
[2026-07-29 11:55:30.660259] lock_for_read errcode=-4009 failed to lock for read txid:2106353433
[2026-07-29 11:55:30.660273] lock_for_read_inner_ errcode=-4009 lock for read failed block_txid:2106353433 request_txidtxi:2106353434
[2026-07-29 11:55:31.029697] handle_timeout errcode=0 clog disk has fatal error, make scheduler retry commit txid:2106353433
[2026-07-29 11:55:31.029897] handle_tx_commit_result_ handle tx commit result txid:2106353433 commit_fin=false, result=-4038
[2026-07-29 11:55:31.129820] handle_tx_commit_timeout handle tx commit timeout(ret=0, tx_id={txid:2106353433}
[2026-07-29 11:55:31.129890] commit errcode=-4023 tx is 2pc logging txid:2106353433
[2026-07-29 11:55:31.129995] commit trx commit failed ret=-4023, ret=“OB_EAGAIN” txid:2106353433 callback_alloc_count=1
[2026-07-29 11:55:31.130050] local_ls_commit_tx_ txid:2106353433
[2026-07-29 11:55:31.130059] handle_trans_commit_request handle trans commit request failed txid:2106353433
[2026-07-29 11:55:31.130074] process errcode=-4023 handle txn message fail ret=-4023, ret=“OB_EAGAIN” txid:2106353433
[2026-07-29 11:55:31.130146] handle_tx_commit_result_ handle tx commit result txid:2106353433 commit_fin=false, result=-4023
[2026-07-29 11:55:31.130198] handle_trans_msg_callback handle trans msg callback ret=0, elapsed_ts=0 txid:2106353433
[2026-07-29 11:55:32.452147] test_lock errcode=-4389 transaction log sync use too much time txid:2106353433 log_sync_used_time=6717853, ctx_lock_wait_time=4