* [Tarantool-patches] [PATCH luajit] dbg: avoid hardcoded enums (get them from target)
@ 2026-09-09 8:49 Mikhail Elhimov via Tarantool-patches
2026-09-17 9:52 ` Sergey Bronnikov via Tarantool-patches
0 siblings, 1 reply; 4+ messages in thread
From: Mikhail Elhimov via Tarantool-patches @ 2026-09-09 8:49 UTC (permalink / raw)
To: Sergey Kaplun, Sergey Bronnikov, Evgeniy Temirgaleev; +Cc: tarantool-patches
Besides reducing lines of code this way the extension become compatible
with various version of luajit because different version might use
different set of enum members (newer version might introduce additional
BC/IR/etc.)
Prior to this patch string was used as a debugger-agnostic way to
specify type, but this way might not work in case of enum because
debugging information might be optimized out if no variable of such enum
type is declared and its members are used only as a predefined
constants. The type of enum is needed to be able to get human readable
value (member name instead of number). The alternative way to get enum
type is get it from the value, i.e. somehow obtain the value that is of
enum type and then get its type object.
To do that, separate method to create enum value was introduced in
Debugger API, LLDB value class was monkey-patched to get value type in
the same way as GDB value class does and 'cast' method was adjusted to
accept also type object, not only string.
Other adjustments:
- [lldb] 'eval' adjusted to return object of the same type as 'cast'
method (this improves consistency)
- [gdb] in 'eval' method dropped check of the value returned by
gdb.parse_and_eval() as it would fail also for expression like '0'
(looks like a kind of legacy code that is not needed now).
- [gdb/lldb] renamed eval argument to reflect its meaning (it is
an expression rather than a command).
Closes tarantool/tarantool#13094
---
This patch is to be applied after the 'fix mapping of FPMATHOP' patch.
Branch: https://github.com/tarantool/luajit/tree/elhimov/gh-13094-dbg-use-enums-from-inferior
Related issue: https://github.com/tarantool/tarantool/issues/13094
src/luajit_dbg.py | 628 +++++++++++-----------------------------------
1 file changed, 141 insertions(+), 487 deletions(-)
diff --git a/src/luajit_dbg.py b/src/luajit_dbg.py
index 76001b7d..1ac1d275 100644
--- a/src/luajit_dbg.py
+++ b/src/luajit_dbg.py
@@ -110,9 +110,21 @@ class Debugger(object):
self.write('{} command initialized\n'.format(name))
self.write('LuaJIT debug extension is successfully loaded\n')
+ def cast(self, tp, val):
+ '''Cast the value to the given type (it is either C type string
+ or the debugger-specific type object).'''
+ if isinstance(tp, str):
+ tp = self._dbgtype(tp)
+ return self._cast(tp, val)
+
+ @abc.abstractmethod
+ def _cast(self, tp, val):
+ '''Cast the value to the debugger-specific type.'''
+ pass
+
@abc.abstractmethod
- def cast(self, typestr, val):
- '''Cast the value to the required C type.'''
+ def _dbgtype(self, typestr):
+ '''Convert C type string into debugger-specific type object.'''
pass
@abc.abstractmethod
@@ -146,8 +158,9 @@ class Debugger(object):
pass
@abc.abstractmethod
- def eval(self, command):
- '''Parse and evaluate the given debugger command.'''
+ def eval(self, expr):
+ '''Parse and evaluate the given debugger expression.
+ Return debugger-specific value.'''
pass
@abc.abstractmethod
@@ -187,6 +200,13 @@ class Debugger(object):
'''Register the command with the corresponding name.'''
pass
+ @abc.abstractmethod
+ def create_enum_value(self, enum_name, enum_member_name):
+ '''Return debugger-specific value object that represents
+ the given enum member.
+ '''
+ pass
+
@abc.abstractproperty
def LJBase(self):
'''Base command class.
@@ -211,8 +231,9 @@ class _GDBDebugger(Debugger):
super(_GDBDebugger, self).__init__()
self.CONNECTED = False
- def cast(self, typestr, val):
- return gdb.Value(val).cast(self._dbgtype(typestr))
+ def _cast(self, tp, val):
+ assert isinstance(tp, gdb.Type)
+ return gdb.Value(val).cast(tp)
def sizeof(self, typestr):
return self._dbgtype(typestr).sizeof
@@ -268,14 +289,11 @@ class _GDBDebugger(Debugger):
else:
return None
- def eval(self, command):
- if not command:
+ def eval(self, expr):
+ if not expr:
return None
- ret = gdb.parse_and_eval(command)
- if not ret:
- raise gdb.GdbError('table argument empty')
- return ret
+ return gdb.parse_and_eval(expr)
def detect_arch(self):
if hasattr(self, 'arch'):
@@ -340,6 +358,9 @@ class _GDBDebugger(Debugger):
def register_command(self, command, name):
command(name)
+ def create_enum_value(self, enum_name, enum_member_name):
+ return self.eval(enum_member_name)
+
class LJBase(gdb and gdb.Command or object):
def __init__(ljbase, name):
# XXX Fragile: Though the command initialization looks
@@ -378,6 +399,10 @@ class _LLDBDebugger(Debugger):
lldb.eBasicTypeInt128
]
+ def _lldb_tp_isenum(self, tp):
+ return tp.GetCanonicalType().GetTypeClass() == \
+ lldb.eTypeClassEnumeration
+
def _lldb_value_from_raw(self, raw_value, size, tp):
isfp = self._lldb_tp_isfp(tp)
if isfp:
@@ -491,8 +516,7 @@ class _LLDBDebugger(Debugger):
# Instead of default GetSummary.
if not lldbval.sbvalue.TypeIsPointerType():
tp = lldbval.sbvalue.GetType()
- is_float = self._lldb_tp_isfp(tp)
- if is_float:
+ if self._lldb_tp_isfp(tp) or self._lldb_tp_isenum(tp):
return lldbval.sbvalue.GetValue()
else:
return str(int(lldbval))
@@ -530,6 +554,9 @@ class _LLDBDebugger(Debugger):
else:
return int(lldbval) - int(other)
+ def lldb_gettype(lldbval):
+ return lldbval.sbvalue.type
+
super(_LLDBDebugger, self).__init__()
self.target = lldb.debugger.GetSelectedTarget()
# Monkey-patch the lldb.value class.
@@ -545,6 +572,7 @@ class _LLDBDebugger(Debugger):
lldb.value.__ror__ = lldb__or__ # Same semantics.
lldb.value.__str__ = lldb__str__
lldb.value.__sub__ = lldb__sub__
+ lldb.value.type = property(lldb_gettype)
def lldb_major_version():
version_string = lldb.SBDebugger.GetVersionString()
@@ -568,11 +596,11 @@ class _LLDBDebugger(Debugger):
self.dbgtype_cache[typestr] = dbgtype
return dbgtype
- def cast(self, typestr, val):
+ def _cast(self, tp, val):
+ assert isinstance(tp, lldb.SBType)
if isinstance(val, lldb.value):
val = val.sbvalue
elif type(val) is int:
- tp = self._dbgtype(typestr)
return self._lldb_value_from_raw(val, tp.GetByteSize(), tp)
elif not isinstance(val, lldb.SBValue):
raise Exception(
@@ -582,7 +610,6 @@ class _LLDBDebugger(Debugger):
# XXX: Simply SBValue.Cast() works incorrectly since it
# may take the 8 bytes of memory instead of 4, before the
# cast. Construct the value on the fly.
- tp = self._dbgtype(typestr)
if self._lldb_tp_isfp(tp):
rawval = float(val.GetValue())
elif self._lldb_tp_issigned(tp):
@@ -670,15 +697,15 @@ class _LLDBDebugger(Debugger):
else:
return None
- def eval(self, command):
- if not command:
+ def eval(self, expr):
+ if not expr:
return None
process = self.target.GetProcess()
thread = process.GetSelectedThread()
frame = thread.GetSelectedFrame()
- ret = frame.EvaluateExpression(command)
- return ret
+ ret = frame.EvaluateExpression(expr)
+ return lldb.value(ret)
def detect_arch(self):
if hasattr(self, 'arch'):
@@ -724,6 +751,37 @@ class _LLDBDebugger(Debugger):
)
)
+ def create_enum_value(self, enum_name, enum_member_name):
+ val = self.eval(enum_name + "::" + enum_member_name)
+ if val.sbvalue.IsValid() and val.sbvalue.error.Success():
+ return val
+
+ # LLDB uses enum name in expression above but debugging information
+ # about enum name migth be optimized out if no variable of the given
+ # enum type is declared and its members are only used as the predefined
+ # constants (like IRFieldID).
+
+ # In this case the above method doesn't work so trying to discover
+ # enum type by the given enum member.
+
+ def find_enum_type_member(enum_type, enum_member_name):
+ # SBTypeEnumMemberList supports members iteration and [] access
+ # (both by index and by member name) only starting from lldb-12
+ # so this implementation is used to handle earlier versions.
+ members = enum_type.GetEnumMembers()
+ for i in range(members.GetSize()):
+ item = members.GetTypeEnumMemberAtIndex(i)
+ if item.name == enum_member_name:
+ return item
+ return None
+
+ for m in self.target.modules:
+ for et in m.GetTypes(lldb.eTypeClassEnumeration):
+ et_member = find_enum_type_member(et, enum_member_name)
+ if et_member is not None:
+ return self.cast(et, et_member.unsigned)
+ return None
+
class LJBase(object):
# Ignore given parameters by LLDB.
def __init__(ljbase, debugger, unused):
@@ -800,6 +858,45 @@ def strx64(val):
return re.sub('L?$', '', hex(int(tou64(val))))
+class EnumBasedList(object):
+ def __init__(self, enum_name, max_enum_member, map_func=None,
+ *map_func_extra_args):
+ self.__enum_name = enum_name
+ self.__max_enum_member = max_enum_member
+ self.__map_func = map_func
+ self.__map_func_extra_args = map_func_extra_args
+ # Lazy initialization (see __get_items method) as the required
+ # information might be unavailable at this moment.
+ self.__items = None
+
+ def __iter__(self):
+ return iter(self.__get_items())
+
+ def __getitem__(self, key):
+ return self.__get_items()[key]
+
+ def __len__(self):
+ return len(self.__get_items())
+
+ def __get_items(self):
+ if self.__items is None:
+ max_enum_value = dbg.create_enum_value(
+ self.__enum_name, self.__max_enum_member
+ )
+ items = []
+ for i in range(dbg.cast('int', max_enum_value)):
+ item = str(dbg.cast(max_enum_value.type, dbg.eval(str(i))))
+ if self.__map_func:
+ item = self.__map_func(item, *self.__map_func_extra_args)
+ items.append(item)
+ self.__items = items
+ return self.__items
+
+
+def cut_prefix(s, prefix):
+ return s[len(prefix):] if s.startswith(prefix) else s
+
+
# Types and TValues.
@@ -877,10 +974,7 @@ def bc_d(ins):
return int(ins) >> 16
-BCMODE = [
- 'none', 'dst', 'base', 'var', 'rbase', 'uv',
- 'lit', 'lits', 'pri', 'num', 'str', 'tab', 'func', 'jump', 'cdata',
-]
+BCMODE = EnumBasedList('BCMode', 'BCM_max', cut_prefix, 'BCM')
lj_bc_mode_ = None
@@ -906,136 +1000,7 @@ def bcmode_cd(op):
return int((lj_bc_mode()[op] >> 7) & 15)
-# Unfortunately, there is no place in the VM except the generated
-# Lua table, where the bytecode names are stored. So duplicate
-# them here.
-BYTECODES = [
- # Comparison ops. ORDER OPR.
- 'ISLT',
- 'ISGE',
- 'ISLE',
- 'ISGT',
-
- 'ISEQV',
- 'ISNEV',
- 'ISEQS',
- 'ISNES',
- 'ISEQN',
- 'ISNEN',
- 'ISEQP',
- 'ISNEP',
-
- # Unary test and copy ops.
- 'ISTC',
- 'ISFC',
- 'IST',
- 'ISF',
- 'ISTYPE',
- 'ISNUM',
- 'MOV',
- 'NOT',
- 'UNM',
- 'LEN',
- 'ADDVN',
- 'SUBVN',
- 'MULVN',
- 'DIVVN',
- 'MODVN',
-
- # Binary ops. ORDER OPR.
- 'ADDNV',
- 'SUBNV',
- 'MULNV',
- 'DIVNV',
- 'MODNV',
-
- 'ADDVV',
- 'SUBVV',
- 'MULVV',
- 'DIVVV',
- 'MODVV',
-
- 'POW',
- 'CAT',
-
- # Constant ops.
- 'KSTR',
- 'KCDATA',
- 'KSHORT',
- 'KNUM',
- 'KPRI',
- 'KNIL',
-
- # Upvalue and function ops.
- 'UGET',
- 'USETV',
- 'USETS',
- 'USETN',
- 'USETP',
- 'UCLO',
- 'FNEW',
-
- # Table ops.
- 'TNEW',
- 'TDUP',
- 'GGET',
- 'GSET',
- 'TGETV',
- 'TGETS',
- 'TGETB',
- 'TGETR',
- 'TSETV',
- 'TSETS',
- 'TSETB',
- 'TSETM',
- 'TSETR',
-
- # Calls and vararg handling. T = tail call.
- 'CALLM',
- 'CALL',
- 'CALLMT',
- 'CALLT',
- 'ITERC',
- 'ITERN',
- 'VARG',
- 'ISNEXT',
-
- # Returns.
- 'RETM',
- 'RET',
- 'RET0',
- 'RET1',
-
- # Loops and branches. I/J = interp/JIT.
- # I/C/L = init/call/loop.
- 'FORI',
- 'JFORI',
-
- 'FORL',
- 'IFORL',
- 'JFORL',
-
- 'ITERL',
- 'IITERL',
- 'JITERL',
-
- 'LOOP',
- 'ILOOP',
- 'JLOOP',
-
- 'JMP',
-
- # Function headers. I/J = interp/JIT.
- # F/V/C = fixarg/vararg/C func.
- 'FUNCF',
- 'IFUNCF',
- 'JFUNCF',
- 'FUNCV',
- 'IFUNCV',
- 'JFUNCV',
- 'FUNCC',
- 'FUNCCW',
-]
+BYTECODES = EnumBasedList('BCOp', 'BC__MAX', cut_prefix, 'BC_')
def proto_bc(proto):
@@ -1190,42 +1155,16 @@ def J(g):
# Matched `MMDEF(_)`.
-MM_NAMES = [
- 'index',
- 'newindex',
- 'gc',
- 'mode',
- 'eq',
- 'len',
- 'lt',
- 'le',
- 'concat',
- 'call',
- 'add',
- 'sub',
- 'mul',
- 'div',
- 'mod',
- 'pow',
- 'unm',
- 'metatable',
- 'tostring',
- # TODO: depends on LJ_HASFFI, see `MMDEF_FFI(_)`.
- 'new',
- # TODO: depends on LJ_52 || LJ_HASFFI, see `MMDEF_PAIRS(_)`.
- 'pairs',
- 'ipairs',
-]
-
-
-GCROOT_MMNAME = 0
-GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
-GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
-GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
+MM_NAMES = EnumBasedList('MMS', 'MM__MAX', cut_prefix, 'MM_')
# Get the name of the index in the predefined arrays.
def idx_name(field_name):
+ GCROOT_MMNAME = 0
+ GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
+ GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
+ GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
+
# Don't use **{ to be compatible with Python 2.
gcroot = {}
gcroot.update({
@@ -1477,140 +1416,7 @@ def cdataptr(cd):
# JIT engine.
-IRS = [
- # Guarded assertions.
- 'LT',
- 'GE',
- 'LE',
- 'GT',
-
- 'ULT',
- 'UGE',
- 'ULE',
- 'UGT',
-
- 'EQ',
- 'NE',
-
- 'ABC',
- 'RETF',
-
- # Miscellaneous ops.
- 'NOP',
- 'BASE',
- 'PVAL',
- 'GCSTEP',
- 'HIOP',
- 'LOOP',
- 'USE',
- 'PHI',
- 'RENAME',
- 'PROF',
-
- # Constants.
- 'KPRI',
- 'KINT',
- 'KGC',
- 'KPTR',
- 'KKPTR',
- 'KNULL',
- 'KNUM',
- 'KINT64',
- 'KSLOT',
-
- # Bit ops.
- 'BNOT',
- 'BSWAP',
- 'BAND',
- 'BOR',
- 'BXOR',
- 'BSHL',
- 'BSHR',
- 'BSAR',
- 'BROL',
- 'BROR',
-
- # Arithmetic ops. ORDER ARITH
- 'ADD',
- 'SUB',
- 'MUL',
- 'DIV',
- 'MOD',
- 'POW',
- 'NEG',
-
- 'ABS',
- 'LDEXP',
- 'MIN',
- 'MAX',
- 'FPMATH',
-
- # Overflow-checking arithmetic ops.
- 'ADDOV',
- 'SUBOV',
- 'MULOV',
-
- # Memory ops. A = array, H = hash, U = upvalue, F = field,
- # S = stack.
-
- # Memory references.
- 'AREF',
- 'HREFK',
- 'HREF',
- 'NEWREF',
- 'UREFO',
- 'UREFC',
- 'FREF',
- 'STRREF',
- 'LREF',
-
- # Loads and Stores. These must be in the same order.
- 'ALOAD',
- 'HLOAD',
- 'ULOAD',
- 'FLOAD',
- 'XLOAD',
- 'SLOAD',
- 'VLOAD',
-
- 'ASTORE',
- 'HSTORE',
- 'USTORE',
- 'FSTORE',
- 'XSTORE',
-
- # Allocations.
- 'SNEW',
- 'XSNEW',
- 'TNEW',
- 'TDUP',
- 'CNEW',
- 'CNEWI',
-
- # Buffer operations.
- 'BUFHDR',
- 'BUFPUT',
- 'BUFSTR',
-
- # Barriers.
- 'TBAR',
- 'OBAR',
- 'XBAR',
-
- # Type conversions.
- 'CONV',
- 'TOBIT',
- 'TOSTR',
- 'STRTO',
-
- # Calls.
- 'CALLN',
- 'CALLA',
- 'CALLL',
- 'CALLS',
- 'CALLXS',
- 'CARG',
-]
+IRS = EnumBasedList('IROp', 'IR__MAX', cut_prefix, 'IR_')
# Mode bits: Commutative, {Normal/Ref, Alloc, Load, Store},
@@ -1662,70 +1468,21 @@ def ir_mode(op):
return mode
-IRTYPES = [
- 'nil',
- 'fal',
- 'tru',
- 'lud',
- 'str',
- 'p32',
- 'thr',
- 'pro',
- 'fun',
- 'p64',
- 'cdt',
- 'tab',
- 'udt',
- 'flt',
- 'num',
- 'i8 ',
- 'u8 ',
- 'i16',
- 'u16',
- 'int',
- 'u32',
- 'i64',
- 'u64',
- 'sfp',
-]
+IRTYPES = EnumBasedList('IRType', 'IRT__MAX', lambda x: {
+ 'IRT_LIGHTUD': 'lud',
+ 'IRT_CDATA': 'cdt',
+ 'IRT_UDATA': 'udt',
+ 'IRT_FLOAT': 'flt',
+ 'IRT_SOFTFP': 'sfp',
+ }.get(x, cut_prefix(x, 'IRT_')[:3].ljust(3).lower()))
-IRT_NUM = 14
-assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
-
-
-IRFIELDS = [
- 'str.len',
- 'func.env',
- 'func.pc',
- 'func.ffid',
- 'thread.env',
- 'tab.meta',
- 'tab.array',
- 'tab.node',
- 'tab.asize',
- 'tab.hmask',
- 'tab.nomm',
- 'udata.meta',
- 'udata.udtype',
- 'udata.file',
- 'cdata.ctypeid',
- 'cdata.ptr',
- 'cdata.int',
- 'cdata.int64',
- 'cdata.int64_4',
-]
+IRFIELDS = EnumBasedList('IRFieldID', 'IRFL__MAX', lambda x:
+ cut_prefix(x, 'IRFL_').lower().replace('_', '.', 1))
-IRFPMS = [
- 'floor',
- 'ceil',
- 'trunc',
- 'sqrt',
- 'log',
- 'log2',
- 'other'
-]
+IRFPMS = EnumBasedList('IRFPMathOp', 'IRFPM__MAX',
+ lambda x: cut_prefix(x, 'IRFPM_').lower())
# Don't use *[ to be compatible with Python 2.
@@ -1754,112 +1511,7 @@ REGISTERS = {
}
-IR_CALLS = [
- 'lj_str_cmp',
- 'lj_str_find',
- 'lj_str_new',
- 'lj_strscan_num',
- 'lj_strfmt_int',
- 'lj_strfmt_num',
- 'lj_strfmt_char',
- 'lj_strfmt_putint',
- 'lj_strfmt_putnum',
- 'lj_strfmt_putquoted',
- 'lj_strfmt_putfxint',
- 'lj_strfmt_putfnum_int',
- 'lj_strfmt_putfnum_uint',
- 'lj_strfmt_putfnum',
- 'lj_strfmt_putfstr',
- 'lj_strfmt_putfchar',
- 'lj_buf_putmem',
- 'lj_buf_putstr',
- 'lj_buf_putchar',
- 'lj_buf_putstr_reverse',
- 'lj_buf_putstr_lower',
- 'lj_buf_putstr_upper',
- 'lj_buf_putstr_rep',
- 'lj_buf_puttab',
- 'lj_buf_tostr',
- 'lj_tab_new_ah',
- 'lj_tab_new1',
- 'lj_tab_dup',
- 'lj_tab_clear',
- 'lj_tab_newkey',
- 'lj_tab_len',
- 'lj_gc_step_jit',
- 'lj_gc_barrieruv',
- 'lj_mem_newgco',
- 'lj_math_random_step',
- 'lj_vm_modi',
- 'log10',
- 'exp',
- 'sin',
- 'cos',
- 'tan',
- 'asin',
- 'acos',
- 'atan',
- 'sinh',
- 'cosh',
- 'tanh',
- 'fputc',
- 'fwrite',
- 'fflush',
- 'lj_vm_floor',
- 'lj_vm_ceil',
- 'lj_vm_trunc',
- 'sqrt',
- 'log',
- 'lj_vm_log2',
- 'pow',
- 'atan2',
- 'ldexp',
- 'lj_vm_tobit',
- 'softfp_add',
- 'softfp_sub',
- 'softfp_mul',
- 'softfp_div',
- 'softfp_cmp',
- 'softfp_i2d',
- 'softfp_d2i',
- 'lj_vm_sfmin',
- 'lj_vm_sfmax',
- 'lj_vm_tointg',
- 'softfp_ui2d',
- 'softfp_f2d',
- 'softfp_d2ui',
- 'softfp_d2f',
- 'softfp_i2f',
- 'softfp_ui2f',
- 'softfp_f2i',
- 'softfp_f2ui',
- 'fp64_l2d',
- 'fp64_ul2d',
- 'fp64_l2f',
- 'fp64_ul2f',
- 'fp64_d2l',
- 'fp64_d2ul',
- 'fp64_f2l',
- 'fp64_f2ul',
- 'lj_carith_divi64',
- 'lj_carith_divu64',
- 'lj_carith_modi64',
- 'lj_carith_modu64',
- 'lj_carith_powi64',
- 'lj_carith_powu64',
- 'lj_cdata_newv',
- 'lj_cdata_setfin',
- 'strlen',
- 'memcpy',
- 'memset',
- 'lj_vm_errno',
- 'lj_carith_mul64',
- 'lj_carith_shl64',
- 'lj_carith_shr64',
- 'lj_carith_sar64',
- 'lj_carith_rol64',
- 'lj_carith_ror64',
-]
+IR_CALLS = EnumBasedList('IRCallID', 'IRCALL__MAX', cut_prefix, 'IRCALL_')
def regname(reg_number):
@@ -1995,6 +1647,8 @@ def irt_isguard(t):
def irt_toitype(irt):
+ IRT_NUM = 14
+ assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
t = irt_type(irt)
if LJ_DUALNUM and t > IRT_NUM:
return LJ_T['NUMX']
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [Tarantool-patches] [PATCH luajit] dbg: avoid hardcoded enums (get them from target)
2026-09-09 8:49 [Tarantool-patches] [PATCH luajit] dbg: avoid hardcoded enums (get them from target) Mikhail Elhimov via Tarantool-patches
@ 2026-09-17 9:52 ` Sergey Bronnikov via Tarantool-patches
2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
0 siblings, 2 replies; 4+ messages in thread
From: Sergey Bronnikov via Tarantool-patches @ 2026-09-17 9:52 UTC (permalink / raw)
To: Mikhail Elhimov, Sergey Kaplun, Evgeniy Temirgaleev; +Cc: tarantool-patches
[-- Attachment #1: Type: text/plain, Size: 23808 bytes --]
Hi, Mikhail,
thanks for the patch! Please see my comments.
Sergey
On 9/9/26 11:49, Mikhail Elhimov wrote:
> Besides reducing lines of code this way the extension become compatible
> with various version of luajit because different version might use
s/luajit/LuaJIT/
> different set of enum members (newer version might introduce additional
> BC/IR/etc.)
>
> Prior to this patch string was used as a debugger-agnostic way to
> specify type, but this way might not work in case of enum because
> debugging information might be optimized out if no variable of such enum
> type is declared and its members are used only as a predefined
> constants. The type of enum is needed to be able to get human readable
> value (member name instead of number). The alternative way to get enum
> type is get it from the value, i.e. somehow obtain the value that is of
> enum type and then get its type object.
>
> To do that, separate method to create enum value was introduced in
> Debugger API, LLDB value class was monkey-patched to get value type in
> the same way as GDB value class does and 'cast' method was adjusted to
> accept also type object, not only string.
>
> Other adjustments:
> - [lldb] 'eval' adjusted to return object of the same type as 'cast'
> method (this improves consistency)
> - [gdb] in 'eval' method dropped check of the value returned by
> gdb.parse_and_eval() as it would fail also for expression like '0'
> (looks like a kind of legacy code that is not needed now).
> - [gdb/lldb] renamed eval argument to reflect its meaning (it is
> an expression rather than a command).
>
> Closes tarantool/tarantool#13094
> ---
> This patch is to be applied after the 'fix mapping of FPMATHOP' patch.
>
> Branch:https://github.com/tarantool/luajit/tree/elhimov/gh-13094-dbg-use-enums-from-inferior
> Related issue:https://github.com/tarantool/tarantool/issues/13094
>
> src/luajit_dbg.py | 628 +++++++++++-----------------------------------
> 1 file changed, 141 insertions(+), 487 deletions(-)
The existing test suite provides good coverage for `BCMODE`,
`BYTECODES`, `IRS`, `IRTYPES`, `IRFIELDS`, and `GCROOT_*`,
but the new code path handling "optimized enum type names" (where the
type is looked up via a member)
is not specifically tested. It is worth verifying that an `IRFieldID` in
the test binary actually triggers
the fallback mechanism; otherwise, the LLDB fallback remains untested.
>
> diff --git a/src/luajit_dbg.py b/src/luajit_dbg.py
> index 76001b7d..1ac1d275 100644
> --- a/src/luajit_dbg.py
> +++ b/src/luajit_dbg.py
> @@ -110,9 +110,21 @@ class Debugger(object):
> self.write('{} command initialized\n'.format(name))
> self.write('LuaJIT debug extension is successfully loaded\n')
>
> + def cast(self, tp, val):
> + '''Cast the value to the given type (it is either C type string
> + or the debugger-specific type object).'''
> + if isinstance(tp, str):
> + tp = self._dbgtype(tp)
> + return self._cast(tp, val)
> +
> + @abc.abstractmethod
> + def _cast(self, tp, val):
> + '''Cast the value to the debugger-specific type.'''
> + pass
> +
> @abc.abstractmethod
> - def cast(self, typestr, val):
> - '''Cast the value to the required C type.'''
> + def _dbgtype(self, typestr):
> + '''Convert C type string into debugger-specific type object.'''
> pass
>
> @abc.abstractmethod
> @@ -146,8 +158,9 @@ class Debugger(object):
> pass
>
> @abc.abstractmethod
> - def eval(self, command):
> - '''Parse and evaluate the given debugger command.'''
> + def eval(self, expr):
> + '''Parse and evaluate the given debugger expression.
> + Return debugger-specific value.'''
> pass
>
> @abc.abstractmethod
> @@ -187,6 +200,13 @@ class Debugger(object):
> '''Register the command with the corresponding name.'''
> pass
>
> + @abc.abstractmethod
> + def create_enum_value(self, enum_name, enum_member_name):
> + '''Return debugger-specific value object that represents
> + the given enum member.
> + '''
> + pass
> +
> @abc.abstractproperty
> def LJBase(self):
> '''Base command class.
> @@ -211,8 +231,9 @@ class _GDBDebugger(Debugger):
> super(_GDBDebugger, self).__init__()
> self.CONNECTED = False
>
> - def cast(self, typestr, val):
> - return gdb.Value(val).cast(self._dbgtype(typestr))
> + def _cast(self, tp, val):
> + assert isinstance(tp, gdb.Type)
> + return gdb.Value(val).cast(tp)
>
> def sizeof(self, typestr):
> return self._dbgtype(typestr).sizeof
> @@ -268,14 +289,11 @@ class _GDBDebugger(Debugger):
> else:
> return None
>
> - def eval(self, command):
> - if not command:
> + def eval(self, expr):
> + if not expr:
> return None
>
> - ret = gdb.parse_and_eval(command)
> - if not ret:
> - raise gdb.GdbError('table argument empty')
> - return ret
> + return gdb.parse_and_eval(expr)
>
> def detect_arch(self):
> if hasattr(self, 'arch'):
> @@ -340,6 +358,9 @@ class _GDBDebugger(Debugger):
> def register_command(self, command, name):
> command(name)
>
> + def create_enum_value(self, enum_name, enum_member_name):
> + return self.eval(enum_member_name)
For GDB, `create_enum_value` returns `self.eval(enum_member_name)`, and
`EnumBasedList` uses `.type` as the enum type.
If, in a certain version of GDB, `parse_and_eval('BC__MAX').type`
evaluates to `int`, `str()` will return numbers,
and the list will be populated with "0", "1", etc., without an explicit
error.
It is worth adding an assertion to verify that the type is an enum
(`gdb.TYPE_CODE_ENUM`).
> +
> class LJBase(gdb and gdb.Command or object):
> def __init__(ljbase, name):
> # XXX Fragile: Though the command initialization looks
> @@ -378,6 +399,10 @@ class _LLDBDebugger(Debugger):
> lldb.eBasicTypeInt128
> ]
>
> + def _lldb_tp_isenum(self, tp):
> + return tp.GetCanonicalType().GetTypeClass() == \
> + lldb.eTypeClassEnumeration
> +
> def _lldb_value_from_raw(self, raw_value, size, tp):
> isfp = self._lldb_tp_isfp(tp)
> if isfp:
> @@ -491,8 +516,7 @@ class _LLDBDebugger(Debugger):
> # Instead of default GetSummary.
> if not lldbval.sbvalue.TypeIsPointerType():
> tp = lldbval.sbvalue.GetType()
> - is_float = self._lldb_tp_isfp(tp)
> - if is_float:
> + if self._lldb_tp_isfp(tp) or self._lldb_tp_isenum(tp):
> return lldbval.sbvalue.GetValue()
> else:
> return str(int(lldbval))
> @@ -530,6 +554,9 @@ class _LLDBDebugger(Debugger):
> else:
> return int(lldbval) - int(other)
>
> + def lldb_gettype(lldbval):
> + return lldbval.sbvalue.type
> +
> super(_LLDBDebugger, self).__init__()
> self.target = lldb.debugger.GetSelectedTarget()
> # Monkey-patch the lldb.value class.
> @@ -545,6 +572,7 @@ class _LLDBDebugger(Debugger):
> lldb.value.__ror__ = lldb__or__ # Same semantics.
> lldb.value.__str__ = lldb__str__
> lldb.value.__sub__ = lldb__sub__
> + lldb.value.type = property(lldb_gettype)
>
> def lldb_major_version():
> version_string = lldb.SBDebugger.GetVersionString()
> @@ -568,11 +596,11 @@ class _LLDBDebugger(Debugger):
> self.dbgtype_cache[typestr] = dbgtype
> return dbgtype
>
> - def cast(self, typestr, val):
> + def _cast(self, tp, val):
> + assert isinstance(tp, lldb.SBType)
> if isinstance(val, lldb.value):
> val = val.sbvalue
> elif type(val) is int:
> - tp = self._dbgtype(typestr)
> return self._lldb_value_from_raw(val, tp.GetByteSize(), tp)
> elif not isinstance(val, lldb.SBValue):
> raise Exception(
> @@ -582,7 +610,6 @@ class _LLDBDebugger(Debugger):
> # XXX: Simply SBValue.Cast() works incorrectly since it
> # may take the 8 bytes of memory instead of 4, before the
> # cast. Construct the value on the fly.
> - tp = self._dbgtype(typestr)
> if self._lldb_tp_isfp(tp):
> rawval = float(val.GetValue())
> elif self._lldb_tp_issigned(tp):
> @@ -670,15 +697,15 @@ class _LLDBDebugger(Debugger):
> else:
> return None
>
> - def eval(self, command):
> - if not command:
> + def eval(self, expr):
> + if not expr:
> return None
>
> process = self.target.GetProcess()
> thread = process.GetSelectedThread()
> frame = thread.GetSelectedFrame()
> - ret = frame.EvaluateExpression(command)
> - return ret
> + ret = frame.EvaluateExpression(expr)
> + return lldb.value(ret)
>
> def detect_arch(self):
> if hasattr(self, 'arch'):
> @@ -724,6 +751,37 @@ class _LLDBDebugger(Debugger):
> )
> )
>
> + def create_enum_value(self, enum_name, enum_member_name):
> + val = self.eval(enum_name + "::" + enum_member_name)
> + if val.sbvalue.IsValid() and val.sbvalue.error.Success():
> + return val
> +
> + # LLDB uses enum name in expression above but debugging information
> + # about enum name migth be optimized out if no variable of the given
s/migth/might/
> + # enum type is declared and its members are only used as the predefined
> + # constants (like IRFieldID).
> +
> + # In this case the above method doesn't work so trying to discover
> + # enum type by the given enum member.
> +
> + def find_enum_type_member(enum_type, enum_member_name):
> + # SBTypeEnumMemberList supports members iteration and [] access
> + # (both by index and by member name) only starting from lldb-12
> + # so this implementation is used to handle earlier versions.
> + members = enum_type.GetEnumMembers()
> + for i in range(members.GetSize()):
> + item = members.GetTypeEnumMemberAtIndex(i)
> + if item.name == enum_member_name:
> + return item
> + return None
> +
> + for m in self.target.modules:
> + for et in m.GetTypes(lldb.eTypeClassEnumeration):
The construct `for et in m.GetTypes(lldb.eTypeClassEnumeration)` relies
on the `__iter__` method of `SBTypeList`.
There is a comment above noting that `SBTypeEnumMemberList` could not be
iterated over
prior to lldb 12 and no such guarantee exists for `SBTypeList`. For the
sake of consistency and
compatibility, it is better to use: list = m.GetTypes(...) followed by
for i in range(list.GetSize()):
et = list.GetTypeAtIndex(i)
> + et_member = find_enum_type_member(et, enum_member_name)
> + if et_member is not None:
> + return self.cast(et, et_member.unsigned)
> + return None
> +
> class LJBase(object):
> # Ignore given parameters by LLDB.
> def __init__(ljbase, debugger, unused):
> @@ -800,6 +858,45 @@ def strx64(val):
> return re.sub('L?$', '', hex(int(tou64(val))))
>
>
> +class EnumBasedList(object):
> + def __init__(self, enum_name, max_enum_member, map_func=None,
> + *map_func_extra_args):
> + self.__enum_name = enum_name
> + self.__max_enum_member = max_enum_member
> + self.__map_func = map_func
> + self.__map_func_extra_args = map_func_extra_args
> + # Lazy initialization (see __get_items method) as the required
> + # information might be unavailable at this moment.
> + self.__items = None
> +
> + def __iter__(self):
> + return iter(self.__get_items())
> +
> + def __getitem__(self, key):
> + return self.__get_items()[key]
> +
> + def __len__(self):
> + return len(self.__get_items())
> +
> + def __get_items(self):
> + if self.__items is None:
> + max_enum_value = dbg.create_enum_value(
> + self.__enum_name, self.__max_enum_member
> + )
> + items = []
> + for i in range(dbg.cast('int', max_enum_value)):
> + item = str(dbg.cast(max_enum_value.type, dbg.eval(str(i))))
> + if self.__map_func:
> + item = self.__map_func(item, *self.__map_func_extra_args)
> + items.append(item)
> + self.__items = items
> + return self.__items
> +
> +
> +def cut_prefix(s, prefix):
> + return s[len(prefix):] if s.startswith(prefix) else s
> +
> +
> # Types and TValues.
>
>
> @@ -877,10 +974,7 @@ def bc_d(ins):
> return int(ins) >> 16
>
>
> -BCMODE = [
> - 'none', 'dst', 'base', 'var', 'rbase', 'uv',
> - 'lit', 'lits', 'pri', 'num', 'str', 'tab', 'func', 'jump', 'cdata',
> -]
> +BCMODE = EnumBasedList('BCMode', 'BCM_max', cut_prefix, 'BCM')
>
>
> lj_bc_mode_ = None
> @@ -906,136 +1000,7 @@ def bcmode_cd(op):
> return int((lj_bc_mode()[op] >> 7) & 15)
>
>
> -# Unfortunately, there is no place in the VM except the generated
> -# Lua table, where the bytecode names are stored. So duplicate
> -# them here.
> -BYTECODES = [
> - # Comparison ops. ORDER OPR.
> - 'ISLT',
> - 'ISGE',
> - 'ISLE',
> - 'ISGT',
> -
> - 'ISEQV',
> - 'ISNEV',
> - 'ISEQS',
> - 'ISNES',
> - 'ISEQN',
> - 'ISNEN',
> - 'ISEQP',
> - 'ISNEP',
> -
> - # Unary test and copy ops.
> - 'ISTC',
> - 'ISFC',
> - 'IST',
> - 'ISF',
> - 'ISTYPE',
> - 'ISNUM',
> - 'MOV',
> - 'NOT',
> - 'UNM',
> - 'LEN',
> - 'ADDVN',
> - 'SUBVN',
> - 'MULVN',
> - 'DIVVN',
> - 'MODVN',
> -
> - # Binary ops. ORDER OPR.
> - 'ADDNV',
> - 'SUBNV',
> - 'MULNV',
> - 'DIVNV',
> - 'MODNV',
> -
> - 'ADDVV',
> - 'SUBVV',
> - 'MULVV',
> - 'DIVVV',
> - 'MODVV',
> -
> - 'POW',
> - 'CAT',
> -
> - # Constant ops.
> - 'KSTR',
> - 'KCDATA',
> - 'KSHORT',
> - 'KNUM',
> - 'KPRI',
> - 'KNIL',
> -
> - # Upvalue and function ops.
> - 'UGET',
> - 'USETV',
> - 'USETS',
> - 'USETN',
> - 'USETP',
> - 'UCLO',
> - 'FNEW',
> -
> - # Table ops.
> - 'TNEW',
> - 'TDUP',
> - 'GGET',
> - 'GSET',
> - 'TGETV',
> - 'TGETS',
> - 'TGETB',
> - 'TGETR',
> - 'TSETV',
> - 'TSETS',
> - 'TSETB',
> - 'TSETM',
> - 'TSETR',
> -
> - # Calls and vararg handling. T = tail call.
> - 'CALLM',
> - 'CALL',
> - 'CALLMT',
> - 'CALLT',
> - 'ITERC',
> - 'ITERN',
> - 'VARG',
> - 'ISNEXT',
> -
> - # Returns.
> - 'RETM',
> - 'RET',
> - 'RET0',
> - 'RET1',
> -
> - # Loops and branches. I/J = interp/JIT.
> - # I/C/L = init/call/loop.
> - 'FORI',
> - 'JFORI',
> -
> - 'FORL',
> - 'IFORL',
> - 'JFORL',
> -
> - 'ITERL',
> - 'IITERL',
> - 'JITERL',
> -
> - 'LOOP',
> - 'ILOOP',
> - 'JLOOP',
> -
> - 'JMP',
> -
> - # Function headers. I/J = interp/JIT.
> - # F/V/C = fixarg/vararg/C func.
> - 'FUNCF',
> - 'IFUNCF',
> - 'JFUNCF',
> - 'FUNCV',
> - 'IFUNCV',
> - 'JFUNCV',
> - 'FUNCC',
> - 'FUNCCW',
> -]
> +BYTECODES = EnumBasedList('BCOp', 'BC__MAX', cut_prefix, 'BC_')
>
>
> def proto_bc(proto):
> @@ -1190,42 +1155,16 @@ def J(g):
>
>
> # Matched `MMDEF(_)`.
> -MM_NAMES = [
> - 'index',
> - 'newindex',
> - 'gc',
> - 'mode',
> - 'eq',
> - 'len',
> - 'lt',
> - 'le',
> - 'concat',
> - 'call',
> - 'add',
> - 'sub',
> - 'mul',
> - 'div',
> - 'mod',
> - 'pow',
> - 'unm',
> - 'metatable',
> - 'tostring',
> - # TODO: depends on LJ_HASFFI, see `MMDEF_FFI(_)`.
> - 'new',
> - # TODO: depends on LJ_52 || LJ_HASFFI, see `MMDEF_PAIRS(_)`.
> - 'pairs',
> - 'ipairs',
> -]
> -
> -
> -GCROOT_MMNAME = 0
> -GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
> -GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
> -GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
> +MM_NAMES = EnumBasedList('MMS', 'MM__MAX', cut_prefix, 'MM_')
>
>
> # Get the name of the index in the predefined arrays.
> def idx_name(field_name):
> + GCROOT_MMNAME = 0
> + GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
> + GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
> + GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
> +
> # Don't use **{ to be compatible with Python 2.
> gcroot = {}
> gcroot.update({
> @@ -1477,140 +1416,7 @@ def cdataptr(cd):
> # JIT engine.
>
>
> -IRS = [
> - # Guarded assertions.
> - 'LT',
> - 'GE',
> - 'LE',
> - 'GT',
> -
> - 'ULT',
> - 'UGE',
> - 'ULE',
> - 'UGT',
> -
> - 'EQ',
> - 'NE',
> -
> - 'ABC',
> - 'RETF',
> -
> - # Miscellaneous ops.
> - 'NOP',
> - 'BASE',
> - 'PVAL',
> - 'GCSTEP',
> - 'HIOP',
> - 'LOOP',
> - 'USE',
> - 'PHI',
> - 'RENAME',
> - 'PROF',
> -
> - # Constants.
> - 'KPRI',
> - 'KINT',
> - 'KGC',
> - 'KPTR',
> - 'KKPTR',
> - 'KNULL',
> - 'KNUM',
> - 'KINT64',
> - 'KSLOT',
> -
> - # Bit ops.
> - 'BNOT',
> - 'BSWAP',
> - 'BAND',
> - 'BOR',
> - 'BXOR',
> - 'BSHL',
> - 'BSHR',
> - 'BSAR',
> - 'BROL',
> - 'BROR',
> -
> - # Arithmetic ops. ORDER ARITH
> - 'ADD',
> - 'SUB',
> - 'MUL',
> - 'DIV',
> - 'MOD',
> - 'POW',
> - 'NEG',
> -
> - 'ABS',
> - 'LDEXP',
> - 'MIN',
> - 'MAX',
> - 'FPMATH',
> -
> - # Overflow-checking arithmetic ops.
> - 'ADDOV',
> - 'SUBOV',
> - 'MULOV',
> -
> - # Memory ops. A = array, H = hash, U = upvalue, F = field,
> - # S = stack.
> -
> - # Memory references.
> - 'AREF',
> - 'HREFK',
> - 'HREF',
> - 'NEWREF',
> - 'UREFO',
> - 'UREFC',
> - 'FREF',
> - 'STRREF',
> - 'LREF',
> -
> - # Loads and Stores. These must be in the same order.
> - 'ALOAD',
> - 'HLOAD',
> - 'ULOAD',
> - 'FLOAD',
> - 'XLOAD',
> - 'SLOAD',
> - 'VLOAD',
> -
> - 'ASTORE',
> - 'HSTORE',
> - 'USTORE',
> - 'FSTORE',
> - 'XSTORE',
> -
> - # Allocations.
> - 'SNEW',
> - 'XSNEW',
> - 'TNEW',
> - 'TDUP',
> - 'CNEW',
> - 'CNEWI',
> -
> - # Buffer operations.
> - 'BUFHDR',
> - 'BUFPUT',
> - 'BUFSTR',
> -
> - # Barriers.
> - 'TBAR',
> - 'OBAR',
> - 'XBAR',
> -
> - # Type conversions.
> - 'CONV',
> - 'TOBIT',
> - 'TOSTR',
> - 'STRTO',
> -
> - # Calls.
> - 'CALLN',
> - 'CALLA',
> - 'CALLL',
> - 'CALLS',
> - 'CALLXS',
> - 'CARG',
> -]
> +IRS = EnumBasedList('IROp', 'IR__MAX', cut_prefix, 'IR_')
>
>
> # Mode bits: Commutative, {Normal/Ref, Alloc, Load, Store},
> @@ -1662,70 +1468,21 @@ def ir_mode(op):
> return mode
>
>
> -IRTYPES = [
> - 'nil',
> - 'fal',
> - 'tru',
> - 'lud',
> - 'str',
> - 'p32',
> - 'thr',
> - 'pro',
> - 'fun',
> - 'p64',
> - 'cdt',
> - 'tab',
> - 'udt',
> - 'flt',
> - 'num',
> - 'i8 ',
> - 'u8 ',
> - 'i16',
> - 'u16',
> - 'int',
> - 'u32',
> - 'i64',
> - 'u64',
> - 'sfp',
> -]
> +IRTYPES = EnumBasedList('IRType', 'IRT__MAX', lambda x: {
> + 'IRT_LIGHTUD': 'lud',
> + 'IRT_CDATA': 'cdt',
> + 'IRT_UDATA': 'udt',
> + 'IRT_FLOAT': 'flt',
> + 'IRT_SOFTFP': 'sfp',
> + }.get(x, cut_prefix(x, 'IRT_')[:3].ljust(3).lower()))
>
>
> -IRT_NUM = 14
> -assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
> -
> -
> -IRFIELDS = [
> - 'str.len',
> - 'func.env',
> - 'func.pc',
> - 'func.ffid',
> - 'thread.env',
> - 'tab.meta',
> - 'tab.array',
> - 'tab.node',
> - 'tab.asize',
> - 'tab.hmask',
> - 'tab.nomm',
> - 'udata.meta',
> - 'udata.udtype',
> - 'udata.file',
> - 'cdata.ctypeid',
> - 'cdata.ptr',
> - 'cdata.int',
> - 'cdata.int64',
> - 'cdata.int64_4',
> -]
> +IRFIELDS = EnumBasedList('IRFieldID', 'IRFL__MAX', lambda x:
> + cut_prefix(x, 'IRFL_').lower().replace('_', '.', 1))
>
>
> -IRFPMS = [
> - 'floor',
> - 'ceil',
> - 'trunc',
> - 'sqrt',
> - 'log',
> - 'log2',
> - 'other'
> -]
> +IRFPMS = EnumBasedList('IRFPMathOp', 'IRFPM__MAX',
> + lambda x: cut_prefix(x, 'IRFPM_').lower())
>
>
> # Don't use *[ to be compatible with Python 2.
> @@ -1754,112 +1511,7 @@ REGISTERS = {
> }
>
>
> -IR_CALLS = [
> - 'lj_str_cmp',
> - 'lj_str_find',
> - 'lj_str_new',
> - 'lj_strscan_num',
> - 'lj_strfmt_int',
> - 'lj_strfmt_num',
> - 'lj_strfmt_char',
> - 'lj_strfmt_putint',
> - 'lj_strfmt_putnum',
> - 'lj_strfmt_putquoted',
> - 'lj_strfmt_putfxint',
> - 'lj_strfmt_putfnum_int',
> - 'lj_strfmt_putfnum_uint',
> - 'lj_strfmt_putfnum',
> - 'lj_strfmt_putfstr',
> - 'lj_strfmt_putfchar',
> - 'lj_buf_putmem',
> - 'lj_buf_putstr',
> - 'lj_buf_putchar',
> - 'lj_buf_putstr_reverse',
> - 'lj_buf_putstr_lower',
> - 'lj_buf_putstr_upper',
> - 'lj_buf_putstr_rep',
> - 'lj_buf_puttab',
> - 'lj_buf_tostr',
> - 'lj_tab_new_ah',
> - 'lj_tab_new1',
> - 'lj_tab_dup',
> - 'lj_tab_clear',
> - 'lj_tab_newkey',
> - 'lj_tab_len',
> - 'lj_gc_step_jit',
> - 'lj_gc_barrieruv',
> - 'lj_mem_newgco',
> - 'lj_math_random_step',
> - 'lj_vm_modi',
> - 'log10',
> - 'exp',
> - 'sin',
> - 'cos',
> - 'tan',
> - 'asin',
> - 'acos',
> - 'atan',
> - 'sinh',
> - 'cosh',
> - 'tanh',
> - 'fputc',
> - 'fwrite',
> - 'fflush',
> - 'lj_vm_floor',
> - 'lj_vm_ceil',
> - 'lj_vm_trunc',
> - 'sqrt',
> - 'log',
> - 'lj_vm_log2',
> - 'pow',
> - 'atan2',
> - 'ldexp',
> - 'lj_vm_tobit',
> - 'softfp_add',
> - 'softfp_sub',
> - 'softfp_mul',
> - 'softfp_div',
> - 'softfp_cmp',
> - 'softfp_i2d',
> - 'softfp_d2i',
> - 'lj_vm_sfmin',
> - 'lj_vm_sfmax',
> - 'lj_vm_tointg',
> - 'softfp_ui2d',
> - 'softfp_f2d',
> - 'softfp_d2ui',
> - 'softfp_d2f',
> - 'softfp_i2f',
> - 'softfp_ui2f',
> - 'softfp_f2i',
> - 'softfp_f2ui',
> - 'fp64_l2d',
> - 'fp64_ul2d',
> - 'fp64_l2f',
> - 'fp64_ul2f',
> - 'fp64_d2l',
> - 'fp64_d2ul',
> - 'fp64_f2l',
> - 'fp64_f2ul',
> - 'lj_carith_divi64',
> - 'lj_carith_divu64',
> - 'lj_carith_modi64',
> - 'lj_carith_modu64',
> - 'lj_carith_powi64',
> - 'lj_carith_powu64',
> - 'lj_cdata_newv',
> - 'lj_cdata_setfin',
> - 'strlen',
> - 'memcpy',
> - 'memset',
> - 'lj_vm_errno',
> - 'lj_carith_mul64',
> - 'lj_carith_shl64',
> - 'lj_carith_shr64',
> - 'lj_carith_sar64',
> - 'lj_carith_rol64',
> - 'lj_carith_ror64',
> -]
> +IR_CALLS = EnumBasedList('IRCallID', 'IRCALL__MAX', cut_prefix, 'IRCALL_')
>
>
> def regname(reg_number):
> @@ -1995,6 +1647,8 @@ def irt_isguard(t):
>
>
> def irt_toitype(irt):
> + IRT_NUM = 14
> + assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
> t = irt_type(irt)
> if LJ_DUALNUM and t > IRT_NUM:
> return LJ_T['NUMX']
[-- Attachment #2: Type: text/html, Size: 25450 bytes --]
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [Tarantool-patches] [PATCH luajit] dbg: avoid hardcoded enums (get them from target)
2026-09-17 9:52 ` Sergey Bronnikov via Tarantool-patches
@ 2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
1 sibling, 0 replies; 4+ messages in thread
From: Mikhail Elhimov via Tarantool-patches @ 2026-09-17 22:42 UTC (permalink / raw)
To: Sergey Bronnikov, Sergey Kaplun, Evgeniy Temirgaleev; +Cc: tarantool-patches
[-- Attachment #1: Type: text/plain, Size: 27176 bytes --]
Hi, Sergey!
Thanks for the review! See my comments below
On 17.09.2026 12:52, Sergey Bronnikov wrote:
>
> Hi, Mikhail,
>
> thanks for the patch! Please see my comments.
>
> Sergey
>
> On 9/9/26 11:49, Mikhail Elhimov wrote:
>> Besides reducing lines of code this way the extension become compatible
>> with various version of luajit because different version might use
> s/luajit/LuaJIT/
Fixed
>> different set of enum members (newer version might introduce additional
>> BC/IR/etc.)
>>
>> Prior to this patch string was used as a debugger-agnostic way to
>> specify type, but this way might not work in case of enum because
>> debugging information might be optimized out if no variable of such enum
>> type is declared and its members are used only as a predefined
>> constants. The type of enum is needed to be able to get human readable
>> value (member name instead of number). The alternative way to get enum
>> type is get it from the value, i.e. somehow obtain the value that is of
>> enum type and then get its type object.
>>
>> To do that, separate method to create enum value was introduced in
>> Debugger API, LLDB value class was monkey-patched to get value type in
>> the same way as GDB value class does and 'cast' method was adjusted to
>> accept also type object, not only string.
>>
>> Other adjustments:
>> - [lldb] 'eval' adjusted to return object of the same type as 'cast'
>> method (this improves consistency)
>> - [gdb] in 'eval' method dropped check of the value returned by
>> gdb.parse_and_eval() as it would fail also for expression like '0'
>> (looks like a kind of legacy code that is not needed now).
>> - [gdb/lldb] renamed eval argument to reflect its meaning (it is
>> an expression rather than a command).
>>
>> Closes tarantool/tarantool#13094
>> ---
>> This patch is to be applied after the 'fix mapping of FPMATHOP' patch.
>>
>> Branch:https://github.com/tarantool/luajit/tree/elhimov/gh-13094-dbg-use-enums-from-inferior
>> Related issue:https://github.com/tarantool/tarantool/issues/13094
>>
>> src/luajit_dbg.py | 628 +++++++++++-----------------------------------
>> 1 file changed, 141 insertions(+), 487 deletions(-)
>
> The existing test suite provides good coverage for `BCMODE`,
> `BYTECODES`, `IRS`, `IRTYPES`, `IRFIELDS`, and `GCROOT_*`,
>
> but the new code path handling "optimized enum type names" (where the
> type is looked up via a member)
>
> is not specifically tested. It is worth verifying that an `IRFieldID`
> in the test binary actually triggers
>
> the fallback mechanism; otherwise, the LLDB fallback remains untested.
>
At the moment the only way I found to verify that the fallback mechanism
was triggered is:
val = dbg.create_enum_value('IRFieldID', 'IRFL__MAX')
if val is not None and val.type.name == '(unnamed enum)':
# fallback triggered
But it does work only from the extension and I see no way to get access
to the enum_type from the tests level. So now we are only assuming that
the fallback mechanism triggers in case of `IRFL__MAX` of `IRFieldID` enum.
>> diff --git a/src/luajit_dbg.py b/src/luajit_dbg.py
>> index 76001b7d..1ac1d275 100644
>> --- a/src/luajit_dbg.py
>> +++ b/src/luajit_dbg.py
>> @@ -110,9 +110,21 @@ class Debugger(object):
>> self.write('{} command initialized\n'.format(name))
>> self.write('LuaJIT debug extension is successfully loaded\n')
>>
>> + def cast(self, tp, val):
>> + '''Cast the value to the given type (it is either C type string
>> + or the debugger-specific type object).'''
>> + if isinstance(tp, str):
>> + tp = self._dbgtype(tp)
>> + return self._cast(tp, val)
>> +
>> + @abc.abstractmethod
>> + def _cast(self, tp, val):
>> + '''Cast the value to the debugger-specific type.'''
>> + pass
>> +
>> @abc.abstractmethod
>> - def cast(self, typestr, val):
>> - '''Cast the value to the required C type.'''
>> + def _dbgtype(self, typestr):
>> + '''Convert C type string into debugger-specific type object.'''
>> pass
>>
>> @abc.abstractmethod
>> @@ -146,8 +158,9 @@ class Debugger(object):
>> pass
>>
>> @abc.abstractmethod
>> - def eval(self, command):
>> - '''Parse and evaluate the given debugger command.'''
>> + def eval(self, expr):
>> + '''Parse and evaluate the given debugger expression.
>> + Return debugger-specific value.'''
>> pass
>>
>> @abc.abstractmethod
>> @@ -187,6 +200,13 @@ class Debugger(object):
>> '''Register the command with the corresponding name.'''
>> pass
>>
>> + @abc.abstractmethod
>> + def create_enum_value(self, enum_name, enum_member_name):
>> + '''Return debugger-specific value object that represents
>> + the given enum member.
>> + '''
>> + pass
>> +
>> @abc.abstractproperty
>> def LJBase(self):
>> '''Base command class.
>> @@ -211,8 +231,9 @@ class _GDBDebugger(Debugger):
>> super(_GDBDebugger, self).__init__()
>> self.CONNECTED = False
>>
>> - def cast(self, typestr, val):
>> - return gdb.Value(val).cast(self._dbgtype(typestr))
>> + def _cast(self, tp, val):
>> + assert isinstance(tp, gdb.Type)
>> + return gdb.Value(val).cast(tp)
>>
>> def sizeof(self, typestr):
>> return self._dbgtype(typestr).sizeof
>> @@ -268,14 +289,11 @@ class _GDBDebugger(Debugger):
>> else:
>> return None
>>
>> - def eval(self, command):
>> - if not command:
>> + def eval(self, expr):
>> + if not expr:
>> return None
>>
>> - ret = gdb.parse_and_eval(command)
>> - if not ret:
>> - raise gdb.GdbError('table argument empty')
>> - return ret
>> + return gdb.parse_and_eval(expr)
>>
>> def detect_arch(self):
>> if hasattr(self, 'arch'):
>> @@ -340,6 +358,9 @@ class _GDBDebugger(Debugger):
>> def register_command(self, command, name):
>> command(name)
>>
>> + def create_enum_value(self, enum_name, enum_member_name):
>> + return self.eval(enum_member_name)
>
> For GDB, `create_enum_value` returns `self.eval(enum_member_name)`,
> and `EnumBasedList` uses `.type` as the enum type.
>
> If, in a certain version of GDB, `parse_and_eval('BC__MAX').type`
> evaluates to `int`, `str()` will return numbers,
>
> and the list will be populated with "0", "1", etc., without an
> explicit error.
>
> It is worth adding an assertion to verify that the type is an enum
> (`gdb.TYPE_CODE_ENUM`).
>
If such kind of expression is recognized as enum member it is evaluated
to the value of enum type. If not then it's just an invalid expression
and `gdb.parse_and_eval()` fails with exception. The mentioned scenario
with int instead of enum looks quite weird for me as it would mean that
gdb recognize enum value (which means that there is information about
corresponding enum type), but for some inexpicable reason "decides" to
cast it to another type. Why? Is that just a concern, or have you seen
that kind of behavior?
However I have add assertions to both gdb and lldb implementation, but
considering the above I would treat them as a verification that correct
arguments were passed into `create_enum_value()` rather than a
verification that such a spontaneous casting-to-int didn't happen.
>> +
>> class LJBase(gdb and gdb.Command or object):
>> def __init__(ljbase, name):
>> # XXX Fragile: Though the command initialization looks
>> @@ -378,6 +399,10 @@ class _LLDBDebugger(Debugger):
>> lldb.eBasicTypeInt128
>> ]
>>
>> + def _lldb_tp_isenum(self, tp):
>> + return tp.GetCanonicalType().GetTypeClass() == \
>> + lldb.eTypeClassEnumeration
>> +
>> def _lldb_value_from_raw(self, raw_value, size, tp):
>> isfp = self._lldb_tp_isfp(tp)
>> if isfp:
>> @@ -491,8 +516,7 @@ class _LLDBDebugger(Debugger):
>> # Instead of default GetSummary.
>> if not lldbval.sbvalue.TypeIsPointerType():
>> tp = lldbval.sbvalue.GetType()
>> - is_float = self._lldb_tp_isfp(tp)
>> - if is_float:
>> + if self._lldb_tp_isfp(tp) or self._lldb_tp_isenum(tp):
>> return lldbval.sbvalue.GetValue()
>> else:
>> return str(int(lldbval))
>> @@ -530,6 +554,9 @@ class _LLDBDebugger(Debugger):
>> else:
>> return int(lldbval) - int(other)
>>
>> + def lldb_gettype(lldbval):
>> + return lldbval.sbvalue.type
>> +
>> super(_LLDBDebugger, self).__init__()
>> self.target = lldb.debugger.GetSelectedTarget()
>> # Monkey-patch the lldb.value class.
>> @@ -545,6 +572,7 @@ class _LLDBDebugger(Debugger):
>> lldb.value.__ror__ = lldb__or__ # Same semantics.
>> lldb.value.__str__ = lldb__str__
>> lldb.value.__sub__ = lldb__sub__
>> + lldb.value.type = property(lldb_gettype)
>>
>> def lldb_major_version():
>> version_string = lldb.SBDebugger.GetVersionString()
>> @@ -568,11 +596,11 @@ class _LLDBDebugger(Debugger):
>> self.dbgtype_cache[typestr] = dbgtype
>> return dbgtype
>>
>> - def cast(self, typestr, val):
>> + def _cast(self, tp, val):
>> + assert isinstance(tp, lldb.SBType)
>> if isinstance(val, lldb.value):
>> val = val.sbvalue
>> elif type(val) is int:
>> - tp = self._dbgtype(typestr)
>> return self._lldb_value_from_raw(val, tp.GetByteSize(), tp)
>> elif not isinstance(val, lldb.SBValue):
>> raise Exception(
>> @@ -582,7 +610,6 @@ class _LLDBDebugger(Debugger):
>> # XXX: Simply SBValue.Cast() works incorrectly since it
>> # may take the 8 bytes of memory instead of 4, before the
>> # cast. Construct the value on the fly.
>> - tp = self._dbgtype(typestr)
>> if self._lldb_tp_isfp(tp):
>> rawval = float(val.GetValue())
>> elif self._lldb_tp_issigned(tp):
>> @@ -670,15 +697,15 @@ class _LLDBDebugger(Debugger):
>> else:
>> return None
>>
>> - def eval(self, command):
>> - if not command:
>> + def eval(self, expr):
>> + if not expr:
>> return None
>>
>> process = self.target.GetProcess()
>> thread = process.GetSelectedThread()
>> frame = thread.GetSelectedFrame()
>> - ret = frame.EvaluateExpression(command)
>> - return ret
>> + ret = frame.EvaluateExpression(expr)
>> + return lldb.value(ret)
>>
>> def detect_arch(self):
>> if hasattr(self, 'arch'):
>> @@ -724,6 +751,37 @@ class _LLDBDebugger(Debugger):
>> )
>> )
>>
>> + def create_enum_value(self, enum_name, enum_member_name):
>> + val = self.eval(enum_name + "::" + enum_member_name)
>> + if val.sbvalue.IsValid() and val.sbvalue.error.Success():
>> + return val
>> +
>> + # LLDB uses enum name in expression above but debugging information
>> + # about enum name migth be optimized out if no variable of the given
> s/migth/might/
Fixed
>> + # enum type is declared and its members are only used as the predefined
>> + # constants (like IRFieldID).
>> +
>> + # In this case the above method doesn't work so trying to discover
>> + # enum type by the given enum member.
>> +
>> + def find_enum_type_member(enum_type, enum_member_name):
>> + # SBTypeEnumMemberList supports members iteration and [] access
>> + # (both by index and by member name) only starting from lldb-12
>> + # so this implementation is used to handle earlier versions.
>> + members = enum_type.GetEnumMembers()
>> + for i in range(members.GetSize()):
>> + item = members.GetTypeEnumMemberAtIndex(i)
>> + if item.name == enum_member_name:
>> + return item
>> + return None
>> +
>> + for m in self.target.modules:
>> + for et in m.GetTypes(lldb.eTypeClassEnumeration):
>
> The construct `for et in m.GetTypes(lldb.eTypeClassEnumeration)`
> relies on the `__iter__` method of `SBTypeList`.
>
> There is a comment above noting that `SBTypeEnumMemberList` could not
> be iterated over
>
> prior to lldb 12 and no such guarantee exists for `SBTypeList`. For
> the sake of consistency and
>
> compatibility, it is better to use: list = m.GetTypes(...) followed by
>
> for i in range(list.GetSize()):
>
> et = list.GetTypeAtIndex(i)
>
I find this kind of iteration as the most convenient so I'm trying to
use it unless there is some reason against it. `SBTypeList` is just
another type that has no relation to `SBTypeEnumMemberList` and it is
not affected with the mentioned "problem". And I'm not sure I understand
correctly what does mean "no guarantee exists for `SBTypeList`" since we
are able to check it in llvm repo. I checked by tag llvmorg-4.0.0
(version `lldb` subdirectory first appeared in) and it looks like
`SBTypeList` had supported iteration initially. Thus it turns out that
the only reason left to change the iteration approach is to make it
similar to iteration over `SBTypeEnumMemberList`. Then what about
iteration over modules a line above? Should it be changed to look
similar as well? Something like this:
for i_module in range(self.target.GetNumModules()):
list = self.target.GetModuleAtIndex(i_module).GetTypes(lldb.eTypeClassEnumeration)
for i in range(list.GetSize()):
et = list.GetTypeAtIndex(i)
>> + et_member = find_enum_type_member(et, enum_member_name)
>> + if et_member is notNone:first ver
>> + return self.cast(et, et_member.unsigned)
>> + return None
>> +
>> class LJBase(object):
>> # Ignore given parameters by LLDB.
>> def __init__(ljbase, debugger, unused):
>> @@ -800,6 +858,45 @@ def strx64(val):
>> return re.sub('L?$', '', hex(int(tou64(val))))
>>
>>
>> +class EnumBasedList(object):
>> + def __init__(self, enum_name, max_enum_member, map_func=None,
>> + *map_func_extra_args):
>> + self.__enum_name = enum_name
>> + self.__max_enum_member = max_enum_member
>> + self.__map_func = map_func
>> + self.__map_func_extra_args = map_func_extra_args
>> + # Lazy initialization (see __get_items method) as the required
>> + # information might be unavailable at this moment.
>> + self.__items = None
>> +
>> + def __iter__(self):
>> + return iter(self.__get_items())
>> +
>> + def __getitem__(self, key):
>> + return self.__get_items()[key]
>> +
>> + def __len__(self):
>> + return len(self.__get_items())
>> +
>> + def __get_items(self):
>> + if self.__items is None:
>> + max_enum_value = dbg.create_enum_value(
>> + self.__enum_name, self.__max_enum_member
>> + )
>> + items = []
>> + for i in range(dbg.cast('int', max_enum_value)):
>> + item = str(dbg.cast(max_enum_value.type, dbg.eval(str(i))))
>> + if self.__map_func:
>> + item = self.__map_func(item, *self.__map_func_extra_args)
>> + items.append(item)
>> + self.__items = items
>> + return self.__items
>> +
>> +
>> +def cut_prefix(s, prefix):
>> + return s[len(prefix):] if s.startswith(prefix) else s
>> +
>> +
>> # Types and TValues.
>>
>>
>> @@ -877,10 +974,7 @@ def bc_d(ins):
>> return int(ins) >> 16
>>
>>
>> -BCMODE = [
>> - 'none', 'dst', 'base', 'var', 'rbase', 'uv',
>> - 'lit', 'lits', 'pri', 'num', 'str', 'tab', 'func', 'jump', 'cdata',
>> -]
>> +BCMODE = EnumBasedList('BCMode', 'BCM_max', cut_prefix, 'BCM')
>>
>>
>> lj_bc_mode_ = None
>> @@ -906,136 +1000,7 @@ def bcmode_cd(op):
>> return int((lj_bc_mode()[op] >> 7) & 15)
>>
>>
>> -# Unfortunately, there is no place in the VM except the generated
>> -# Lua table, where the bytecode names are stored. So duplicate
>> -# them here.
>> -BYTECODES = [
>> - # Comparison ops. ORDER OPR.
>> - 'ISLT',
>> - 'ISGE',
>> - 'ISLE',
>> - 'ISGT',
>> -
>> - 'ISEQV',
>> - 'ISNEV',
>> - 'ISEQS',
>> - 'ISNES',
>> - 'ISEQN',
>> - 'ISNEN',
>> - 'ISEQP',
>> - 'ISNEP',
>> -
>> - # Unary test and copy ops.
>> - 'ISTC',
>> - 'ISFC',
>> - 'IST',
>> - 'ISF',
>> - 'ISTYPE',
>> - 'ISNUM',
>> - 'MOV',
>> - 'NOT',
>> - 'UNM',
>> - 'LEN',
>> - 'ADDVN',
>> - 'SUBVN',
>> - 'MULVN',
>> - 'DIVVN',
>> - 'MODVN',
>> -
>> - # Binary ops. ORDER OPR.
>> - 'ADDNV',
>> - 'SUBNV',
>> - 'MULNV',
>> - 'DIVNV',
>> - 'MODNV',
>> -
>> - 'ADDVV',
>> - 'SUBVV',
>> - 'MULVV',
>> - 'DIVVV',
>> - 'MODVV',
>> -
>> - 'POW',
>> - 'CAT',
>> -
>> - # Constant ops.
>> - 'KSTR',
>> - 'KCDATA',
>> - 'KSHORT',
>> - 'KNUM',
>> - 'KPRI',
>> - 'KNIL',
>> -
>> - # Upvalue and function ops.
>> - 'UGET',
>> - 'USETV',
>> - 'USETS',
>> - 'USETN',
>> - 'USETP',
>> - 'UCLO',
>> - 'FNEW',
>> -
>> - # Table ops.
>> - 'TNEW',
>> - 'TDUP',
>> - 'GGET',
>> - 'GSET',
>> - 'TGETV',
>> - 'TGETS',
>> - 'TGETB',
>> - 'TGETR',
>> - 'TSETV',
>> - 'TSETS',
>> - 'TSETB',
>> - 'TSETM',
>> - 'TSETR',
>> -
>> - # Calls and vararg handling. T = tail call.
>> - 'CALLM',
>> - 'CALL',
>> - 'CALLMT',
>> - 'CALLT',
>> - 'ITERC',
>> - 'ITERN',
>> - 'VARG',
>> - 'ISNEXT',
>> -
>> - # Returns.
>> - 'RETM',
>> - 'RET',
>> - 'RET0',
>> - 'RET1',
>> -
>> - # Loops and branches. I/J = interp/JIT.
>> - # I/C/L = init/call/loop.
>> - 'FORI',
>> - 'JFORI',
>> -
>> - 'FORL',
>> - 'IFORL',
>> - 'JFORL',
>> -
>> - 'ITERL',
>> - 'IITERL',
>> - 'JITERL',
>> -
>> - 'LOOP',
>> - 'ILOOP',
>> - 'JLOOP',
>> -
>> - 'JMP',
>> -
>> - # Function headers. I/J = interp/JIT.
>> - # F/V/C = fixarg/vararg/C func.
>> - 'FUNCF',
>> - 'IFUNCF',
>> - 'JFUNCF',
>> - 'FUNCV',
>> - 'IFUNCV',
>> - 'JFUNCV',
>> - 'FUNCC',
>> - 'FUNCCW',
>> -]
>> +BYTECODES = EnumBasedList('BCOp', 'BC__MAX', cut_prefix, 'BC_')
>>
>>
>> def proto_bc(proto):
>> @@ -1190,42 +1155,16 @@ def J(g):
>>
>>
>> # Matched `MMDEF(_)`.
>> -MM_NAMES = [
>> - 'index',
>> - 'newindex',
>> - 'gc',
>> - 'mode',
>> - 'eq',
>> - 'len',
>> - 'lt',
>> - 'le',
>> - 'concat',
>> - 'call',
>> - 'add',
>> - 'sub',
>> - 'mul',
>> - 'div',
>> - 'mod',
>> - 'pow',
>> - 'unm',
>> - 'metatable',
>> - 'tostring',
>> - # TODO: depends on LJ_HASFFI, see `MMDEF_FFI(_)`.
>> - 'new',
>> - # TODO: depends on LJ_52 || LJ_HASFFI, see `MMDEF_PAIRS(_)`.
>> - 'pairs',
>> - 'ipairs',
>> -]
>> -
>> -
>> -GCROOT_MMNAME = 0
>> -GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
>> -GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
>> -GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
>> +MM_NAMES = EnumBasedList('MMS', 'MM__MAX', cut_prefix, 'MM_')
>>
>>
>> # Get the name of the index in the predefined arrays.
>> def idx_name(field_name):
>> + GCROOT_MMNAME = 0
>> + GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
>> + GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
>> + GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
>> +
>> # Don't use **{ to be compatible with Python 2.
>> gcroot = {}
>> gcroot.update({
>> @@ -1477,140 +1416,7 @@ def cdataptr(cd):
>> # JIT engine.
>>
>>
>> -IRS = [
>> - # Guarded assertions.
>> - 'LT',
>> - 'GE',
>> - 'LE',
>> - 'GT',
>> -
>> - 'ULT',
>> - 'UGE',
>> - 'ULE',
>> - 'UGT',
>> -
>> - 'EQ',
>> - 'NE',
>> -
>> - 'ABC',
>> - 'RETF',
>> -
>> - # Miscellaneous ops.
>> - 'NOP',
>> - 'BASE',
>> - 'PVAL',
>> - 'GCSTEP',
>> - 'HIOP',
>> - 'LOOP',
>> - 'USE',
>> - 'PHI',
>> - 'RENAME',
>> - 'PROF',
>> -
>> - # Constants.
>> - 'KPRI',
>> - 'KINT',
>> - 'KGC',
>> - 'KPTR',
>> - 'KKPTR',
>> - 'KNULL',
>> - 'KNUM',
>> - 'KINT64',
>> - 'KSLOT',
>> -
>> - # Bit ops.
>> - 'BNOT',
>> - 'BSWAP',
>> - 'BAND',
>> - 'BOR',
>> - 'BXOR',
>> - 'BSHL',
>> - 'BSHR',
>> - 'BSAR',
>> - 'BROL',
>> - 'BROR',
>> -
>> - # Arithmetic ops. ORDER ARITH
>> - 'ADD',
>> - 'SUB',
>> - 'MUL',
>> - 'DIV',
>> - 'MOD',
>> - 'POW',
>> - 'NEG',
>> -
>> - 'ABS',
>> - 'LDEXP',
>> - 'MIN',
>> - 'MAX',
>> - 'FPMATH',
>> -
>> - # Overflow-checking arithmetic ops.
>> - 'ADDOV',
>> - 'SUBOV',
>> - 'MULOV',
>> -
>> - # Memory ops. A = array, H = hash, U = upvalue, F = field,
>> - # S = stack.
>> -
>> - # Memory references.
>> - 'AREF',
>> - 'HREFK',
>> - 'HREF',
>> - 'NEWREF',
>> - 'UREFO',
>> - 'UREFC',
>> - 'FREF',
>> - 'STRREF',
>> - 'LREF',
>> -
>> - # Loads and Stores. These must be in the same order.
>> - 'ALOAD',
>> - 'HLOAD',
>> - 'ULOAD',
>> - 'FLOAD',
>> - 'XLOAD',
>> - 'SLOAD',
>> - 'VLOAD',
>> -
>> - 'ASTORE',
>> - 'HSTORE',
>> - 'USTORE',
>> - 'FSTORE',
>> - 'XSTORE',
>> -
>> - # Allocations.
>> - 'SNEW',
>> - 'XSNEW',
>> - 'TNEW',
>> - 'TDUP',
>> - 'CNEW',
>> - 'CNEWI',
>> -
>> - # Buffer operations.
>> - 'BUFHDR',
>> - 'BUFPUT',
>> - 'BUFSTR',
>> -
>> - # Barriers.
>> - 'TBAR',
>> - 'OBAR',
>> - 'XBAR',
>> -
>> - # Type conversions.
>> - 'CONV',
>> - 'TOBIT',
>> - 'TOSTR',
>> - 'STRTO',
>> -
>> - # Calls.
>> - 'CALLN',
>> - 'CALLA',
>> - 'CALLL',
>> - 'CALLS',
>> - 'CALLXS',
>> - 'CARG',
>> -]
>> +IRS = EnumBasedList('IROp', 'IR__MAX', cut_prefix, 'IR_')
>>
>>
>> # Mode bits: Commutative, {Normal/Ref, Alloc, Load, Store},
>> @@ -1662,70 +1468,21 @@ def ir_mode(op):
>> return mode
>>
>>
>> -IRTYPES = [
>> - 'nil',
>> - 'fal',
>> - 'tru',
>> - 'lud',
>> - 'str',
>> - 'p32',
>> - 'thr',
>> - 'pro',
>> - 'fun',
>> - 'p64',
>> - 'cdt',
>> - 'tab',
>> - 'udt',
>> - 'flt',
>> - 'num',
>> - 'i8 ',
>> - 'u8 ',
>> - 'i16',
>> - 'u16',
>> - 'int',
>> - 'u32',
>> - 'i64',
>> - 'u64',
>> - 'sfp',
>> -]
>> +IRTYPES = EnumBasedList('IRType', 'IRT__MAX', lambda x: {
>> + 'IRT_LIGHTUD': 'lud',
>> + 'IRT_CDATA': 'cdt',
>> + 'IRT_UDATA': 'udt',
>> + 'IRT_FLOAT': 'flt',
>> + 'IRT_SOFTFP': 'sfp',
>> + }.get(x, cut_prefix(x, 'IRT_')[:3].ljust(3).lower()))
>>
>>
>> -IRT_NUM = 14
>> -assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
>> -
>> -
>> -IRFIELDS = [
>> - 'str.len',
>> - 'func.env',
>> - 'func.pc',
>> - 'func.ffid',
>> - 'thread.env',
>> - 'tab.meta',
>> - 'tab.array',
>> - 'tab.node',
>> - 'tab.asize',
>> - 'tab.hmask',
>> - 'tab.nomm',
>> - 'udata.meta',
>> - 'udata.udtype',
>> - 'udata.file',
>> - 'cdata.ctypeid',
>> - 'cdata.ptr',
>> - 'cdata.int',
>> - 'cdata.int64',
>> - 'cdata.int64_4',
>> -]
>> +IRFIELDS = EnumBasedList('IRFieldID', 'IRFL__MAX', lambda x:
>> + cut_prefix(x, 'IRFL_').lower().replace('_', '.', 1))
>>
>>
>> -IRFPMS = [
>> - 'floor',
>> - 'ceil',
>> - 'trunc',
>> - 'sqrt',
>> - 'log',
>> - 'log2',
>> - 'other'
>> -]
>> +IRFPMS = EnumBasedList('IRFPMathOp', 'IRFPM__MAX',
>> + lambda x: cut_prefix(x, 'IRFPM_').lower())
>>
>>
>> # Don't use *[ to be compatible with Python 2.
>> @@ -1754,112 +1511,7 @@ REGISTERS = {
>> }
>>
>>
>> -IR_CALLS = [
>> - 'lj_str_cmp',
>> - 'lj_str_find',
>> - 'lj_str_new',
>> - 'lj_strscan_num',
>> - 'lj_strfmt_int',
>> - 'lj_strfmt_num',
>> - 'lj_strfmt_char',
>> - 'lj_strfmt_putint',
>> - 'lj_strfmt_putnum',
>> - 'lj_strfmt_putquoted',
>> - 'lj_strfmt_putfxint',
>> - 'lj_strfmt_putfnum_int',
>> - 'lj_strfmt_putfnum_uint',
>> - 'lj_strfmt_putfnum',
>> - 'lj_strfmt_putfstr',
>> - 'lj_strfmt_putfchar',
>> - 'lj_buf_putmem',
>> - 'lj_buf_putstr',
>> - 'lj_buf_putchar',
>> - 'lj_buf_putstr_reverse',
>> - 'lj_buf_putstr_lower',
>> - 'lj_buf_putstr_upper',
>> - 'lj_buf_putstr_rep',
>> - 'lj_buf_puttab',
>> - 'lj_buf_tostr',
>> - 'lj_tab_new_ah',
>> - 'lj_tab_new1',
>> - 'lj_tab_dup',
>> - 'lj_tab_clear',
>> - 'lj_tab_newkey',
>> - 'lj_tab_len',
>> - 'lj_gc_step_jit',
>> - 'lj_gc_barrieruv',
>> - 'lj_mem_newgco',
>> - 'lj_math_random_step',
>> - 'lj_vm_modi',
>> - 'log10',
>> - 'exp',
>> - 'sin',
>> - 'cos',
>> - 'tan',
>> - 'asin',
>> - 'acos',
>> - 'atan',
>> - 'sinh',
>> - 'cosh',
>> - 'tanh',
>> - 'fputc',
>> - 'fwrite',
>> - 'fflush',
>> - 'lj_vm_floor',
>> - 'lj_vm_ceil',
>> - 'lj_vm_trunc',
>> - 'sqrt',
>> - 'log',
>> - 'lj_vm_log2',
>> - 'pow',
>> - 'atan2',
>> - 'ldexp',
>> - 'lj_vm_tobit',
>> - 'softfp_add',
>> - 'softfp_sub',
>> - 'softfp_mul',
>> - 'softfp_div',
>> - 'softfp_cmp',
>> - 'softfp_i2d',
>> - 'softfp_d2i',
>> - 'lj_vm_sfmin',
>> - 'lj_vm_sfmax',
>> - 'lj_vm_tointg',
>> - 'softfp_ui2d',
>> - 'softfp_f2d',
>> - 'softfp_d2ui',
>> - 'softfp_d2f',
>> - 'softfp_i2f',
>> - 'softfp_ui2f',
>> - 'softfp_f2i',
>> - 'softfp_f2ui',
>> - 'fp64_l2d',
>> - 'fp64_ul2d',
>> - 'fp64_l2f',
>> - 'fp64_ul2f',
>> - 'fp64_d2l',
>> - 'fp64_d2ul',
>> - 'fp64_f2l',
>> - 'fp64_f2ul',
>> - 'lj_carith_divi64',
>> - 'lj_carith_divu64',
>> - 'lj_carith_modi64',
>> - 'lj_carith_modu64',
>> - 'lj_carith_powi64',
>> - 'lj_carith_powu64',
>> - 'lj_cdata_newv',
>> - 'lj_cdata_setfin',
>> - 'strlen',
>> - 'memcpy',
>> - 'memset',
>> - 'lj_vm_errno',
>> - 'lj_carith_mul64',
>> - 'lj_carith_shl64',
>> - 'lj_carith_shr64',
>> - 'lj_carith_sar64',
>> - 'lj_carith_rol64',
>> - 'lj_carith_ror64',
>> -]
>> +IR_CALLS = EnumBasedList('IRCallID', 'IRCALL__MAX', cut_prefix, 'IRCALL_')
>>
>>
>> def regname(reg_number):
>> @@ -1995,6 +1647,8 @@ def irt_isguard(t):
>>
>>
>> def irt_toitype(irt):
>> + IRT_NUM = 14
>> + assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
>> t = irt_type(irt)
>> if LJ_DUALNUM and t > IRT_NUM:
>> return LJ_T['NUMX']
--
Best regards,
Mikhail Elhimov
[-- Attachment #2: Type: text/html, Size: 30070 bytes --]
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [Tarantool-patches] [PATCH luajit] dbg: avoid hardcoded enums (get them from target)
2026-09-17 9:52 ` Sergey Bronnikov via Tarantool-patches
2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
@ 2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
1 sibling, 0 replies; 4+ messages in thread
From: Mikhail Elhimov via Tarantool-patches @ 2026-09-17 22:42 UTC (permalink / raw)
To: Sergey Bronnikov, Sergey Kaplun, Evgeniy Temirgaleev; +Cc: tarantool-patches
[-- Attachment #1: Type: text/plain, Size: 27176 bytes --]
Hi, Sergey!
Thanks for the review! See my comments below
On 17.09.2026 12:52, Sergey Bronnikov wrote:
>
> Hi, Mikhail,
>
> thanks for the patch! Please see my comments.
>
> Sergey
>
> On 9/9/26 11:49, Mikhail Elhimov wrote:
>> Besides reducing lines of code this way the extension become compatible
>> with various version of luajit because different version might use
> s/luajit/LuaJIT/
Fixed
>> different set of enum members (newer version might introduce additional
>> BC/IR/etc.)
>>
>> Prior to this patch string was used as a debugger-agnostic way to
>> specify type, but this way might not work in case of enum because
>> debugging information might be optimized out if no variable of such enum
>> type is declared and its members are used only as a predefined
>> constants. The type of enum is needed to be able to get human readable
>> value (member name instead of number). The alternative way to get enum
>> type is get it from the value, i.e. somehow obtain the value that is of
>> enum type and then get its type object.
>>
>> To do that, separate method to create enum value was introduced in
>> Debugger API, LLDB value class was monkey-patched to get value type in
>> the same way as GDB value class does and 'cast' method was adjusted to
>> accept also type object, not only string.
>>
>> Other adjustments:
>> - [lldb] 'eval' adjusted to return object of the same type as 'cast'
>> method (this improves consistency)
>> - [gdb] in 'eval' method dropped check of the value returned by
>> gdb.parse_and_eval() as it would fail also for expression like '0'
>> (looks like a kind of legacy code that is not needed now).
>> - [gdb/lldb] renamed eval argument to reflect its meaning (it is
>> an expression rather than a command).
>>
>> Closes tarantool/tarantool#13094
>> ---
>> This patch is to be applied after the 'fix mapping of FPMATHOP' patch.
>>
>> Branch:https://github.com/tarantool/luajit/tree/elhimov/gh-13094-dbg-use-enums-from-inferior
>> Related issue:https://github.com/tarantool/tarantool/issues/13094
>>
>> src/luajit_dbg.py | 628 +++++++++++-----------------------------------
>> 1 file changed, 141 insertions(+), 487 deletions(-)
>
> The existing test suite provides good coverage for `BCMODE`,
> `BYTECODES`, `IRS`, `IRTYPES`, `IRFIELDS`, and `GCROOT_*`,
>
> but the new code path handling "optimized enum type names" (where the
> type is looked up via a member)
>
> is not specifically tested. It is worth verifying that an `IRFieldID`
> in the test binary actually triggers
>
> the fallback mechanism; otherwise, the LLDB fallback remains untested.
>
At the moment the only way I found to verify that the fallback mechanism
was triggered is:
val = dbg.create_enum_value('IRFieldID', 'IRFL__MAX')
if val is not None and val.type.name == '(unnamed enum)':
# fallback triggered
But it does work only from the extension and I see no way to get access
to the enum_type from the tests level. So now we are only assuming that
the fallback mechanism triggers in case of `IRFL__MAX` of `IRFieldID` enum.
>> diff --git a/src/luajit_dbg.py b/src/luajit_dbg.py
>> index 76001b7d..1ac1d275 100644
>> --- a/src/luajit_dbg.py
>> +++ b/src/luajit_dbg.py
>> @@ -110,9 +110,21 @@ class Debugger(object):
>> self.write('{} command initialized\n'.format(name))
>> self.write('LuaJIT debug extension is successfully loaded\n')
>>
>> + def cast(self, tp, val):
>> + '''Cast the value to the given type (it is either C type string
>> + or the debugger-specific type object).'''
>> + if isinstance(tp, str):
>> + tp = self._dbgtype(tp)
>> + return self._cast(tp, val)
>> +
>> + @abc.abstractmethod
>> + def _cast(self, tp, val):
>> + '''Cast the value to the debugger-specific type.'''
>> + pass
>> +
>> @abc.abstractmethod
>> - def cast(self, typestr, val):
>> - '''Cast the value to the required C type.'''
>> + def _dbgtype(self, typestr):
>> + '''Convert C type string into debugger-specific type object.'''
>> pass
>>
>> @abc.abstractmethod
>> @@ -146,8 +158,9 @@ class Debugger(object):
>> pass
>>
>> @abc.abstractmethod
>> - def eval(self, command):
>> - '''Parse and evaluate the given debugger command.'''
>> + def eval(self, expr):
>> + '''Parse and evaluate the given debugger expression.
>> + Return debugger-specific value.'''
>> pass
>>
>> @abc.abstractmethod
>> @@ -187,6 +200,13 @@ class Debugger(object):
>> '''Register the command with the corresponding name.'''
>> pass
>>
>> + @abc.abstractmethod
>> + def create_enum_value(self, enum_name, enum_member_name):
>> + '''Return debugger-specific value object that represents
>> + the given enum member.
>> + '''
>> + pass
>> +
>> @abc.abstractproperty
>> def LJBase(self):
>> '''Base command class.
>> @@ -211,8 +231,9 @@ class _GDBDebugger(Debugger):
>> super(_GDBDebugger, self).__init__()
>> self.CONNECTED = False
>>
>> - def cast(self, typestr, val):
>> - return gdb.Value(val).cast(self._dbgtype(typestr))
>> + def _cast(self, tp, val):
>> + assert isinstance(tp, gdb.Type)
>> + return gdb.Value(val).cast(tp)
>>
>> def sizeof(self, typestr):
>> return self._dbgtype(typestr).sizeof
>> @@ -268,14 +289,11 @@ class _GDBDebugger(Debugger):
>> else:
>> return None
>>
>> - def eval(self, command):
>> - if not command:
>> + def eval(self, expr):
>> + if not expr:
>> return None
>>
>> - ret = gdb.parse_and_eval(command)
>> - if not ret:
>> - raise gdb.GdbError('table argument empty')
>> - return ret
>> + return gdb.parse_and_eval(expr)
>>
>> def detect_arch(self):
>> if hasattr(self, 'arch'):
>> @@ -340,6 +358,9 @@ class _GDBDebugger(Debugger):
>> def register_command(self, command, name):
>> command(name)
>>
>> + def create_enum_value(self, enum_name, enum_member_name):
>> + return self.eval(enum_member_name)
>
> For GDB, `create_enum_value` returns `self.eval(enum_member_name)`,
> and `EnumBasedList` uses `.type` as the enum type.
>
> If, in a certain version of GDB, `parse_and_eval('BC__MAX').type`
> evaluates to `int`, `str()` will return numbers,
>
> and the list will be populated with "0", "1", etc., without an
> explicit error.
>
> It is worth adding an assertion to verify that the type is an enum
> (`gdb.TYPE_CODE_ENUM`).
>
If such kind of expression is recognized as enum member it is evaluated
to the value of enum type. If not then it's just an invalid expression
and `gdb.parse_and_eval()` fails with exception. The mentioned scenario
with int instead of enum looks quite weird for me as it would mean that
gdb recognize enum value (which means that there is information about
corresponding enum type), but for some inexpicable reason "decides" to
cast it to another type. Why? Is that just a concern, or have you seen
that kind of behavior?
However I have add assertions to both gdb and lldb implementation, but
considering the above I would treat them as a verification that correct
arguments were passed into `create_enum_value()` rather than a
verification that such a spontaneous casting-to-int didn't happen.
>> +
>> class LJBase(gdb and gdb.Command or object):
>> def __init__(ljbase, name):
>> # XXX Fragile: Though the command initialization looks
>> @@ -378,6 +399,10 @@ class _LLDBDebugger(Debugger):
>> lldb.eBasicTypeInt128
>> ]
>>
>> + def _lldb_tp_isenum(self, tp):
>> + return tp.GetCanonicalType().GetTypeClass() == \
>> + lldb.eTypeClassEnumeration
>> +
>> def _lldb_value_from_raw(self, raw_value, size, tp):
>> isfp = self._lldb_tp_isfp(tp)
>> if isfp:
>> @@ -491,8 +516,7 @@ class _LLDBDebugger(Debugger):
>> # Instead of default GetSummary.
>> if not lldbval.sbvalue.TypeIsPointerType():
>> tp = lldbval.sbvalue.GetType()
>> - is_float = self._lldb_tp_isfp(tp)
>> - if is_float:
>> + if self._lldb_tp_isfp(tp) or self._lldb_tp_isenum(tp):
>> return lldbval.sbvalue.GetValue()
>> else:
>> return str(int(lldbval))
>> @@ -530,6 +554,9 @@ class _LLDBDebugger(Debugger):
>> else:
>> return int(lldbval) - int(other)
>>
>> + def lldb_gettype(lldbval):
>> + return lldbval.sbvalue.type
>> +
>> super(_LLDBDebugger, self).__init__()
>> self.target = lldb.debugger.GetSelectedTarget()
>> # Monkey-patch the lldb.value class.
>> @@ -545,6 +572,7 @@ class _LLDBDebugger(Debugger):
>> lldb.value.__ror__ = lldb__or__ # Same semantics.
>> lldb.value.__str__ = lldb__str__
>> lldb.value.__sub__ = lldb__sub__
>> + lldb.value.type = property(lldb_gettype)
>>
>> def lldb_major_version():
>> version_string = lldb.SBDebugger.GetVersionString()
>> @@ -568,11 +596,11 @@ class _LLDBDebugger(Debugger):
>> self.dbgtype_cache[typestr] = dbgtype
>> return dbgtype
>>
>> - def cast(self, typestr, val):
>> + def _cast(self, tp, val):
>> + assert isinstance(tp, lldb.SBType)
>> if isinstance(val, lldb.value):
>> val = val.sbvalue
>> elif type(val) is int:
>> - tp = self._dbgtype(typestr)
>> return self._lldb_value_from_raw(val, tp.GetByteSize(), tp)
>> elif not isinstance(val, lldb.SBValue):
>> raise Exception(
>> @@ -582,7 +610,6 @@ class _LLDBDebugger(Debugger):
>> # XXX: Simply SBValue.Cast() works incorrectly since it
>> # may take the 8 bytes of memory instead of 4, before the
>> # cast. Construct the value on the fly.
>> - tp = self._dbgtype(typestr)
>> if self._lldb_tp_isfp(tp):
>> rawval = float(val.GetValue())
>> elif self._lldb_tp_issigned(tp):
>> @@ -670,15 +697,15 @@ class _LLDBDebugger(Debugger):
>> else:
>> return None
>>
>> - def eval(self, command):
>> - if not command:
>> + def eval(self, expr):
>> + if not expr:
>> return None
>>
>> process = self.target.GetProcess()
>> thread = process.GetSelectedThread()
>> frame = thread.GetSelectedFrame()
>> - ret = frame.EvaluateExpression(command)
>> - return ret
>> + ret = frame.EvaluateExpression(expr)
>> + return lldb.value(ret)
>>
>> def detect_arch(self):
>> if hasattr(self, 'arch'):
>> @@ -724,6 +751,37 @@ class _LLDBDebugger(Debugger):
>> )
>> )
>>
>> + def create_enum_value(self, enum_name, enum_member_name):
>> + val = self.eval(enum_name + "::" + enum_member_name)
>> + if val.sbvalue.IsValid() and val.sbvalue.error.Success():
>> + return val
>> +
>> + # LLDB uses enum name in expression above but debugging information
>> + # about enum name migth be optimized out if no variable of the given
> s/migth/might/
Fixed
>> + # enum type is declared and its members are only used as the predefined
>> + # constants (like IRFieldID).
>> +
>> + # In this case the above method doesn't work so trying to discover
>> + # enum type by the given enum member.
>> +
>> + def find_enum_type_member(enum_type, enum_member_name):
>> + # SBTypeEnumMemberList supports members iteration and [] access
>> + # (both by index and by member name) only starting from lldb-12
>> + # so this implementation is used to handle earlier versions.
>> + members = enum_type.GetEnumMembers()
>> + for i in range(members.GetSize()):
>> + item = members.GetTypeEnumMemberAtIndex(i)
>> + if item.name == enum_member_name:
>> + return item
>> + return None
>> +
>> + for m in self.target.modules:
>> + for et in m.GetTypes(lldb.eTypeClassEnumeration):
>
> The construct `for et in m.GetTypes(lldb.eTypeClassEnumeration)`
> relies on the `__iter__` method of `SBTypeList`.
>
> There is a comment above noting that `SBTypeEnumMemberList` could not
> be iterated over
>
> prior to lldb 12 and no such guarantee exists for `SBTypeList`. For
> the sake of consistency and
>
> compatibility, it is better to use: list = m.GetTypes(...) followed by
>
> for i in range(list.GetSize()):
>
> et = list.GetTypeAtIndex(i)
>
I find this kind of iteration as the most convenient so I'm trying to
use it unless there is some reason against it. `SBTypeList` is just
another type that has no relation to `SBTypeEnumMemberList` and it is
not affected with the mentioned "problem". And I'm not sure I understand
correctly what does mean "no guarantee exists for `SBTypeList`" since we
are able to check it in llvm repo. I checked by tag llvmorg-4.0.0
(version `lldb` subdirectory first appeared in) and it looks like
`SBTypeList` had supported iteration initially. Thus it turns out that
the only reason left to change the iteration approach is to make it
similar to iteration over `SBTypeEnumMemberList`. Then what about
iteration over modules a line above? Should it be changed to look
similar as well? Something like this:
for i_module in range(self.target.GetNumModules()):
list = self.target.GetModuleAtIndex(i_module).GetTypes(lldb.eTypeClassEnumeration)
for i in range(list.GetSize()):
et = list.GetTypeAtIndex(i)
>> + et_member = find_enum_type_member(et, enum_member_name)
>> + if et_member is notNone:first ver
>> + return self.cast(et, et_member.unsigned)
>> + return None
>> +
>> class LJBase(object):
>> # Ignore given parameters by LLDB.
>> def __init__(ljbase, debugger, unused):
>> @@ -800,6 +858,45 @@ def strx64(val):
>> return re.sub('L?$', '', hex(int(tou64(val))))
>>
>>
>> +class EnumBasedList(object):
>> + def __init__(self, enum_name, max_enum_member, map_func=None,
>> + *map_func_extra_args):
>> + self.__enum_name = enum_name
>> + self.__max_enum_member = max_enum_member
>> + self.__map_func = map_func
>> + self.__map_func_extra_args = map_func_extra_args
>> + # Lazy initialization (see __get_items method) as the required
>> + # information might be unavailable at this moment.
>> + self.__items = None
>> +
>> + def __iter__(self):
>> + return iter(self.__get_items())
>> +
>> + def __getitem__(self, key):
>> + return self.__get_items()[key]
>> +
>> + def __len__(self):
>> + return len(self.__get_items())
>> +
>> + def __get_items(self):
>> + if self.__items is None:
>> + max_enum_value = dbg.create_enum_value(
>> + self.__enum_name, self.__max_enum_member
>> + )
>> + items = []
>> + for i in range(dbg.cast('int', max_enum_value)):
>> + item = str(dbg.cast(max_enum_value.type, dbg.eval(str(i))))
>> + if self.__map_func:
>> + item = self.__map_func(item, *self.__map_func_extra_args)
>> + items.append(item)
>> + self.__items = items
>> + return self.__items
>> +
>> +
>> +def cut_prefix(s, prefix):
>> + return s[len(prefix):] if s.startswith(prefix) else s
>> +
>> +
>> # Types and TValues.
>>
>>
>> @@ -877,10 +974,7 @@ def bc_d(ins):
>> return int(ins) >> 16
>>
>>
>> -BCMODE = [
>> - 'none', 'dst', 'base', 'var', 'rbase', 'uv',
>> - 'lit', 'lits', 'pri', 'num', 'str', 'tab', 'func', 'jump', 'cdata',
>> -]
>> +BCMODE = EnumBasedList('BCMode', 'BCM_max', cut_prefix, 'BCM')
>>
>>
>> lj_bc_mode_ = None
>> @@ -906,136 +1000,7 @@ def bcmode_cd(op):
>> return int((lj_bc_mode()[op] >> 7) & 15)
>>
>>
>> -# Unfortunately, there is no place in the VM except the generated
>> -# Lua table, where the bytecode names are stored. So duplicate
>> -# them here.
>> -BYTECODES = [
>> - # Comparison ops. ORDER OPR.
>> - 'ISLT',
>> - 'ISGE',
>> - 'ISLE',
>> - 'ISGT',
>> -
>> - 'ISEQV',
>> - 'ISNEV',
>> - 'ISEQS',
>> - 'ISNES',
>> - 'ISEQN',
>> - 'ISNEN',
>> - 'ISEQP',
>> - 'ISNEP',
>> -
>> - # Unary test and copy ops.
>> - 'ISTC',
>> - 'ISFC',
>> - 'IST',
>> - 'ISF',
>> - 'ISTYPE',
>> - 'ISNUM',
>> - 'MOV',
>> - 'NOT',
>> - 'UNM',
>> - 'LEN',
>> - 'ADDVN',
>> - 'SUBVN',
>> - 'MULVN',
>> - 'DIVVN',
>> - 'MODVN',
>> -
>> - # Binary ops. ORDER OPR.
>> - 'ADDNV',
>> - 'SUBNV',
>> - 'MULNV',
>> - 'DIVNV',
>> - 'MODNV',
>> -
>> - 'ADDVV',
>> - 'SUBVV',
>> - 'MULVV',
>> - 'DIVVV',
>> - 'MODVV',
>> -
>> - 'POW',
>> - 'CAT',
>> -
>> - # Constant ops.
>> - 'KSTR',
>> - 'KCDATA',
>> - 'KSHORT',
>> - 'KNUM',
>> - 'KPRI',
>> - 'KNIL',
>> -
>> - # Upvalue and function ops.
>> - 'UGET',
>> - 'USETV',
>> - 'USETS',
>> - 'USETN',
>> - 'USETP',
>> - 'UCLO',
>> - 'FNEW',
>> -
>> - # Table ops.
>> - 'TNEW',
>> - 'TDUP',
>> - 'GGET',
>> - 'GSET',
>> - 'TGETV',
>> - 'TGETS',
>> - 'TGETB',
>> - 'TGETR',
>> - 'TSETV',
>> - 'TSETS',
>> - 'TSETB',
>> - 'TSETM',
>> - 'TSETR',
>> -
>> - # Calls and vararg handling. T = tail call.
>> - 'CALLM',
>> - 'CALL',
>> - 'CALLMT',
>> - 'CALLT',
>> - 'ITERC',
>> - 'ITERN',
>> - 'VARG',
>> - 'ISNEXT',
>> -
>> - # Returns.
>> - 'RETM',
>> - 'RET',
>> - 'RET0',
>> - 'RET1',
>> -
>> - # Loops and branches. I/J = interp/JIT.
>> - # I/C/L = init/call/loop.
>> - 'FORI',
>> - 'JFORI',
>> -
>> - 'FORL',
>> - 'IFORL',
>> - 'JFORL',
>> -
>> - 'ITERL',
>> - 'IITERL',
>> - 'JITERL',
>> -
>> - 'LOOP',
>> - 'ILOOP',
>> - 'JLOOP',
>> -
>> - 'JMP',
>> -
>> - # Function headers. I/J = interp/JIT.
>> - # F/V/C = fixarg/vararg/C func.
>> - 'FUNCF',
>> - 'IFUNCF',
>> - 'JFUNCF',
>> - 'FUNCV',
>> - 'IFUNCV',
>> - 'JFUNCV',
>> - 'FUNCC',
>> - 'FUNCCW',
>> -]
>> +BYTECODES = EnumBasedList('BCOp', 'BC__MAX', cut_prefix, 'BC_')
>>
>>
>> def proto_bc(proto):
>> @@ -1190,42 +1155,16 @@ def J(g):
>>
>>
>> # Matched `MMDEF(_)`.
>> -MM_NAMES = [
>> - 'index',
>> - 'newindex',
>> - 'gc',
>> - 'mode',
>> - 'eq',
>> - 'len',
>> - 'lt',
>> - 'le',
>> - 'concat',
>> - 'call',
>> - 'add',
>> - 'sub',
>> - 'mul',
>> - 'div',
>> - 'mod',
>> - 'pow',
>> - 'unm',
>> - 'metatable',
>> - 'tostring',
>> - # TODO: depends on LJ_HASFFI, see `MMDEF_FFI(_)`.
>> - 'new',
>> - # TODO: depends on LJ_52 || LJ_HASFFI, see `MMDEF_PAIRS(_)`.
>> - 'pairs',
>> - 'ipairs',
>> -]
>> -
>> -
>> -GCROOT_MMNAME = 0
>> -GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
>> -GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
>> -GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
>> +MM_NAMES = EnumBasedList('MMS', 'MM__MAX', cut_prefix, 'MM_')
>>
>>
>> # Get the name of the index in the predefined arrays.
>> def idx_name(field_name):
>> + GCROOT_MMNAME = 0
>> + GCROOT_BASEMT = GCROOT_MMNAME + len(MM_NAMES)
>> + GCROOT_IO_INPUT = GCROOT_BASEMT + i2notu32(LJ_T['NUMX']) + 1
>> + GCROOT_IO_OUTPUT = GCROOT_IO_INPUT + 1
>> +
>> # Don't use **{ to be compatible with Python 2.
>> gcroot = {}
>> gcroot.update({
>> @@ -1477,140 +1416,7 @@ def cdataptr(cd):
>> # JIT engine.
>>
>>
>> -IRS = [
>> - # Guarded assertions.
>> - 'LT',
>> - 'GE',
>> - 'LE',
>> - 'GT',
>> -
>> - 'ULT',
>> - 'UGE',
>> - 'ULE',
>> - 'UGT',
>> -
>> - 'EQ',
>> - 'NE',
>> -
>> - 'ABC',
>> - 'RETF',
>> -
>> - # Miscellaneous ops.
>> - 'NOP',
>> - 'BASE',
>> - 'PVAL',
>> - 'GCSTEP',
>> - 'HIOP',
>> - 'LOOP',
>> - 'USE',
>> - 'PHI',
>> - 'RENAME',
>> - 'PROF',
>> -
>> - # Constants.
>> - 'KPRI',
>> - 'KINT',
>> - 'KGC',
>> - 'KPTR',
>> - 'KKPTR',
>> - 'KNULL',
>> - 'KNUM',
>> - 'KINT64',
>> - 'KSLOT',
>> -
>> - # Bit ops.
>> - 'BNOT',
>> - 'BSWAP',
>> - 'BAND',
>> - 'BOR',
>> - 'BXOR',
>> - 'BSHL',
>> - 'BSHR',
>> - 'BSAR',
>> - 'BROL',
>> - 'BROR',
>> -
>> - # Arithmetic ops. ORDER ARITH
>> - 'ADD',
>> - 'SUB',
>> - 'MUL',
>> - 'DIV',
>> - 'MOD',
>> - 'POW',
>> - 'NEG',
>> -
>> - 'ABS',
>> - 'LDEXP',
>> - 'MIN',
>> - 'MAX',
>> - 'FPMATH',
>> -
>> - # Overflow-checking arithmetic ops.
>> - 'ADDOV',
>> - 'SUBOV',
>> - 'MULOV',
>> -
>> - # Memory ops. A = array, H = hash, U = upvalue, F = field,
>> - # S = stack.
>> -
>> - # Memory references.
>> - 'AREF',
>> - 'HREFK',
>> - 'HREF',
>> - 'NEWREF',
>> - 'UREFO',
>> - 'UREFC',
>> - 'FREF',
>> - 'STRREF',
>> - 'LREF',
>> -
>> - # Loads and Stores. These must be in the same order.
>> - 'ALOAD',
>> - 'HLOAD',
>> - 'ULOAD',
>> - 'FLOAD',
>> - 'XLOAD',
>> - 'SLOAD',
>> - 'VLOAD',
>> -
>> - 'ASTORE',
>> - 'HSTORE',
>> - 'USTORE',
>> - 'FSTORE',
>> - 'XSTORE',
>> -
>> - # Allocations.
>> - 'SNEW',
>> - 'XSNEW',
>> - 'TNEW',
>> - 'TDUP',
>> - 'CNEW',
>> - 'CNEWI',
>> -
>> - # Buffer operations.
>> - 'BUFHDR',
>> - 'BUFPUT',
>> - 'BUFSTR',
>> -
>> - # Barriers.
>> - 'TBAR',
>> - 'OBAR',
>> - 'XBAR',
>> -
>> - # Type conversions.
>> - 'CONV',
>> - 'TOBIT',
>> - 'TOSTR',
>> - 'STRTO',
>> -
>> - # Calls.
>> - 'CALLN',
>> - 'CALLA',
>> - 'CALLL',
>> - 'CALLS',
>> - 'CALLXS',
>> - 'CARG',
>> -]
>> +IRS = EnumBasedList('IROp', 'IR__MAX', cut_prefix, 'IR_')
>>
>>
>> # Mode bits: Commutative, {Normal/Ref, Alloc, Load, Store},
>> @@ -1662,70 +1468,21 @@ def ir_mode(op):
>> return mode
>>
>>
>> -IRTYPES = [
>> - 'nil',
>> - 'fal',
>> - 'tru',
>> - 'lud',
>> - 'str',
>> - 'p32',
>> - 'thr',
>> - 'pro',
>> - 'fun',
>> - 'p64',
>> - 'cdt',
>> - 'tab',
>> - 'udt',
>> - 'flt',
>> - 'num',
>> - 'i8 ',
>> - 'u8 ',
>> - 'i16',
>> - 'u16',
>> - 'int',
>> - 'u32',
>> - 'i64',
>> - 'u64',
>> - 'sfp',
>> -]
>> +IRTYPES = EnumBasedList('IRType', 'IRT__MAX', lambda x: {
>> + 'IRT_LIGHTUD': 'lud',
>> + 'IRT_CDATA': 'cdt',
>> + 'IRT_UDATA': 'udt',
>> + 'IRT_FLOAT': 'flt',
>> + 'IRT_SOFTFP': 'sfp',
>> + }.get(x, cut_prefix(x, 'IRT_')[:3].ljust(3).lower()))
>>
>>
>> -IRT_NUM = 14
>> -assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
>> -
>> -
>> -IRFIELDS = [
>> - 'str.len',
>> - 'func.env',
>> - 'func.pc',
>> - 'func.ffid',
>> - 'thread.env',
>> - 'tab.meta',
>> - 'tab.array',
>> - 'tab.node',
>> - 'tab.asize',
>> - 'tab.hmask',
>> - 'tab.nomm',
>> - 'udata.meta',
>> - 'udata.udtype',
>> - 'udata.file',
>> - 'cdata.ctypeid',
>> - 'cdata.ptr',
>> - 'cdata.int',
>> - 'cdata.int64',
>> - 'cdata.int64_4',
>> -]
>> +IRFIELDS = EnumBasedList('IRFieldID', 'IRFL__MAX', lambda x:
>> + cut_prefix(x, 'IRFL_').lower().replace('_', '.', 1))
>>
>>
>> -IRFPMS = [
>> - 'floor',
>> - 'ceil',
>> - 'trunc',
>> - 'sqrt',
>> - 'log',
>> - 'log2',
>> - 'other'
>> -]
>> +IRFPMS = EnumBasedList('IRFPMathOp', 'IRFPM__MAX',
>> + lambda x: cut_prefix(x, 'IRFPM_').lower())
>>
>>
>> # Don't use *[ to be compatible with Python 2.
>> @@ -1754,112 +1511,7 @@ REGISTERS = {
>> }
>>
>>
>> -IR_CALLS = [
>> - 'lj_str_cmp',
>> - 'lj_str_find',
>> - 'lj_str_new',
>> - 'lj_strscan_num',
>> - 'lj_strfmt_int',
>> - 'lj_strfmt_num',
>> - 'lj_strfmt_char',
>> - 'lj_strfmt_putint',
>> - 'lj_strfmt_putnum',
>> - 'lj_strfmt_putquoted',
>> - 'lj_strfmt_putfxint',
>> - 'lj_strfmt_putfnum_int',
>> - 'lj_strfmt_putfnum_uint',
>> - 'lj_strfmt_putfnum',
>> - 'lj_strfmt_putfstr',
>> - 'lj_strfmt_putfchar',
>> - 'lj_buf_putmem',
>> - 'lj_buf_putstr',
>> - 'lj_buf_putchar',
>> - 'lj_buf_putstr_reverse',
>> - 'lj_buf_putstr_lower',
>> - 'lj_buf_putstr_upper',
>> - 'lj_buf_putstr_rep',
>> - 'lj_buf_puttab',
>> - 'lj_buf_tostr',
>> - 'lj_tab_new_ah',
>> - 'lj_tab_new1',
>> - 'lj_tab_dup',
>> - 'lj_tab_clear',
>> - 'lj_tab_newkey',
>> - 'lj_tab_len',
>> - 'lj_gc_step_jit',
>> - 'lj_gc_barrieruv',
>> - 'lj_mem_newgco',
>> - 'lj_math_random_step',
>> - 'lj_vm_modi',
>> - 'log10',
>> - 'exp',
>> - 'sin',
>> - 'cos',
>> - 'tan',
>> - 'asin',
>> - 'acos',
>> - 'atan',
>> - 'sinh',
>> - 'cosh',
>> - 'tanh',
>> - 'fputc',
>> - 'fwrite',
>> - 'fflush',
>> - 'lj_vm_floor',
>> - 'lj_vm_ceil',
>> - 'lj_vm_trunc',
>> - 'sqrt',
>> - 'log',
>> - 'lj_vm_log2',
>> - 'pow',
>> - 'atan2',
>> - 'ldexp',
>> - 'lj_vm_tobit',
>> - 'softfp_add',
>> - 'softfp_sub',
>> - 'softfp_mul',
>> - 'softfp_div',
>> - 'softfp_cmp',
>> - 'softfp_i2d',
>> - 'softfp_d2i',
>> - 'lj_vm_sfmin',
>> - 'lj_vm_sfmax',
>> - 'lj_vm_tointg',
>> - 'softfp_ui2d',
>> - 'softfp_f2d',
>> - 'softfp_d2ui',
>> - 'softfp_d2f',
>> - 'softfp_i2f',
>> - 'softfp_ui2f',
>> - 'softfp_f2i',
>> - 'softfp_f2ui',
>> - 'fp64_l2d',
>> - 'fp64_ul2d',
>> - 'fp64_l2f',
>> - 'fp64_ul2f',
>> - 'fp64_d2l',
>> - 'fp64_d2ul',
>> - 'fp64_f2l',
>> - 'fp64_f2ul',
>> - 'lj_carith_divi64',
>> - 'lj_carith_divu64',
>> - 'lj_carith_modi64',
>> - 'lj_carith_modu64',
>> - 'lj_carith_powi64',
>> - 'lj_carith_powu64',
>> - 'lj_cdata_newv',
>> - 'lj_cdata_setfin',
>> - 'strlen',
>> - 'memcpy',
>> - 'memset',
>> - 'lj_vm_errno',
>> - 'lj_carith_mul64',
>> - 'lj_carith_shl64',
>> - 'lj_carith_shr64',
>> - 'lj_carith_sar64',
>> - 'lj_carith_rol64',
>> - 'lj_carith_ror64',
>> -]
>> +IR_CALLS = EnumBasedList('IRCallID', 'IRCALL__MAX', cut_prefix, 'IRCALL_')
>>
>>
>> def regname(reg_number):
>> @@ -1995,6 +1647,8 @@ def irt_isguard(t):
>>
>>
>> def irt_toitype(irt):
>> + IRT_NUM = 14
>> + assert IRTYPES[IRT_NUM] == 'num', 'incorrect IRT_NUM definition'
>> t = irt_type(irt)
>> if LJ_DUALNUM and t > IRT_NUM:
>> return LJ_T['NUMX']
--
Best regards,
Mikhail Elhimov
[-- Attachment #2: Type: text/html, Size: 30070 bytes --]
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-17 22:43 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-09 8:49 [Tarantool-patches] [PATCH luajit] dbg: avoid hardcoded enums (get them from target) Mikhail Elhimov via Tarantool-patches
2026-09-17 9:52 ` Sergey Bronnikov via Tarantool-patches
2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
2026-09-17 22:42 ` Mikhail Elhimov via Tarantool-patches
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox