Follow-up to the scalar-only fix in #8754, per aothms's direct request on
that PR ("Please do make all int types consistent") and his own original
2023 design intent on issue #3058 ("make all integers (incl. schema
namespaces) an int64_t"). Widens the remaining inconsistent spots now that
compatibility isn't a constraint on this v0.9 branch:
- Integer aggregates (IfcTriangulatedFaceSet.CoordIndex and similar
List<int> attributes), including the SWIG to_vec_int/to_vec_vec_int
helpers, which previously silently truncated via static_cast<int> on the
Python-set path - the same bug class as the original scalar issue.
- The schema code generator (express/mapping.py's integer type mapping),
and all 12 generated schema header/source pairs regenerated to match, so
every schema-typed getter/setter (e.g. IfcOwnerHistory::CreationDate) is
int64_t end to end, not just the dynamic attribute-value path.
Instance/reference identifiers (STEP #123 ids) are deliberately left at
32-bit: they're a file-local index into internal maps, not an EXPRESS
domain value an application chooses, and no realistic STEP file has
billions of entities. The lexer's Token_IDENTIFIER parsing still funnels
through a 32-bit int for this reason - flagged as a known, low-risk gap
rather than fixed, since fixing it would mean touching indexing/hashing
code for no realistic benefit.
Verified: original PR's round-trip tests extended with aggregate cases
(IfcTriangulatedFaceSet.CoordIndex, InnerCoordIndices) at 64-bit boundary
values, in memory and through STEP text, IFC2X3 and IFC4. A standalone C++
program exercising the generated schema API directly (Ifc4::IfcOwnerHistory
::setCreationDate/CreationDate, IfcTriangulatedFaceSet::setCoordIndex/
CoordIndex) confirms int64_t end to end, bypassing SWIG. Full build
(BUILD_IFCGEOM, WITH_OPENCASCADE, BUILD_IFCPYTHON, IFC2X3+IFC4) clean.
test/util/test_attribute.py and test_file.py pass unchanged.
This contribution was produced with the assistance of an AI coding tool.
Setting an IfcInteger/IfcTimeStamp typed attribute (e.g. IfcOwnerHistory.CreationDate)
outside the signed 32-bit range corrupted the value instead of raising, since the
Python wrapper's set_attribute_value_py() truncated it with a plain static_cast<int>
before handing it to the C++ storage. Unix timestamps before 1901-12-13 or after
2038-01-19 silently wrapped around (e.g. 3000000000 became -1294967296) rather than
being rejected or stored correctly. Fixes#3058, equivalent to PR #8683 but ported to
this branch's rewritten ifcparse (snake_case files, variant_array/instance_data
storage, SWIG PyObject-based attribute setter) instead of the old IfcEntityInstanceData
sources, which no longer exist here.
The scalar slot of the attribute variant (Argument_INT) becomes int64_t. Integer
aggregates (Argument_AGGREGATE_OF_INT, e.g. CoordIndex) and instance/reference
identifiers stay 32-bit, since neither is the value that overflows here; this narrow
scope is kept on its own technical merits (aggregates and identifiers were never the
source of the bug, and widening them would be a much larger, riskier change for no
benefit) even though aothms said compatibility isn't a concern on this v0.9-track
branch. express::Base::set_attribute_value promotes the schema-generated int to
int64_t at a single choke point, so the generated setters keep compiling unchanged.
The STEP lexer, writer, and SWIG wrapper (set_attribute_value_py, pythonize) are all
widened together, since widening only the Python-facing setter would have silently
wrapped the value on file write instead of raising.
Verified in a build (IFC2X3 and IFC4, BUILD_IFCGEOM off, no kernels): pre-1901,
post-2038, both 32-bit boundaries, and a 9e12 value all round trip exactly both in
memory and through STEP text serialization (write then reopen). A value outside the
64-bit range now raises a clean exception instead of corrupting data. Ordinary
in-range integers and integer aggregates (e.g. IfcTriangulatedFaceSet.CoordIndex) are
unaffected. The existing util/test_attribute.py and test_file.py suites pass
unchanged; test_entity_instance.py has 5 pre-existing failures unrelated to this
change (confirmed identical on an unfixed build of this branch, caused by a missing
get_info_2 binding and _patch_swig_comparisons never being implemented here).
Generated with the assistance of an AI coding tool.
parse_num_ used std::from_chars for both integers and doubles, but the
floating-point from_chars overload is =deleted in Apple clang's libc++, so
the macOS build failed to compile (parse.cpp:136, instantiated for double).
Split parse_num_ with `if constexpr`: integers keep std::from_chars
everywhere; on macOS, doubles parse via strtod_l with a cached "C" locale
(locale-independent, restoring the pre-charconv Apple path). libstdc++ and
the MSVC STL have working float from_chars and are left unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
file.write() streamed directly onto the target path, so a crash mid-write
left a truncated file with dangling STEP references. Serialize to a temp
file in the same directory, then atomically rename it onto the target.
- New IfcUtil::path::atomic_rename_file: std::rename on POSIX, MoveFileExW
with MOVEFILE_REPLACE_EXISTING on Windows. Unlike rename_file it never
unlinks the destination first, so there is no window where it goes missing.
- Fully in C++/swig (per aothms), so the FILE_NAME header is untouched: it
comes from the model header, not the output path (verified empirically).
- Temp lives next to the target so the rename stays on one filesystem.
- Stream is closed before the rename (Windows cannot move an open file).
- On any write error the temp is removed and the original target is intact.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
std::to_chars is locale-independent, so the ostringstream and imbue(locale)
are no longer needed. Build the REAL string with plain std::string operations.
Output is unchanged (verified in standalone compile: same shortest values, all
round-trip). Addresses review feedback on #8309.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
format_double formatted doubles with setprecision(max_digits10) (17 digits),
which padded clean values with noise: 0.0174532925199433 was rewritten as
0.017453292519943299 and 1.E-05 as 1.0000000000000001E-05. Every REAL in a file
changed on save, producing enormous diffs for anyone version-controlling IFC.
Use std::to_chars, which emits the shortest string that round-trips exactly
(like Python's repr), then keep the existing mantissa/exponent formatting.
Verified in a standalone compile of the exact function logic: the reporter's
values become 0.0174532925199433 and 1.E-05, 0.1 stays 0.1, and every tested
value (including a denormal) round-trips back to the identical double.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
asStringRef removes the first and last characters of a string, enumeration
or binary token to drop the delimiters, guarded only by !str.empty(). A
malformed single-character token (e.g. a bare '.' left when a fuzzer turns
'.PHYSICAL.' into '.)HYSICAL.') has length 1, so the first erase empties the
string and the second erase(str.begin()) runs on an empty string. That is
undefined behaviour: benign on a normal build, but it aborts (or throws
std::length_error from a later append) under a hardened libstdc++ with
_GLIBCXX_ASSERTIONS, which is why this file only segfaulted on the Fedora
build. Require at least two characters before stripping.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
When ADD_COMMIT_SHA is off (the default for release tarballs), buildinfo.cpp
fell back to a hardcoded "0.8.0", so a 0.8.5/0.8.6 build reported 0.8.0 from
IfcConvert --version and in written file headers. Pass CMake's RELEASE_VERSION
(read from the VERSION file) to IfcParse as IFCOPENSHELL_VERSION_STRING and use
it as the fallback, mirroring how the branch/commit defines are handled. The
commit-sha build and the last-resort literal are unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Some distributions (e.g. Fedora) ship only a shared RocksDB that exports
RocksDB::rocksdb-shared rather than RocksDB::rocksdb. The CMake target
selection now falls back to the shared target when the static one is absent.
Newer RocksDB also changed DB::Open and DB::OpenForReadOnly to take
std::unique_ptr<DB>* instead of DB**. IfcFile.cpp uses SFINAE tag dispatch
to build against both old and new APIs without version detection.
Fixes builds with newer GCC/libstdc++ that no longer provide <cstdint>,
<cstring>, <cfloat>, <memory>, <algorithm> etc. transitively. Also
disambiguates visit<> calls in taxonomy.h with the full namespace and
casts the character value in IfcCharacterDecoder to uint32_t to silence
ambiguous overload warnings.
Some distributions (e.g. Fedora) ship only a shared RocksDB that exports
RocksDB::rocksdb-shared rather than RocksDB::rocksdb. The CMake target
selection now falls back to the shared target when the static one is absent.
Newer RocksDB also changed DB::Open and DB::OpenForReadOnly to take
std::unique_ptr<DB>* instead of DB**. IfcFile.cpp uses SFINAE tag dispatch
to build against both old and new APIs without version detection.
Fixes builds with newer GCC/libstdc++ that no longer provide <cstdint>,
<cstring>, <cfloat>, <memory>, <algorithm> etc. transitively. Also
disambiguates visit<> calls in taxonomy.h with the full namespace and
casts the character value in IfcCharacterDecoder to uint32_t to silence
ambiguous overload warnings.
# What
Two upstream-shaped fixes to ifcopenshell's runtime plug-in discovery
so the schema/serializer plug-ins are findable in deployment layouts
other than a flat \`<prefix>/lib/\` (specifically: macOS .app bundles).
## (1) Stable anchor variable instead of a function pointer
\`schema_plugin_directory()\` used \`&load_schema_plugins\` as the
anchor whose containing module \`dladdr\` is asked to resolve. Function
addresses are not reliably equal to a single canonical location across
toolchains — on macOS arm64 with BonsaiViewer.app, \`&load_schema_plugins\`
took the address of a PLT/stub inside the consumer binary rather than
the actual symbol inside \`libIfcParse.dylib\`. \`dladdr\` then dutifully
returned the consumer's path and the loader started searching
\`BonsaiViewer.app/Contents/MacOS/\` for plug-ins that were never
installed there.
A variable doesn't suffer from this — it has exactly one canonical
address inside its defining dylib. Add \`ifcopenshell_libifcparse_anchor\`
(exported via IFC_PARSE_API) and use \`&that\` instead. Standard pattern
used by Boost.DLL, GStreamer, \`_dyld_get_image_*\`, etc.
## (2) Bundle-aware fallback search paths
The primary search path is \`dirname(libIfcParse)\`. That works for
flat installs (Linux \`lib/\`, Windows \`bin/\`) where plug-ins are
siblings of libIfcParse. macOS app bundles split the layout:
\`macdeployqt\` puts non-Qt @rpath deps in \`Contents/Frameworks/\`,
Apple convention asks for \`Contents/PlugIns/\`, and some install rules
co-locate libIfcParse with the exe in \`Contents/MacOS/\`. Plug-ins
typically end up in a sibling directory, not the same one.
In \`add_search_paths_or_default\`, after registering the primary path,
also register \`<parent>/PlugIns\`, \`<parent>/Frameworks\`, and
\`<parent>/MacOS\` on Apple platforms. \`discover_exact\` short-circuits
on the first hit so duplicates and missing directories are harmless.
# What this does NOT do
The plug-in dylibs still need to actually be inside the app bundle
somewhere for these fallbacks to find them — the upstream install
rules (\`install(TARGETS …)\` in \`src/ifcparse/CMakeLists.txt\`,
\`src/serializers/CMakeLists.txt\`, etc.) put them in \`<prefix>/lib/\`
which lives outside \`BonsaiViewer.app\`. That side of the fix is a
follow-up — either an explicit bundle-aware install destination on
the plug-in targets, or an install(CODE) sweep that mirrors them
into the bundle.
# What this also un-does
Reverts the BonsaiViewer-only \`install(CODE)\` hack that was about to
copy \`ifcopenshell.*.dylib\` from \`lib/\` into
\`BonsaiViewer.app/Contents/MacOS/\` — superseded by the loader-side
fix above, which lets us put the plug-ins anywhere sane inside the
bundle without further consumer-side stitching.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Loading a federated project (.ifcfed) with several models segfaults
non-deterministically on a fresh start. The viewer's SceneLoader spawns
one background std::thread per model in startDataSourceLoad() to
construct an ifcopenshell::file; with cached sidecars all models reach
that point near-simultaneously, so multiple threads parse different IFC
files at once. Parsing touches the process-wide schema singleton, which
was not thread-safe in two places.
Race 1 — concurrent schema population
-------------------------------------
schema_registry::get() lazily runs the schema's get_() function (e.g.
Ifc4::get_schema() -> IFC4_populate_schema()) and mutates entries_ with
no lock. Two threads calling schema_by_name("IFC4") at once both run
IFC4_populate_schema() concurrently, which fills global arrays
(IFC4_types[], strings[]). One thread reads a slot the other is still
writing.
Core-dump evidence (gdb thread apply all bt):
Thread 1 SIGSEGV in IFC4_populate_schema Ifc4-schema.cpp:1989
<- Ifc4::get_schema
<- schema_registry::get schema.cpp:241
<- schema_by_name("IFC4")
<- ifcopenshell::file::file (NWCH-PIR-SS...ifc)
<- SceneLoader::startDataSourceLoad lambda SceneLoader.cpp:315
Thread 3 also in IFC4_populate_schema (entity ctor for
"IfcMaterialProfileSetUsageTapering")
<- Ifc4::get_schema
<- schema_registry::get schema.cpp:241
<- ifcopenshell::file::file (NWCH-PIR-PT...ifc)
<- SceneLoader::startDataSourceLoad lambda
Two threads inside IFC4_populate_schema() at the same time is the race.
Fix: guard schema_registry's bind()/get()/names()/clear() with a
recursive_mutex (recursive because get() re-enters bind() via
load_schema_plugin(), and a freshly populated schema registers itself
through register_schema()). get() is serialized, so only the first
thread populates the schema; the rest block briefly and then observe
the finished result. Returned schema pointers are stable for the
process lifetime, so holding the lock only across get() is sufficient.
Race 2 — lazy all_attributes_ cache filled during parsing
---------------------------------------------------------
entity::all_attributes() lazily fills a `mutable` optional cache on the
shared schema entity the first time it is accessed — and that first
access happens during parsing (parse_context::construct), not during
schema population. With race 1 fixed, two parser threads still raced
here: both saw the cache empty, both did all_attributes_.emplace() and
std::copy() into it, corrupting the vector.
Core-dump evidence after the race-1 fix:
Thread 1 SIGSEGV in attribute::type_of_attribute (this=0xe130...55c)
<- std::transform(first=0x4, last=0xb0d1...) <-- garbage
iterators into a corrupt std::vector
<- parse_context::construct over
decl->as_entity()->all_attributes() file.cpp:249
<- instance_streamer::read_instance
<- ifcopenshell::file::file (NWCH-PIR-PT...ifc)
<- SceneLoader::startDataSourceLoad lambda
The begin pointer 0x4 is a half-written vector being read mid-resize by
another thread.
Fix: force every entity's all_attributes_ cache in the
schema_definition constructor, while construction is still
single-threaded. The schema is then genuinely immutable after
construction, so concurrent parsing needs no hot-path lock.
Both crashes reproduce reliably on a fresh start at native speed but
vanish under gdb (which serializes thread scheduling) — the classic
signature of a data race. With both fixes the federated load completes
cleanly.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Read-only handles reject Flush/CompactRange, so the destructor's status
assertion always fired on shutdown when the streamer's sidecar was
opened with read_only=true. Track the flag and skip the write path; also
guard against a null db when the initial open failed.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>