ORCA-C

mirror of https://github.com/byteworksinc/ORCA-C.git synced 2024-12-21 16:29:31 +00:00

Author	SHA1	Message	Date
Stephen Heumann	05ecf5eef3	Add option to use the declared type for float/double/comp params. This differs from the usual ORCA/C behavior of treating all floating-point parameters as extended. With the option enabled, they will still be passed in the extended format, but will be converted to their declared type at the start of the function. This is needed for strict standards conformance, because you should be able to take the address of a parameter and get a usable pointer to its declared type. The difference in types can also affect the behavior of _Generic expressions. The implementation of this is based on ORCA/Pascal, which already did the same thing (unconditionally) with real/double/comp parameters.	2022-09-18 21:16:46 -05:00
Stephen Heumann	95ad02f0b9	Detect various errors in macro definitions. These changes detect violations of several constraints in C17 section 6.10.3 and subsections.	2022-07-28 20:49:22 -05:00
Stephen Heumann	6e3fca8b82	Implement strict type checking for enum types. If strict type checking is enabled, this will prohibit redefinition of enums, like: enum E {a,b,c}; enum E {x,y,z}; It also prohibits use of an "enum E" type specifier if the enum has not been previously declared (with its constants). These things were historically supported by ORCA/C, but they are prohibited by constraints in section 6.7.2.3 of C99 and later. (The C90 wording was different and less clear, but I think they were not intended to be valid there either.)	2022-07-19 20:35:44 -05:00
Stephen Heumann	8406921147	Parse command-line macros more consistently with macros in code. This makes a macro defined on the command line like -Dfoo=-1 consist of two tokens, the same as it would if defined in code. (Previously, it was just one token.) This also somewhat expands the set of macros accepted on the command line. A prefix of +, -, *, &, ~, or ! (the one-character unary operators) can now be used ahead of any identifier, number, or string. Empty macro definitions like -Dfoo= are also permitted.	2022-06-15 21:52:35 -05:00
Stephen Heumann	3c2b492618	Add support for compound literals within functions. The basic approach is to generate a single expression tree containing the code for the initialization plus the reference to the compound literal (or its address). The various subexpressions are joined together with pc_bno pcodes, similar to the code generated for the comma operator. The initializer expressions are placed in a balanced binary tree, so that it is not excessively deep. Note: Common subexpression elimination has poor performance for very large trees. This is not specific to compound literals, but compound literals for relatively large arrays can run into this issue. It will eventually complete and generate a correct program, but it may be quite slow. To avoid this, turn off CSE.	2022-06-08 21:34:12 -05:00
Stephen Heumann	58771ec71c	Do not do macro expansion after each ## operator is evaluated. It should only be done after all the ## operators in the macro have been evaluated, potentially merging together several tokens via successive ## operators. Here is an example illustrating the problem: #define merge(a,b,c) a##b##c #define foobar #define foobarbaz a int merge(foo,bar,baz) = 42; int main(void) { return a; }	2022-05-24 22:38:56 -05:00
Stephen Heumann	deca73d233	Properly expand macros that have the same name as a keyword or typedef. If such macros were used within other macros, they would generally not be expanded, due to the order in which operations were evaluated during preprocessing. This is actually an issue that was fixed by the changes from ORCA/C 2.1.0 to 2.1.1 B3, but then broken again by commit `d0b4b75970`. Here is an example with the name of a keyword: #define X long int #define long X x; int main(void) { return sizeof(x); /* should be sizeof(int) / } Here is an example with the name of a typedef: typedef short T; #define T long #define X T X x; int main(void) { return sizeof(x); / should be sizeof(long) */ }	2022-05-24 22:22:37 -05:00
Stephen Heumann	21f266c5df	Require use of digraphs in macro redefinitions to match the original. This is part of the general requirement that macro redefinitions be "identical" as defined in the standard. This affects code like: #define x [ #define x <:	2022-04-05 19:47:22 -05:00
Stephen Heumann	a1d57c4db3	Allow ORCA/C-specific keywords to be disabled via a new pragma. This allows those tokens (asm, comp, extended, pascal, and segment) to be used as identifiers, consistent with the C standards. A new pragma (#pragma extensions) is introduced to control this. It might also be used for other things in the future.	2022-03-26 18:45:47 -05:00
Stephen Heumann	b2edeb4ad1	Properly stringize tokens that start with a trigraph. This did not work correctly before, because such tokens were recorded as starting with the third character of the trigraph. Here is an example affected by this: #define mkstr(a) # a #include <stdio.h> int main(void) { puts(mkstr(??!)); puts(mkstr(??!??!)); puts(mkstr('??<')); puts(mkstr(+??!)); puts(mkstr(+??')); }	2022-03-25 18:10:13 -05:00
Stephen Heumann	f531f38463	Use suffixes on numeric constants in #pragma expand output. A suffix will now be printed on any integer constant with a type other than int, or any floating constant with a type other than double. This ensures that all constants have the correct types, and also serves as documentation of the types.	2022-03-01 19:46:14 -06:00
Stephen Heumann	182cf66754	Properly stringize tokens with line continuations or non-initial trigraphs. Previously, continuations or trigraphs would be included in the string as-is, which should not be the case because they are (conceptually) processed in earlier compilation phases. Initial trigraphs still do not get stringized properly, because the token starting position is not recorded correctly for them. This fixes code like the following: #define mkstr(a) # a #include <stdio.h> int main(void) { puts(mkstr(a\ bc)); puts(mkstr(qr\ )); puts(mkstr(\ xy)); puts(mkstr(12??/ 34)); puts(mkstr('??<')); }	2022-03-01 19:01:11 -06:00
Stephen Heumann	fec7b57ec2	Generate a string representation of tokens merged with ##. This is necessary for correct behavior if such tokens are subsequently stringized with #. Previously, only the first half of the token would be produced. Here is an example demonstrating the issue: #define mkstr(a) # a #define in_between(a) mkstr(a) #define joinstr(a,b) in_between(a ## b) #include <stdio.h> int main(void) { puts(joinstr(123,456)); puts(joinstr(abc,def)); puts(joinstr(dou,ble)); puts(joinstr(+,=)); puts(joinstr(:,>)); }	2022-02-22 18:48:34 -06:00
Stephen Heumann	6cfe8cc886	Remove an unused string representation of macro tokens. The string representation of macro tokens is needed for some preprocessor operations, but we get this in other ways (e.g. based on tokenStart/tokenEnd).	2022-02-21 18:39:39 -06:00
Stephen Heumann	8f27b8abdb	Print any ## tokens in #pragma expand output. Note that ## will not currently be recognized as a token in some contexts, leading to it not being printed.	2022-02-20 20:53:37 -06:00
Stephen Heumann	bf7a6fa5db	Use separate functions for merging tokens with ## and merging adjacent strings. These are conceptually separate operations occurring in different phases of the translation process. This change means that ## can no longer merge string constants: such operations will give an error about an illegal token. Cases like this are technically undefined behavior, so the old behavior could have been permitted, but it is clearer and more consistent with other compilers to treat this as an error.	2022-02-20 20:16:08 -06:00
Stephen Heumann	26e1bfc253	Allow generation of digraphs via ## token merging.	2022-02-20 18:57:03 -06:00
Stephen Heumann	2b062a8392	Make ## token merging on character constants give an error. This ultimately should be supported, but that will be more work. For now, we just set the string representation to '?', which will usually give an error when merged. (Previously, whatever was at memory location 0 would be treated as the string representation of the token. Frequently this would just be an empty string, leading to no error but incorrect results.)	2022-02-20 16:19:00 -06:00
Stephen Heumann	da978932bf	Save string representation of macros defined on command line. This is necessary for correct operation of the # and ## preprocessor operators on the tokens from such macros. Integers with a sign character still have the non-standard property of being treated as a single token, so they cannot be used with ##, but in most cases such uses will now give an error.	2022-02-20 15:35:49 -06:00
Stephen Heumann	aabbadb34b	Terminate header generation if #warning is encountered. This is necessary to ensure that the warning message is printed on subsequent compiles.	2022-02-19 14:06:15 -06:00
Stephen Heumann	a73dce103b	Terminate PCH generation if an #append is encountered. If the appended file was another C file and that file contained an #include, this would create an invalid record in the sym file. It would record memory from the buffer holding the original file to the buffer holding the appended file. In general, these are not contiguous, so superfluous data from other parts of memory would be included in the sym file. This record would normally just be treated as invalid on subsequent compiles, but it could theoretically be very large (depending on the memory layout) and might contain sensitive data from other parts of memory.	2022-02-19 14:05:07 -06:00
Stephen Heumann	f2d6625300	Save #pragma path directives in sym files. They were not being saved, which would result in ORCA/C not searching the proper paths when looking for an include file after the sym file had ended. Here is an example showing the problem: #pragma path "include" #include <stdio.h> int k = 50; #include "n.h" /* will not find include:n.h */	2022-02-15 21:27:35 -06:00
Stephen Heumann	3893db1346	Make sure #pragma expand is properly applied in all cases. There were various places where the flag for macro expansions was saved, set to false, and then later restored. If #pragma expand was used within those areas, it would not be properly applied. Here is an example showing that problem: void f(void #pragma expand 1 ) {} This could also affect some uses of #pragma expand within precompiled headers, e.g.: #pragma expand 1 #include "a.h" #undef foobar #include "b.h" ... Also, add a note saying that code in precompiled headers will not be expanded. (This has always been the case, but was not clearly documented.)	2022-02-15 20:50:02 -06:00
Stephen Heumann	c96cf4f1dd	Do not save predefined and command-line macros in the sym file. Previously, these might or might not be saved (based on the contents of uninitialized memory), but in many cases they were. This was unnecessary, since these macros are automatically defined when the scanner is initialized. Reading them from the sym file could result in duplicate copies of them in the macro list. This is usually harmless, but might result in #undefs of macros from the command line not working properly.	2022-02-13 20:17:33 -06:00
Stephen Heumann	b493dcb1da	Add lint check to require whitespace after names of object-like macros. This is a requirement added in C99, so it is added as part of the C99 syntax checks. This affects definitions like: #define foo;	2022-02-13 19:44:56 -06:00
Stephen Heumann	c169c2bf92	Fully prohibit redefinition of predefined macros. Code like the following was previously being allowed: #define __STDC__ /* no tokens */	2022-02-13 18:10:45 -06:00
Stephen Heumann	5d7c002819	Fix bug causing some #undefs to be ignored when using a sym file. This would occur if the macro had already been saved in the sym file and the #undef occurred before a subsequent #include that was also recorded in the sym file. The solution is simply to terminate sym file generation if an #undef of an already-saved macro is encountered. Here is an example showing the problem: test.c: #include "test1.h" #undef x #include "test2.h" int main(void) { #ifdef x return x; #else return y; #endif } test1.h: #define x 27 test2.h: #define y 6	2022-02-13 16:33:43 -06:00
Stephen Heumann	b231782442	Add option to use a custom pre-include file. This is a file that will be included before the source file is processed. If specified, it is used instead of the default .h file.	2022-02-12 21:36:39 -06:00
Stephen Heumann	bd811559d6	Fix issues with keep names in sym files. There were a couple issues that could occur with #pragma keep and sym files: If a source file used #pragma keep but it was overridden by KEEP= on the command line or {KeepName} in the shell, then the overriding keep name would be saved to the sym file. It would therefore be applied to subsequent compilations even if it was no longer specified in the command line or shell variable. If a source file used #pragma keep, that keep name would be recorded in the sym file. On subsequent compilations, it would always be used, overriding any keep name specified by the command line or shell, contrary to the usual rule that the name on the command line takes priority. With this patch, the keep name recorded in the sym file (if any) should always be the one specified by #pragma keep, but it can be overridden as usual.	2022-02-06 21:49:08 -06:00
Stephen Heumann	5f03dee66a	Allow negated long long constants in cc= defines. These are still treated as one token, like other negated numbers specified in cc=(-d...).	2022-02-06 15:33:42 -06:00
Stephen Heumann	efb363a04d	Update a comment.	2022-02-06 15:08:04 -06:00
Stephen Heumann	7d4f923470	Improve error handling for cc= options on command line.	2022-02-06 14:24:22 -06:00
Stephen Heumann	785a6997de	Record source file changes within a function as part of debug info. This affects functions whose body spans multiple files due to includes, or is treated as doing so due to #line directives. ORCA/C will now generate a COP 6 instruction to record each source file change, allowing debuggers to properly track the flow of execution across files.	2022-02-05 18:32:11 -06:00
Stephen Heumann	7322428e1d	Add an option to print file names in error messages. This can help identify if an error is in the main source file or an include file.	2022-02-04 22:10:50 -06:00
Stephen Heumann	4cb2106ee4	Change the name of the current source file on an #include or #append. This causes __FILE__ to give the name of an include file if used within it, which seems to be what the standards intend (and what other compilers do). It also affects the file name recorded in debugging information for functions declared in an include file. (Note that occ will generate a #line directive before an #append, essentially to work around the problem this patch fixes. After the patch, such a #line directive is effectively ignored. This should be OK, although it may result in a difference in whether a full or partial pathname is used for __FILE__ and in debug info.)	2022-02-03 22:22:33 -06:00
Stephen Heumann	dce9d36edd	Comment out unused error messages and update docs about errors.	2022-02-01 22:16:57 -06:00
Stephen Heumann	b43036409e	Add a new optimize flag for FP math optimizations that break IEEE rules. There were several existing optimizations that could change behavior in ways that violated the IEEE standard with regard to infinities, NaNs, or signed zeros. They are now gated behind a new #pragma optimize flag. This change allows intermediate code peephole optimization and common subexpression elimination to be used while maintaining IEEE conformance, but also keeps the rule-breaking optimizations available if desired. See section F.9.2 of recent C standards for a discussion of how these optimizations violate IEEE rules.	2021-11-29 20:31:15 -06:00
Stephen Heumann	73d194c12f	Allow string constants with up to 32760 bytes. This allows the length of the string plus a few extra bytes used internally to be represented by a 16-bit integer. Since the size limit for memory allocations has been raised, there is no good reason to impose a shorter limit on strings. Note that C99 and later specify a minimum translation limit for string constants of at least 4095 characters.	2021-10-24 21:43:43 -05:00
Stephen Heumann	f567d60429	Allow bit-fields in unions. All versions of standard C allow this, but ORCA/C previously did not.	2021-10-18 21:48:18 -05:00
Stephen Heumann	692ebaba85	Structs or arrays may not contain structs with a flexible array member. We previously ignored this, but it is a constraint violation under the C standards, so it should be reported as an error. GCC and Clang allow this as an extension, as we were effectively doing previously. We will follow the standards for now, but if there was demand for such an extension in ORCA/C, it could be re-introduced subject to a #pragma ignore flag.	2021-10-17 22:22:42 -05:00
Stephen Heumann	ad5063a9a3	Support hexadecimal floating-point constants.	2021-10-17 18:19:29 -05:00
Stephen Heumann	5871820e0c	Support UTF-8/16/32 string literals and character constants (C11). These have u8, u, or U prefixes, respectively. The types char16_t and char32_t (defined in <uchar.h>) are used for UTF-16 and UTF-32 code points.	2021-10-11 20:54:37 -05:00
Stephen Heumann	b076f85149	Avoid possible stack overflow when merging adjacent string literals. The code for this was recursive and could overflow if there were several dozen consecutive string literals. It has been changed to only use one level of recursion, avoiding the problem.	2021-10-11 18:55:10 -05:00
Stephen Heumann	7ae830ae7e	Initial support for compound literals. Compound literals outside of functions should work at this point. Compound literals inside of functions are not fully implemented, so they are disabled for now. (There is some code to support them, but the code to actually initialize them at the appropriate time is not written yet.)	2021-09-16 18:34:55 -05:00
Stephen Heumann	a8682e28d3	Give an error for pointer assignments that discard qualifiers. This is controlled by #pragma ignore bit 5, which is now a more general "loose type checks" bit.	2021-09-10 17:58:20 -05:00
Stephen Heumann	9c04b94093	Allow invalid escape sequences and UCN-like sequences in skipped code. The standard wording is not always clear on these cases, but I think at least some of them should be allowed and others may be undefined behavior (which we can choose to allow). At any rate, this allows non-standard escape sequences targeted at other compilers to appear in skipped-over code. There probably ought to be similar handling for #defines that are never expanded, but that would require more code changes.	2021-09-06 20:37:17 -05:00
Stephen Heumann	ea461dba7b	Give clearer error messages for errors in the command line.	2021-08-31 19:23:10 -05:00
Stephen Heumann	b8c332deeb	Treat invalid escape sequences as errors. This applies to octal and hexadecimal sequences with out-of-range values, and also to unrecognized escape characters. The C standards say both of these cases are syntax/constraint violations requiring a diagnostic.	2021-08-31 18:36:06 -05:00
Stephen Heumann	2b9d332580	Give an appropriate error for an illegal operator in a constant expression. This was being reported as an "illegal type cast".	2021-08-22 20:33:34 -05:00
Stephen Heumann	5faf219eff	Update comments about pragma flags.	2021-08-22 17:35:16 -05:00

1 2 3

148 Commits